User portrait label generation method and system based on time-aligned transfer model

By building a state transfer probability matrix of user behavior based on the time-horizontal transfer model, a dynamically updated user portrait tag is solved, and the problem of insufficient dynamicity of user behavior relationships in the prior art is solved, and a more accurate and scientific user portrait generation is achieved.

CN119989032AActive Publication Date: 2025-05-13CHINA RONGXIN CLOUD TECH CO LTD

Patent Information

Application Number
CN202510466876.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-15
Publication Date
2025-05-13
Estimated Expiration
2045-04-15

AI Technical Summary

Technical Problem

The prior art fails to fully explore the complex relationship between user behavior operation types when generating user portraits, and fails to capture dynamic changes in user behavior in a timely manner, resulting in the generated portrait labels being too one-sided and cannot accurately reflect the real patterns and dynamic changes of user behavior.

Method used

Using a method based on the time-horizontal transfer model, a state transition probability matrix is ​​constructed by obtaining the target user's user behavior trajectory data, determining the absorbed state behavior operation type set, generating a picture tag set, and dynamically updating the picture tag sequence.

Benefits of technology

This method can fully capture the timing correlation between user behaviors, accurately describe the user's stable characteristics in the specified behavior dimension, avoid misjudgments caused by short-term behavior fluctuations in users, and improve the scientificity and reliability of the image tag generation process.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119989032A_ABST
    Figure CN119989032A_ABST
Patent Text Reader

Abstract

The invention provides a user portrait label generation method and system based on a time-aligned transition model, and the method comprises the steps: firstly obtaining behavior track data, containing behavior operation types and context attributes, of a target user in a preset time period, and then constructing a state transition probability matrix according to the time sequence relevance of the behavior operation types; then determining an absorption state behavior operation type set meeting specific conditions according to the steady-state convergence of a matrix state node, generating a portrait label set based on the absorption state behavior operation type set and corresponding context attributes, and describing stable characteristics of a target user in a specified behavior dimension; finally, matching verification is conducted on the portrait label set and a preset knowledge base, the dynamically-updated portrait label sequence is output, portrait labels reflecting stable features of the user can be comprehensively and accurately generated, and the portrait quality of the user is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of digital service technology, and in particular to a method and system for generating user portrait labels based on a time-aligned transfer model. Background Art

[0002] In today's digital age, user portrait technology plays a vital role in helping companies gain a deeper understanding of user needs, provide personalized services, and develop precise marketing strategies. By building user portraits, companies can abstract and quantify complex and diverse user behaviors and characteristics, thereby better understanding user behavior patterns and preferences.

[0003] However, many existing methods simply collect and count user behavior data, and fail to fully explore the potential complex relationships between user behavior operation types. For example, they only focus on the frequency of a single behavior, while ignoring the correlation between different behaviors in time series, making the generated user portrait too one-sided and unable to accurately reflect the real pattern and dynamic changes of user behavior, resulting in an in-depth and incomplete understanding of user behavior.

[0004] In addition, most existing technologies do not take into account the dynamic evolution of user behavior over time. User behavior patterns are not static, but will change with time, environment and other factors. However, traditional methods often use static methods to generate portrait tags, which cannot capture the dynamic changes of user behavior in a timely manner, causing the generated portrait tags to lose their accurate description of the user's real behavior after a certain period of time, reducing the effectiveness and value of user portraits in practical applications. Summary of the invention

[0005] In view of the above-mentioned problems, in combination with the first aspect of the present invention, an embodiment of the present invention provides a method for generating user portrait labels based on a time-aligned transfer model, the method comprising: Acquire user behavior trajectory data of a target user within a preset time period, wherein the user behavior trajectory data includes behavior operation types and corresponding context attributes at multiple discrete time points; Based on the temporal correlation of each behavior operation type in the user behavior trajectory data, a state transition probability matrix corresponding to the target user is constructed, wherein each element in the state transition probability matrix represents a transition probability of the target user migrating from a first behavior operation type to a second behavior operation type; Determine, according to the steady-state convergence of each state node in the state transition probability matrix, a set of absorbing state behavior operation types corresponding to the target user, each behavior operation type in the set of absorbing state behavior operation types satisfies a state transition probability threshold condition within a continuous time period; Based on the behavior operation types in the absorbing state behavior operation type set and the corresponding context attributes, a portrait tag set of the target user is generated, wherein each portrait tag in the portrait tag set is used to describe the stable characteristics of the target user in a specified behavior dimension; After matching and verifying the portrait tag set with a preset portrait tag knowledge base, a dynamically updated portrait tag sequence corresponding to the target user is output.

[0006] On the other hand, an embodiment of the present invention also provides a user portrait label generation system based on a time-aligned transfer model, including a processor and a machine-readable storage medium, wherein the machine-readable storage medium is connected to the processor, the machine-readable storage medium is used to store programs, instructions or codes, and the processor is used to execute the programs, instructions or codes in the machine-readable storage medium to implement the above method.

[0007] Based on the above aspects, the embodiment of the present application can fully capture the temporal correlation between user behaviors by acquiring the detailed behavior trajectory data of the target user within a preset time period, including the behavior operation type and its contextual attributes, and constructing a state transition probability matrix based on a time-homogeneous transition model. On this basis, the set of absorbing state behavior operation types is determined according to the state transition probability matrix, and then a portrait label set is generated, which can accurately describe the stable characteristics of the user in the specified behavior dimension, and effectively avoid the misjudgment caused by the short-term behavior fluctuation of the user. Next, the user behavior data is converted into a state transition probability matrix using a time-homogeneous transition model, and the absorbing state behavior operation type is determined by analyzing the steady-state convergence of the state nodes in the state transition probability matrix, which greatly improves the scientificity and reliability of the portrait label generation process, so that the generated portrait label can more accurately reflect the user's real behavior characteristics. Finally, after the portrait label set is generated, it is matched and verified with the preset portrait label knowledge base, and then the dynamically updated portrait label sequence is output, which ensures that the portrait label can not only reflect the latest behavior characteristics of the user at the moment, but also conform to the existing knowledge base standards and knowledge system. On the one hand, the dynamic update mechanism enables portrait labels to be adjusted in a timely manner as user behavior changes, always maintaining real-time tracking and accurate characterization of user behavior; on the other hand, the matching verification process ensures the consistency and accuracy of the generated portrait labels within the entire system knowledge framework, improving the versatility and usability of portrait labels in different application scenarios. BRIEF DESCRIPTION OF THE DRAWINGS

[0008] Figure 1 It is a schematic diagram of the execution flow of the method for generating user portrait labels based on the time-aligned transfer model provided in an embodiment of the present invention.

[0009] Figure 2It is a schematic diagram of exemplary hardware and software components of a user portrait label generation system based on a time-aligned transfer model provided in an embodiment of the present invention. DETAILED DESCRIPTION

[0010] The present invention will be described in detail below with reference to the accompanying drawings. Figure 1 It is a flow chart of a method for generating user portrait labels based on a time-aligned transfer model provided by an embodiment of the present invention. The method for generating user portrait labels based on a time-aligned transfer model is introduced in detail below.

[0011] Step S110, obtaining user behavior trajectory data of a target user within a preset time period, wherein the user behavior trajectory data includes behavior operation types and corresponding context attributes at multiple discrete time points.

[0012] For example, taking a financial service platform as an example, financial institutions usually record various user operations. Assume that the preset time period is the past year and the target user is a credit card user.

[0013] First, the original behavioral event sequence can be extracted from the user's credit card interaction log. For example, the event trigger timestamp may include the specific date and time of each transaction the user makes with the credit card, such as 10:30:00 on March 15, 2022; the event operation type identifier can be the transaction type, such as consumption, repayment, cash withdrawal, etc.; the event context parameter can include the transaction amount, transaction location (such as the name or address code of a shopping mall), the industry to which the transaction merchant belongs (such as catering, retail, etc.), etc.

[0014] Then, the original behavior event sequence is cleaned. If an event record with a timestamp earlier than the system record start time (assuming that the system starts recording from January 1, 2020) or later than the current system time (assuming that the current time is July 2023) is detected, it is marked as an abnormal timestamp event, for example, there may be an erroneous time record caused by a system failure. Identify event records with empty event operation type identifiers (possibly due to data transmission errors, some transaction records do not have a clear operation type) or do not meet the preset coding specifications (such as tampering with or incorrect entry of the code), and mark them as invalid operation type events. Traverse every two consecutive event records. If the timestamp interval between the two consecutive event records is less than the preset minimum operation interval (for example, for credit card transactions, the interval between two transactions will not be less than 1 minute under normal circumstances, but if it is less than 1 minute, it may be a duplicate record), it is marked as a repeated trigger event. These abnormal timestamp events, invalid operation type events, and repeated trigger events are removed from the original behavior event sequence to generate an intermediate cleaning event sequence. Then, the integrity of the remaining event records in the intermediate cleaning event sequence is checked. For example, if a transaction record is found to be missing the context parameter of the transaction amount, its default value is filled in (it can be filled in according to the average transaction amount of the user or the common amount of similar transactions) to generate a post-cleaning behavior event sequence.

[0015] The cleaned behavioral event sequence is segmented according to the preset time window granularity. Assuming that according to the distribution density of the event trigger timestamps in the cleaned behavioral event sequence, during the peak period of credit card use (such as holidays or promotional seasons), the number of event triggers exceeds the density threshold (assuming 50 transactions per day), the time window granularity is reduced to the first sub-granularity (for example, from the original weekly window to the daily window); during the low period of credit card use (such as the non-peak period of weekdays), if the number of event triggers is lower than the sparse threshold (assuming 5 transactions per day), the time window granularity is expanded to the second sub-granularity (for example, from the daily window to the two-week window). Based on the adjusted time window granularity, the cleaned behavioral event sequence is divided into multiple continuous or partially overlapping time window intervals. The frequency of the event operation type identifier in each time window interval is counted, and the operation type with the highest frequency is taken as the main behavior operation type. For example, if the frequency of consumption is the highest in a certain time window, then consumption is the main behavior operation type of the time window.

[0016] Perform cluster analysis on the event operation type identifier in each behavior subsequence. Convert the behavior subsequence corresponding to each time window into a sequence vector of the event operation type identifier, for example, [consumption, repayment, consumption, withdrawal, repayment] is converted into a corresponding vector representation. Use a preset clustering model (such as the K-means clustering algorithm) to perform unsupervised clustering on the sequence vector to obtain multiple candidate behavior clusters. Calculate the similarity between the center vector of each candidate behavior cluster and the preset benchmark behavior pattern vector (such as the normal credit card usage behavior pattern vector, including the proportion of consumption, repayment, withdrawal, etc.). Determine the candidate behavior cluster with the highest similarity as the main behavior cluster, and use the operation type corresponding to the main behavior cluster as the main behavior operation type. Add the types of operation types in other candidate behavior clusters whose occurrence frequency is greater than the auxiliary frequency threshold (assuming 3 times) to the auxiliary behavior operation type set. For example, in addition to consumption as the main behavior operation type, if the frequency of withdrawal in a candidate behavior cluster is greater than 3 times, then withdraw will be added to the auxiliary behavior operation type set. Finally, based on the event context parameters, the main behavior operation type and the corresponding auxiliary behavior operation type set, the behavior operation type and context attributes corresponding to each discrete time point in the user behavior trajectory data are generated. For example, the behavior operation type at a discrete time point is consumption, the context attribute is the transaction amount of 500 yuan, the transaction location is a large shopping mall, and the industry to which the transaction merchant belongs is retail.

[0017] Step S120, based on the temporal correlation of each behavior operation type in the user behavior trajectory data, construct a state transition probability matrix corresponding to the target user, each element in the state transition probability matrix represents the transition probability of the target user migrating from the first behavior operation type to the second behavior operation type.

[0018] In detail, the state space dimension of the state transition probability matrix can be defined based on the unique identifier set of all behavior operation types in the user behavior trajectory data, such as consumption (identified as C), repayment (identified as R), cash withdrawal (identified as W), etc. Here, the state transition probability matrix is ​​a 3×3 matrix (assuming that there are only these three behavior operation types), and the rows and columns correspond to the three behavior operation types respectively.

[0019] Align the discrete time points in the user behavior trajectory data in chronological order to generate the behavior state transition chain of the target user. For example, the sequence of behavior operation types recorded in chronological order is [C, R, C, W, R], which constitutes the behavior state transition chain. Count the number of migrations between every two adjacent behavior operation types in the behavior state transition chain to generate the initial state transition number matrix. Assume that in the entire behavior state transition chain, the number of migrations from consumption (C) to repayment (R) is 5 times, the number of migrations from consumption (C) to cash withdrawal (W) is 3 times, the number of migrations from repayment (R) to consumption (C) is 4 times, etc., and construct the initial state transition number matrix based on these data.

[0020] Normalize the initial state transition times matrix. Traverse each row of elements in the initial state transition times matrix and calculate the sum of all elements in the row. Assume that the sum of elements in a row (corresponding to consumption C) is 8 (3 times to repay + 5 times to withdraw cash), compare the sum with the preset minimum row sum threshold (assuming it is 5). If the sum is less than the preset minimum row sum threshold, add a pseudo-count compensation value to the row (assuming the compensation value is 2) to obtain the compensated elements in each row. Normalize the probability of each row element after compensation so that the sum of all elements in the row is 1, and obtain the normalized probability value. Round the normalized probability value to the preset decimal precision (assuming two decimal places) to generate a standardized transition probability matrix. Replace the elements in the standardized transition probability matrix that are lower than the preset probability lower limit (assuming it is 0.05) with the lower limit value to generate the final state transition probability matrix. For example, in this 3×3 matrix, the transition probability corresponding to the element (C, R) is 0.38, the transition probability corresponding to the element (C, W) is 0.62, etc., indicating the transition probability from consumption (C) to repayment (R) and cash withdrawal (W). Replace the elements with zero transition probability values ​​in the state transition probability matrix with the preset minimum probability value (assuming it is 0.01) to generate the target state transition probability matrix after smoothing to avoid calculation problems caused by zero probability in subsequent calculations.

[0021] Step S130, determining an absorbing state behavior operation type set corresponding to the target user according to the steady-state convergence of each state node in the state transition probability matrix, wherein each behavior operation type in the absorbing state behavior operation type set satisfies a state transition probability threshold condition within a continuous time period.

[0022] In detail, the state transition probability matrix can be calculated by power iteration. Assume that the initial state transition probability matrix is:

[0023]

[0024] The above initial state transition probability matrix is ​​subjected to power iteration calculation until the maximum difference of each row element in the state transition probability matrix is ​​less than the preset convergence threshold (assuming it is 0.001), and the steady-state distribution vector is obtained. After multiple iterations, the steady-state distribution vector is [0.33, 0.33, 0.34].

[0025] Filter out the state nodes whose probability values ​​are greater than the preset absorbing state threshold (assuming it is 0.3) from the steady-state distribution vector to generate a set of candidate absorbing state behavior operation types. Here, the probability values ​​corresponding to consumption (C) and cash withdrawal (W) meet the conditions and enter the candidate absorbing state behavior operation type set.

[0026] For each behavior operation type in the candidate absorbing behavior operation type set, verify whether the state self-transition probability of the behavior operation type is continuously greater than the preset stability threshold (assuming it is 0.6) within a preset number of consecutive time periods (assuming it is 3 consecutive months). Assume that for consumption (C), the self-transition probabilities of each month in these 3 months are 0.7, 0.65, and 0.62 respectively, which meets the conditions; for cash withdrawal (W), the self-transition probabilities are 0.5, 0.55, and 0.58 respectively, which does not meet the conditions. The behavior operation types that are continuously greater than the preset stability threshold (here only consumption C) are merged into the absorbing behavior operation type set, and the behavior operation types that have not passed the verification in the candidate absorbing behavior operation type set (here cash withdrawal W) are deleted.

[0027] The self-transition probability of the state node corresponding to each behavior operation type in the set of absorbing behavior operation types (here is consumption C) is increased by the preset reinforcement coefficient (assuming it is 1.2), the transition probability value from the absorbing behavior operation type to the non-absorbing behavior operation type is reduced, and the difference is evenly distributed to the migration path of other absorbing behavior operation types (here there is only consumption C, no other absorbing behavior operation types, so this step is just a self-reinforcement in this example). The adjusted state transition probability matrix is ​​obtained. For example, the original transition probability from consumption (C) to repayment (R) is 0.3, which is adjusted to 0.2 (reduced by 0.1), and the transition probability from consumption (C) to itself (C) is changed from 0.5 to 0.6 (increased by 0.1). The adjusted state transition probability matrix is ​​re-normalized so that the sum of the elements in each column is 1, and the re-normalized state transition probability matrix is ​​obtained. The re-normalized state transition probability matrix is ​​input as the updated state transition probability matrix into the portrait label generation process of the next time period. Record the update history of the absorbing behavior operation type set for dynamic correction of subsequent portrait label sets.

[0028] Step S140, based on the behavior operation types in the absorbing behavior operation type set and the corresponding context attributes, generate a portrait tag set of the target user, each portrait tag in the portrait tag set is used to describe the stable characteristics of the target user in a specified behavior dimension.

[0029] In detail, the context attribute feature vector corresponding to each behavior operation type can be extracted from the set of absorbing behavior operation types (here, consumption C). Assume that the operation frequency of consumption (C) is 10 times per month, the operation duration (here, it can be understood as the average processing time for each consumption) is 2 minutes, and the associated resource identifier is the bank account A to which the credit card belongs.

[0030] Match the context attribute feature vector with the label definition rules in the preset portrait label knowledge base. In the field of default risk in financial services, the portrait label knowledge base may have label definition rules for consumption behavior, such as high-frequency consumption (consumption frequency greater than 8 times per month), short-term consumption (consumption duration less than 3 minutes each time), and association with specific accounts (the associated account is bank account A). Determine the candidate portrait label list corresponding to each behavior operation type, which may include labels such as high-frequency consumption, short-term consumption, and association with specific accounts.

[0031] Perform semantic aggregation on the tags in the candidate profile tag list and merge multiple sub-tags that describe the same behavior dimension. For example, high-frequency consumption and short-duration consumption both describe the time dimension-related characteristics of consumption and can be merged into the aggregated tag of high-frequency short-duration consumption. Generate an aggregated tag set.

[0032] Based on the frequency of occurrence of each portrait tag in the aggregated tag set (here, it is assumed that the high-frequency short-duration consumption tag has appeared 5 times in the past analysis) and the time distribution density (for example, evenly distributed in the past year), the confidence weight of each portrait tag is calculated. Assume that according to a specific calculation rule (such as weighted calculation based on the frequency of occurrence and time distribution density), the confidence weight of the high-frequency short-duration consumption tag is 0.8.

[0033] The aggregated tag set is sorted and filtered according to the confidence weight, and tags with weight values ​​greater than the preset threshold (assuming 0.6) are retained to generate a portrait tag set. Here, the weight of the high-frequency short-duration consumption tag 0.8 is greater than 0.6, so this tag is included in the portrait tag set.

[0034] Step S150, after matching and verifying the portrait tag set with a preset portrait tag knowledge base, output a dynamically updated portrait tag sequence corresponding to the target user.

[0035] In detail, the portrait tag set (here, high-frequency, short-duration consumption) can be matched and verified with the preset portrait tag knowledge base. In the field of default risk of financial services, the portrait tag knowledge base may contain risk assessment information corresponding to different tag combinations. For example, high-frequency, short-duration consumption tags may be associated with a certain default risk, which may indicate that the user's consumption habits are more impulsive or there are potential capital turnover problems. After matching and verification, the dynamically updated portrait tag sequence corresponding to the target user is output. The portrait tag sequence may be dynamically updated with the user's subsequent behavior. For example, if the user changes his consumption habits in the future, reduces the frequency of consumption or increases the duration of consumption, then the portrait tag sequence will be regenerated and updated according to the new behavior data to accurately reflect the user's characteristics in terms of financial service default risk.

[0036] Based on the above steps, the embodiment of the present application can fully capture the temporal correlation between user behaviors by acquiring the detailed behavior trajectory data of the target user within a preset time period, including the behavior operation type and its contextual attributes, and constructing a state transition probability matrix based on the time-homogeneous transition model. On this basis, the set of absorbing state behavior operation types is determined according to the state transition probability matrix, and then a portrait label set is generated, which can accurately describe the stable characteristics of the user in the specified behavior dimension, and effectively avoid the misjudgment caused by the short-term behavior fluctuations of the user. Next, the user behavior data is converted into a state transition probability matrix using the time-homogeneous transition model, and the absorbing state behavior operation type is determined by analyzing the steady-state convergence of the state nodes in the state transition probability matrix, which greatly improves the scientificity and reliability of the portrait label generation process, so that the generated portrait label can more accurately reflect the real behavior characteristics of the user. Finally, after the portrait label set is generated, it is matched and verified with the preset portrait label knowledge base, and then the dynamically updated portrait label sequence is output, which ensures that the portrait label can not only reflect the latest behavior characteristics of the user at the moment, but also conform to the existing knowledge base standards and knowledge system. On the one hand, the dynamic update mechanism enables portrait labels to be adjusted in a timely manner as user behavior changes, always maintaining real-time tracking and accurate characterization of user behavior; on the other hand, the matching verification process ensures the consistency and accuracy of the generated portrait labels within the entire system knowledge framework, improving the versatility and usability of portrait labels in different application scenarios.

[0037] In a possible implementation, step S110 includes: Step S111 : extracting an original behavior event sequence from the target user's interaction log, wherein the original behavior event sequence includes an event triggering timestamp, an event operation type identifier, and event context parameters.

[0038] In detail, the event trigger timestamp accurately records the time of each user's credit card operation, such as a transaction at 13:15:00 on May 10, 2022, and a repayment operation at 15:30:00 on May 15, 2022; the event operation type identifier clarifies the type of operation, such as consumption, repayment, cash withdrawal, etc.; the event context parameters contain rich information, such as the transaction amount is 500 yuan when consuming, the transaction location is a large shopping mall, the address code of the shopping mall in the credit card system is 1234, the industry of the transaction merchant is retail, and the repayment amount during the repayment operation is 800 yuan.

[0039] Step S112, performing data cleaning on the original behavior event sequence, removing duplicate event records and event records with abnormal timestamps, and obtaining a cleaned behavior event sequence.

[0040] In detail, for timestamps, check whether there are event records that are earlier than the system record start time (such as the system starts recording on January 1, 2020) or later than the current system time (assuming the current time is July 2023). If so, mark them as abnormal timestamp events. For example, due to system failure, there may be transaction records recorded on December 30, 2019, which obviously does not meet the requirements. For event operation type identifiers, if they are empty (it may be that part of the data is lost during data transmission, resulting in a transaction without an operation type identifier) ​​or event records that do not meet the preset coding specifications (such as the code is maliciously tampered with or entered in the wrong format), mark them as invalid operation type events. Then traverse every two consecutive event records. If the time stamp interval between the two is less than the preset minimum operation interval (the normal transaction interval of a credit card is at least 1 minute, if it is less than 1 minute, it may be a duplicate record), for example, if two transaction records with the same amount and the same merchant are found at 13:15:00 and 13:15:10, mark them as duplicate trigger events. These abnormal timestamp events, invalid operation type events, and repeated triggering events are removed from the original behavior event sequence to obtain the intermediate cleaning event sequence. After that, the integrity of the remaining event records in the intermediate cleaning event sequence is checked. If the transaction amount of a transaction record is missing, the default value is filled by querying the average amount of the same type of transactions of the user or the common amount of transactions of the same merchant, thereby generating a post-cleaning behavior event sequence.

[0041] Step S113, segmenting the cleaned behavior event sequence according to a preset time window granularity to generate behavior subsequences corresponding to multiple time windows, wherein the event triggering timestamp in each of the behavior subsequences satisfies the time range constraint of the time window.

[0042] In detail, for example, during the credit card promotion season (such as November-December 2022), transactions are frequent. If the number of event triggers exceeds the density threshold (assuming 60 transactions per day), the time window granularity is reduced to the first sub-granularity, and the window is changed from being divided by week to being divided by day; and during non-peak hours on weekdays (such as some weekdays in March 2023), transactions are sparse. If the number of event triggers is lower than the sparse threshold (assuming 3 transactions per day), the time window granularity is expanded to the second sub-granularity, and the window is changed from being divided by day to being divided by two weeks. Based on the adjusted time window granularity, the cleaned behavioral event sequence is divided into multiple continuous or partially overlapping time window intervals. The frequency statistics of the event operation type identifier in each time window interval are performed to determine the main behavior operation type. For example, there are 20 transactions in a time window, including 12 consumption transactions, 5 repayment transactions, and 3 cash withdrawal transactions, then consumption is the main behavior operation type of the time window.

[0043] Step S114: performing cluster analysis on the event operation type identifiers in each of the behavior subsequences to obtain a set of main behavior operation types and auxiliary behavior operation types corresponding to the time window.

[0044] Specifically, the behavior subsequence corresponding to each time window is first converted into a sequence vector of event operation type identification, for example, the operation sequence in a time window is [consumption, repayment, consumption, cash withdrawal, repayment], which is converted into a corresponding vector representation. The sequence vector is unsupervised clustered using a preset clustering model (such as a hierarchical clustering algorithm) to obtain multiple candidate behavior clusters. The similarity between the center vector of each candidate behavior cluster and the preset benchmark behavior pattern vector (such as the proportional relationship vector of consumption, repayment, and cash withdrawal under the normal use mode of a credit card) is calculated. Assuming that there are three candidate behavior clusters, the first candidate behavior cluster is calculated to have the highest similarity with the benchmark behavior pattern vector, and it is determined to be the main behavior cluster, and its corresponding operation type consumption is the main behavior operation type of the time window. For other candidate behavior clusters, if the frequency of occurrence of the operation type therein is greater than the auxiliary frequency threshold (assuming 3 times), then it is added to the auxiliary behavior operation type set. For example, if the frequency of cash withdrawal in a candidate behavior cluster is greater than 3 times, then cash withdrawal is added to the auxiliary behavior operation type set.

[0045] Step S115, based on the event context parameters and the main behavior operation type and the corresponding auxiliary behavior operation type set, generate the behavior operation type and context attributes corresponding to each discrete time point in the user behavior trajectory data.

[0046] For example, the main behavior operation type corresponding to a discrete time point is consumption, the transaction amount is 300 yuan, the transaction location is a supermarket, its address code is 5678, and the industry is retail. The auxiliary behavior operation type set includes cash withdrawal, and the average amount of cash withdrawal is 200 yuan (obtained by counting the cash withdrawal operations in the time window), etc., thereby generating the behavior operation type and context attributes corresponding to each discrete time point.

[0047] In a possible implementation, step S112 includes: Step S1121 , detecting whether there is an event record in the original behavior event sequence whose timestamp is earlier than the system record start time or later than the current system time, and marking it as an abnormal timestamp event.

[0048] In the credit card system, the original behavior event sequence is extracted from the target user's interaction log. The original behavior event sequence contains a lot of key information, such as the event trigger timestamp that accurately records the time when each credit card operation occurs, the event operation type identifier that clarifies the type of operation, and the event context parameter that contains additional information related to the operation.

[0049] In detail, the credit card system has clear time range regulations for credit card operation records. The starting time of the credit card system record is set to January 1, 2020, and the current system time is assumed to be July 2023. In the original behavior event sequence, the timestamps of the events are checked one by one. When a transaction record with a timestamp of December 30, 2019 is found, since this time is earlier than the starting time of the credit card system record, according to the rules, this event record is marked as an abnormal timestamp event. Similarly, if there is a record with a timestamp of August 1, 2023 (assumed), it is also marked as an abnormal timestamp event because it is later than the current system time.

[0050] Step S1122, identifying event records in the original behavior event sequence whose event operation type identifier is empty or does not comply with a preset coding specification, and marking them as invalid operation type events.

[0051] For example, you may find that the operation type identifier in a transaction record is empty, which may be caused by partial loss or writing errors during data transmission. In addition, the credit card operation type identifier has a preset coding specification, such as consumption operation code 1, repayment is 2, cash withdrawal is 3, etc. If the operation type identifier in a record is found to be 5 (not in accordance with the preset coding specification), then this event record is marked as an invalid operation type event.

[0052] Step S1123, traversing every two consecutive event records in the original behavior event sequence, if the time stamp interval between the two consecutive event records is less than a preset minimum operation interval, marking them as repeated triggering events.

[0053] In detail, the normal operation of a credit card has certain time interval requirements, assuming that the preset minimum operation interval is 1 minute. For example, there are two consecutive transaction records. The first transaction occurred at 13:15:00 on May 10, 2022, with an amount of 500 yuan. The second transaction occurred at 13:15:10 on May 10, 2022, with an amount of 500 yuan. The transaction location, merchant and other context parameters are exactly the same. This indicates that the two events are likely to be triggered repeatedly, so they are marked as repeated triggering events.

[0054] Step S1124: remove the abnormal timestamp events, invalid operation type events and repeated triggering events from the original behavior event sequence to generate an intermediate cleaning event sequence.

[0055] For example, the original behavior event sequence has a total of 100 records. After inspection, 5 abnormal timestamp events, 3 invalid operation type events and 2 repeated trigger events are found. After removing these 10 records, the remaining 90 records constitute the intermediate cleaning event sequence.

[0056] Step S1125 , performing integrity check on the remaining event records in the intermediate cleaning event sequence, completing the missing context parameter default values, and generating the post-cleaning behavior event sequence.

[0057] For example, check the event context parameters of each transaction record. If a consumption transaction record lacks an important context parameter such as the transaction amount, then the missing default value needs to be filled in. This can be done by querying the average amount of the same type of transactions by the user (such as consumption at the same merchant or in the same time period). Assuming that the average amount of consumption at the same merchant in the past 10 times is 300 yuan, then the amount of this transaction is filled in as 300 yuan. After the above integrity check, all event records have complete information, thus generating a cleaned behavioral event sequence.

[0058] In a possible implementation, step S113 includes: Step S1131 : dynamically adjusting the length of the time window granularity according to the distribution density of event triggering timestamps in the post-cleaning behavior event sequence.

[0059] Step S1132: if the number of event triggering times exceeds a density threshold within a preset first time period, the time window granularity is reduced to a first sub-granularity.

[0060] Step S1133: if the number of event triggering times in the first time period is lower than a sparse threshold, the time window granularity is expanded to a second sub-granularity.

[0061] Taking credit card transactions as an example, the transaction frequency varies greatly in different time periods. Assume that a first time period is preset, such as November 1, 2022 to December 31, 2022 (this is a credit card promotion season), and the density threshold is set to 60 transactions per day. If during this time period, it is found that the number of times a user's credit card transaction event is triggered exceeds the density threshold, then the time window granularity needs to be reduced to the first sub-granularity. The original time window may be divided by week, but now it is divided by day. On the contrary, if the sparse threshold is set to 3 transactions per day from March 1, 2023 to March 31, 2023 (off-peak hours on weekdays), when it is found that the number of times the user's credit card transaction event is triggered during this time period is lower than the sparse threshold, the time window granularity is expanded to the second sub-granularity. For example, the time window originally divided by day is now divided by two weeks.

[0062] Step S1134: based on the adjusted time window granularity, the post-cleaning behavior event sequence is divided into a plurality of continuous or partially overlapping time window intervals.

[0063] For example, after adjustment, the time window granularity is divided by day. Starting from November 1, 2022, the time window intervals are divided into November 1, 2022, November 2, 2022, etc.

[0064] Step S1135: Perform frequency statistics on the event operation type identifiers within each of the time window intervals, and use the operation type with the highest frequency as the main behavior operation type.

[0065] For example, within the time window of November 1, 2022, there are 20 transaction records, including 12 consumption operations, 5 repayment operations, and 3 cash withdrawal operations. Since consumption operations have the highest frequency, consumption operations are determined as the main behavior operation type corresponding to the time window.

[0066] In a possible implementation, step S114 includes: Step S1141: Convert the behavior subsequence corresponding to each of the time windows into a sequence vector of event operation type identifiers.

[0067] For example, for an operation sequence within a time window of [consumption, repayment, consumption, withdrawal, repayment], if the consumption operation is encoded as 1, repayment is 2, and withdrawal is 3, then the behavior subsequence is converted into a sequence vector of [1, 2, 1, 3, 2].

[0068] Step S1142: Use a preset clustering model to perform unsupervised clustering on the sequence vector to obtain a plurality of candidate behavior clusters.

[0069] Assuming that there are sequence vectors corresponding to multiple time windows, these vectors are divided into multiple candidate behavior clusters according to their feature similarities through the K-means clustering algorithm. For example, three candidate behavior clusters may be obtained, cluster 1 contains some sequence vectors that are mainly consumption-oriented with high consumption frequency and relatively few repayments; cluster 2 contains sequence vectors with relatively balanced consumption and repayment; cluster 3 contains sequence vectors with relatively more cash withdrawal operations.

[0070] Step S1143, calculating the similarity between the center vector of each candidate behavior cluster and a preset reference behavior pattern vector.

[0071] Assume that the benchmark behavior pattern vector is set according to the proportion of operation types in the normal use mode of the credit card, for example, consumption accounts for 60%, repayment accounts for 30%, and cash withdrawal accounts for 10%. For cluster 1, calculate the similarity between its center vector and the benchmark behavior pattern vector. The calculation process is as follows: First, determine the center vector of cluster 1. Assume that the average proportion of consumption operations in cluster 1 is 70%, repayment accounts for 20%, and cash withdrawal accounts for 10%. Calculate the similarity between the two by calculating the sum of the squares of the difference in the proportion of each operation type, taking the square root to measure the distance, and then subtracting the distance from 1 to get the similarity. For consumption operations, the difference is 70% - 60% = 10%; for repayment operations, the difference is 30% - 20% = 10%; for cash withdrawal operations, the difference is 10% - 10% = 0%. The sum of the squares of the differences is (10%)²+(10%)²+0² = 0.02, and the square root is approximately 0.1414, so the similarity is 1 - 0.1414 = 0.8586. The similarity between cluster 2 and cluster 3 and the benchmark behavior pattern vector is calculated in the same way.

[0072] Step S1144: determine the candidate behavior cluster with the highest similarity as the main behavior cluster, and use the operation type corresponding to the main behavior cluster as the main behavior operation type.

[0073] For example, by comparing the calculated similarities, it is found that cluster 1 has the highest similarity, so cluster 1 is determined as the main behavior cluster. Since the operation type corresponding to cluster 1 is mainly consumption, the main behavior operation type corresponding to the time window is consumption.

[0074] Step S1145 , adding the operation types in other candidate behavior clusters whose occurrence frequencies are greater than the auxiliary frequency threshold to the auxiliary behavior operation type set.

[0075] For example, for other candidate behavior clusters (Cluster 2 and Cluster 3), check the frequency of operation types. Assume that the auxiliary frequency threshold is set to 3 times. In Cluster 2, although consumption and repayment are relatively balanced, the cash withdrawal operation occurs 4 times, which is greater than the auxiliary frequency threshold. In Cluster 3, the cash withdrawal operation occurs 5 times, which is also greater than the auxiliary frequency threshold. Then add the cash withdrawal operation to the auxiliary behavior operation type set.

[0076] In a possible implementation, step S120 includes: Step S121, defining the state space dimension of the state transition probability matrix according to the unique identification set of all behavior operation types in the user behavior trajectory data.

[0077] In the credit card usage scenario, assuming that there are three types of behavioral operations: consumption, repayment, and cash withdrawal, which are represented by labels C, R, and W respectively, then the state transition probability matrix is ​​a 3×3 matrix, with rows and columns corresponding to these three types of behavioral operations.

[0078] Step S122: align the discrete time points in the user behavior trajectory data in chronological order to generate a behavior state transition chain of the target user.

[0079] For example, the user's operation sequence is sorted out in chronological order from the credit card interaction log as [C, R, C, W, R]. This constitutes a behavior state transition chain, indicating that the user first performed a consumption operation, then a repayment operation, then another consumption operation, then a cash withdrawal operation, and finally a repayment operation.

[0080] Step S123, counting the number of transitions between every two adjacent behavior operation types in the behavior state transition chain to generate an initial state transition number matrix.

[0081] In the above behavior state transition chain [C, R, C, W, R], the transition from C to R occurs once, the transition from R to C occurs once, the transition from C to W occurs once, and the transition from W to R occurs once. Since there is no transition from R to W and from W to C, the initial state transition count matrix is:

[0082]

[0083] The first row corresponds to the migration from C, the first column to C, and so on.

[0084] Step S124, normalizing the initial state transition frequency matrix to obtain transition probability values ​​between each pair of behavior operation types, and filling the values ​​into corresponding positions of the state transition probability matrix.

[0085] In a possible implementation, step S124 includes: Step S1241, traverse each row of elements in the initial state transition frequency matrix, and calculate the sum of all elements in the row.

[0086] Step S1242, comparing the sum value with a preset minimum row sum threshold, if the sum value is less than the preset minimum row sum threshold, adding a pseudo count compensation value to the row to obtain compensated elements in each row.

[0087] For the first row, the elements are 0, 1, 1, and the sum is 2; for the second row, the elements are 1, 0, 0, and the sum is 1; for the third row, the elements are 0, 1, 0, and the sum is 1. Assume that the preset minimum row sum threshold is 2. For the second and third rows, their sums are less than the preset minimum row sum threshold, so a pseudo count compensation value needs to be added to these two rows. Assuming the compensation value is 1, the second row becomes 2, 0, 0, and the third row becomes 1, 1, 0.

[0088] Step S1243, normalize the probability of each row of elements after compensation so that the sum of all elements in the row is 1, and obtain a normalized probability value.

[0089] Step S1244, rounding the normalized probability values ​​to a preset decimal precision to generate a standardized transition probability matrix.

[0090] For the first row, the original elements are 0, 1, 1, and the sum is 2. After normalization, they are 0.0, 0.5, 0.5; for the second row, the original elements are 2, 0, 0, and the sum is 2. After normalization, they are 1.0, 0.0, 0.0; for the third row, the original elements are 1, 1, 0, and the sum is 2. After normalization, they are 0.5, 0.5, 0.0. The normalized probability values ​​are rounded to the preset decimal precision. Assuming that the preset decimal precision is two decimal places, the obtained standardized transition probability matrix is:

[0091]

[0092] Step S1245: replace the elements in the standardized transition probability matrix that are lower than the preset probability lower limit value with the lower limit value to generate a final state transition probability matrix.

[0093] Assuming that the preset probability lower limit is 0.05, since 0.00 in the matrix is ​​lower than the lower limit, it is replaced with 0.05, and the final state transition probability matrix is: [0.05 0.50 0.45 0.95 0.05 0.00 0.45 0.50 0.05 ] Step S125, replacing the elements with transition probability values ​​of zero in the state transition probability matrix with a preset minimum probability value, and generating a target state transition probability matrix after smoothing.

[0094] Assuming that the preset minimum probability value is 0.01, if there are elements with a probability value of zero in the matrix, they are replaced with 0.01 to obtain the target state transition probability matrix after smoothing, so as to avoid calculation problems caused by zero probability in subsequent calculations, thereby completing the process of constructing a state transition probability matrix based on credit card user behavior trajectory data.

[0095] In a possible implementation, step S130 includes: Step S131, performing power iteration calculation on the state transition probability matrix until the maximum difference of each row element in the state transition probability matrix is ​​less than a preset convergence threshold, thereby obtaining a steady-state distribution vector.

[0096] Assume that the initial state transition probability matrix is ​​constructed based on the three types of credit card user behavior operations (marked as C, R, and W, respectively), as shown below: [0.3 0.5 0.2 0.4 0.3 0.3 0.2 0.4 0.4] Perform power iterations, where each iteration multiplies the matrix by itself. For example, the first iteration calculates the new matrix:

[0097]

[0098]

[0099] After detailed calculation (the detailed calculation process of multiple iterations is omitted here), until the maximum difference of each row element in the state transition probability matrix is ​​less than the preset convergence threshold (assuming the convergence threshold is 0.001), the steady-state distribution vector is obtained. For example, after multiple iterations, the steady-state distribution vector is [0.33, 0.33, 0.34].

[0100] Step S132, filtering out state nodes whose probability values ​​are greater than a preset absorbing state threshold from the steady-state distribution vector, and generating a set of candidate absorbing state behavior operation types.

[0101] Specifically, the state nodes whose probability values ​​are greater than the preset absorption state threshold (assuming the absorption state threshold is 0.3) can be screened out from the steady-state distribution vector to generate a set of candidate absorption state behavior operation types. In this example, the probability values ​​corresponding to consumption (C) and cash withdrawal (W) meet the conditions and enter the candidate absorption state behavior operation type set.

[0102] Step S133: for each behavior operation type in the candidate absorbing state behavior operation type set, verify whether the state self-transition probability of the behavior operation type within a preset number of consecutive time periods is continuously greater than a preset stability threshold.

[0103] Step S134, merging the behavior operation types that are continuously greater than the preset stability threshold into the absorbing state behavior operation type set, and deleting the behavior operation types that have not passed the verification in the candidate absorbing state behavior operation type set.

[0104] In detail, for each behavior operation type in the candidate absorbing behavior operation type set, verify whether the state self-transition probability of the behavior operation type is continuously greater than the preset stability threshold (assuming the stability threshold is 0.6) within a preset number of consecutive time periods (assuming 3 consecutive months). Take consumption (C) as an example, check the value of consumption to consumption (i.e., self-transition probability) in the state transition probability matrix of each month. Assume that the self-transition probability from consumption to consumption is 0.7 in the first month, 0.65 in the second month, and 0.62 in the third month, all greater than 0.6, meeting the condition; for cash withdrawal (W), assume that its self-transition probability is 0.5 in the first month, 0.55 in the second month, and 0.58 in the third month, which does not meet the condition. Merge the behavior operation types that are continuously greater than the preset stability threshold (here only consumption C) into the absorbing behavior operation type set, and delete the behavior operation types that have not passed the verification in the candidate absorbing behavior operation type set (here cash withdrawal W).

[0105] Step S135, increasing the self-transition probability of the state node corresponding to each behavior operation type in the absorbing state behavior operation type set by a preset reinforcement coefficient, reducing the transition probability value from the absorbing state behavior operation type to the non-absorbing state behavior operation type, and evenly distributing the difference to the migration paths of other absorbing state behavior operation types, to obtain an adjusted state transition probability matrix.

[0106] Wherein, step S135 includes: Step S1351, traverse each behavior operation type in the absorbing state behavior operation type set, and identify the row index of the state node corresponding to the behavior operation type in the state transition probability matrix.

[0107] Step S1352: In the state transition probability matrix, for the row vector corresponding to the row index, extract the self-transition probability value of the state node corresponding to the absorbing state behavior operation type in the row vector.

[0108] Step S1353: multiply the self-transition probability value by a preset enhancement coefficient to obtain an enhanced self-transition probability value, and replace the original self-transition probability value in the row vector with the enhanced self-transition probability value.

[0109] Step S1354, in the row vector, identify the column indices corresponding to all non-absorbing state behavior operation types, extract the transition probability values ​​corresponding to the column indices, and reduce the transition probability values ​​corresponding to the column indices to the product of the original transition probability values ​​and the preset attenuation coefficient to obtain the attenuated non-absorbing state transition probability values.

[0110] Step S1355, calculating the difference between the original transition probability value corresponding to the column index and the attenuated non-absorbing state transition probability value, generating a difference component for each non-absorbing state transition path, and summing the difference components of all non-absorbing state transfer paths to obtain a total difference.

[0111] Step S1356, identify the column indexes corresponding to other absorbing state behavior operation types in the absorbing state behavior operation type set except the current behavior operation type, and count the number of the other absorbing state behavior operation types, divide the total difference by the number of the other absorbing state behavior operation types, and obtain the average distributed difference.

[0112] Step S1357: In the row vector, the transition probability value corresponding to each of the other absorbing state behavior operation types is increased by the average distribution difference amount.

[0113] Step S1358, repeat the above steps until the row vectors corresponding to all behavior operation types in the absorbing state behavior operation type set have completed the adjustment of the transfer probability values, recombining all the adjusted row vectors into an intermediate state transfer probability matrix, and performing column normalization on each column of the intermediate state transfer probability matrix so that the sum of all elements in each column is equal to 1, thereby generating an adjusted state transfer probability matrix.

[0114] Step S1359, manually calibrate the elements in the adjusted state transition probability matrix whose floating-point errors due to column normalization processing exceed the preset error threshold, ensure that the error value of each column is less than the preset error threshold, and output the calibrated state transition probability matrix as the final adjusted state transition probability matrix to the next processing node.

[0115] In this embodiment, the self-transition probability of the state node corresponding to each behavior operation type in the absorbing state behavior operation type set (here is consumption C) is increased by a preset reinforcement coefficient (assuming the reinforcement coefficient is 1.2), the transition probability value of the absorbing state behavior operation type to the non-absorbing state behavior operation type is reduced, and the difference is evenly distributed to the migration path of other absorbing state behavior operation types (here there is only consumption C, no other absorbing state behavior operation types, so this step is just a self-reinforcement in this example, but it is explained according to the complete process). In the state transition probability matrix, first traverse each behavior operation type in the absorbing state behavior operation type set (here there is only consumption C), and identify the row index of the state node corresponding to consumption C in the state transition probability matrix (assuming it is the first row). In the row vector, extract the self-transition probability value of the state node corresponding to consumption C (assuming the initial self-transition probability value is 0.5). Multiply the self-transition probability value by the preset reinforcement coefficient 1.2 to obtain the enhanced self-transition probability value of 0.6 (0.5×1.2 = 0.6), and replace the enhanced self-transition probability value with the original self-transition probability value in the row vector.

[0116] In this row vector, identify the column indices (the second and third columns, respectively) corresponding to all non-absorbing behavior operation types (here, repayment R and cash withdrawal W), extract the transition probability values ​​corresponding to the column indices (0.3 and 0.2, respectively), and reduce the transition probability values ​​corresponding to the column indices to the product of the original transition probability value and the preset attenuation coefficient (assuming the attenuation coefficient is 0.8) to obtain the attenuated non-absorbing state transition probability value. For repayment R, the attenuated transition probability value is 0.24 (0.3×0.8= 0.24); for cash withdrawal W, the attenuated transition probability value is 0.16 (0.2×0.8 = 0.16).

[0117] Calculate the difference between the original transition probability value corresponding to the column index and the attenuated non-absorbing state transition probability value, and generate the difference component of each non-absorbing state transition path. For repayment R, the difference component is 0.06 (0.3 - 0.24 = 0.06); for cash withdrawal W, the difference component is 0.04 (0.2 - 0.16 = 0.04). Sum the difference components of all non-absorbing state transfer paths to get a total difference of 0.1 (0.06 + 0.04 = 0.1). Since there is only one absorbing state behavior operation type, consumption C, and there are no other absorbing state behavior operation types to allocate the difference, this step is calculated in this example but not actually allocated.

[0118] Repeat the above steps until the row vectors corresponding to all behavior operation types in the absorbing state behavior operation type set have completed the adjustment of the transition probability values, and then reassemble all the adjusted row vectors into the intermediate state transition probability matrix. Since there is only one row adjusted here, the intermediate state transition probability matrix is ​​the matrix composed of this adjusted row (assuming it is [0.6, 0.24, 0.16]). Then, each column of the intermediate state transition probability matrix is ​​normalized so that the sum of all elements in each column is equal to 1. Calculate the sum of the elements in each column. The first column has only one element 0.6, and no adjustment is required; the second column element is 0.24, and the sum is 0.24. To make it sum to 1, divide 0.24 by 0.24 to get 1; the third column element is 0.16, and the sum is 0.16. Divide 0.16 by 0.16 to get 1. The renormalized state transition probability matrix is ​​[0.6, 1, 1] (here, in order to simplify the explanation, the actual situation will be calculated based on the complete matrix situation).

[0119] In this process, if the floating point error caused by column normalization in the renormalized state transition probability matrix exceeds the preset error threshold (assuming the error threshold is 0.0001), the elements are manually calibrated. For example, if an element should be 0.99995 after calculation, but due to the floating point error, it is displayed as 0.9998, and the error with the theoretical value exceeds 0.0001, it needs to be manually calibrated to 0.9999 to ensure that the error value of each column is less than the preset error threshold. The calibrated state transition probability matrix is ​​output to the next processing node as the final adjusted state transition probability matrix.

[0120] Step S136, re-perform column normalization processing on the adjusted state transition probability matrix so that the sum of elements in each column is 1, thereby obtaining a re-normalized state transition probability matrix.

[0121] Step S137, input the renormalized state transition probability matrix as the updated state transition probability matrix into the portrait label generation process of the next time period.

[0122] Step S138, recording the update history of the absorbing state behavior operation type set for dynamic correction of subsequent portrait label sets.

[0123] For example, record the time when consumption C is determined to be an absorbing behavior operation type, related probability value changes, and other information, so as to serve as a reference for dynamically revising the portrait label set based on new behavior data in the future.

[0124] In a possible implementation, step S140 includes: Step S141 : extracting a context attribute feature vector corresponding to each behavior operation type from the absorbing state behavior operation type set, wherein the context attribute feature vector includes an operation frequency, an operation duration, and an associated resource identifier.

[0125] In this embodiment, it is assumed that there is only one behavior operation type, consumption, in the set of absorbing behavior operation types. For consumption operations, in terms of operation frequency, by counting the credit card interaction logs, it is found that the user has performed consumption operations an average of 15 times per month in the past year; the operation duration here can be understood as the average time from initiation to completion of each consumption transaction, which is calculated to be 2.5 minutes; the associated resource identifier is the specific bank account number 123456 to which the credit card belongs, thus forming a context attribute feature vector corresponding to the behavior operation type of consumption, that is, the operation frequency is 15 times per month, the operation duration is 2.5 minutes, and the associated resource identifier is the bank account number 123456.

[0126] Step S142, matching the context attribute feature vector with the label definition rules in a preset portrait label knowledge base to determine a candidate portrait label list corresponding to each behavior operation type.

[0127] In detail, in the portrait label knowledge base, there are the following definition rules for consumption operation frequency: more than 10 times of consumption per month is high-frequency consumption, 5-10 times is medium-frequency consumption, and less than 5 times is low-frequency consumption. Since the user consumes an average of 15 times per month, it meets the label definition of high-frequency consumption. For operation duration, if the duration of each consumption is less than 3 minutes, it is short-duration consumption. The average consumption duration of this user is 2.5 minutes, which meets the label definition of short-duration consumption. For associated resource identifiers, if the bank account number 123456 belongs to a specific high-risk account association type (this is based on the bank's internal account risk classification setting, for example, the account has had some suspected risky transaction history), it meets the definition of high-risk account associated consumption label. Therefore, the candidate portrait label list corresponding to the consumption behavior operation type is high-frequency consumption, short-duration consumption, and high-risk account associated consumption.

[0128] Step S143, semantically aggregate the tags in the candidate portrait tag list, merge multiple sub-tags describing the same behavior dimension, and generate an aggregated tag set.

[0129] In detail, in this example, both high-frequency consumption and short-duration consumption describe the time and frequency characteristics of consumption to a certain extent and belong to the same behavioral dimension. They are combined into the aggregated label of high-frequency short-duration consumption. As a result, the aggregated label set includes the two labels of high-frequency short-duration consumption and high-risk account-associated consumption.

[0130] Step S144, calculating the confidence weight of each portrait tag based on the occurrence frequency and time distribution density of each portrait tag in the aggregated tag set.

[0131] For example, for high-frequency, short-duration consumption tags, check their frequency of occurrence in past data analysis. Assume that this tag appeared 7 times in the past 10 similar analyses. In terms of its time distribution density, these 7 occurrences are relatively evenly distributed in the past 10 analyses. For high-risk account-associated consumption tags, assume that they appeared 3 times in the past 10 analyses, and the distribution is relatively concentrated in the most recent 3 analyses. When calculating the confidence weight, a calculation method that comprehensively considers the frequency of occurrence and the time distribution density is used. For high-frequency, short-duration consumption tags, first calculate the frequency weight. It appears 7 times, the total number of analyses is 10, and the frequency weight is 7÷10 = 0.7. Because its distribution is relatively uniform, the time distribution density weight is set to 0.8 (here, uniform distribution is given a higher weight, and the specific weight setting can be adjusted according to actual business needs). Multiplying the frequency weight and the time distribution weight gives a confidence weight of 0.7×0.8 = 0.56. For consumption labels associated with high-risk accounts, the frequency weight is 3÷10 = 0.3. Since their distribution is concentrated in the recent period, the time distribution density weight is set to 0.6 (recent concentrated appearances are given a set weight, but it is lower than the weight of uniform distribution), and the confidence weight is 0.3×0.6 = 0.18.

[0132] Step S145, sorting and screening the aggregated tag set according to the confidence weight, retaining tags with weight values ​​greater than a preset threshold, and generating the portrait tag set.

[0133] Assume that the preset threshold is 0.3. Since the confidence weight of the high-frequency short-duration consumption label is 0.56, which is greater than 0.3, and the confidence weight of the high-risk account-associated consumption label is 0.18, which is less than 0.3, only the high-frequency short-duration consumption label is retained in the portrait label set. This portrait label set can reflect that the credit card user has a stable feature of high frequency and short duration in consumption behavior, and this feature is of great significance in the construction of portraits related to the default risk of financial services.

[0134] Figure 2 A schematic diagram of exemplary hardware and software components of a user portrait tag generation system 100 based on a time-aligned transfer model that can implement the concept of the present application is shown in some embodiments of the present application. For example, the processor 120 can be used in the user portrait tag generation system 100 based on a time-aligned transfer model and used to perform the functions in the present application.

[0135] The user portrait tag generation system 100 based on the time-aligned transfer model can be a general server or a special-purpose server, both of which can be used to implement the user portrait tag generation method based on the time-aligned transfer model of the present application. Although only one server is shown in the present application, for convenience, the functions described in the present application can be implemented in a distributed manner on multiple similar platforms to balance the processing load.

[0136] For example, the user portrait label generation system 100 based on the time-aligned transfer model may include a network port 110 connected to the network, one or more processors 120 for executing program instructions, a communication bus 130, and storage media 140 in different forms, such as disks, ROMs, or RAMs, or any combination thereof. Exemplarily, the user portrait label generation system 100 based on the time-aligned transfer model may also include program instructions stored in ROM, RAM, or other types of non-temporary storage media, or any combination thereof. The method of the present application can be implemented according to these program instructions. The user portrait label generation system 100 based on the time-aligned transfer model also includes an input / output (I / O) interface 150 between the computer and other input / output devices.

[0137] For ease of explanation, only one processor is described in the user portrait label generation system 100 based on the time-aligned transfer model. However, it should be noted that the user portrait label generation system 100 based on the time-aligned transfer model in the present application may also include multiple processors, so the steps performed by one processor described in the present application may also be performed jointly or individually by multiple processors. For example, if the processor of the user portrait label generation system 100 based on the time-aligned transfer model executes steps A and B, it should be understood that steps A and B may also be executed jointly by two different processors or individually in one processor. For example, the first processor executes step A, the second processor executes step B, or the first processor and the second processor execute steps A and B together.

[0138] In addition, an embodiment of the present invention further provides a readable storage medium, in which computer executable instructions are preset. When the processor executes the computer executable instructions, the user portrait label generation method based on the time-aligned transfer model as described above is implemented.

[0139] It should be noted that in order to simplify the description of the present invention and thus help understand one or more embodiments of the invention, in the foregoing description of the embodiments of the present invention, various features are sometimes combined into one embodiment, drawing or description thereof.

Claims

1. A method for generating user portrait labels based on a time-aligned transfer model, characterized in that: The method comprises: Acquire user behavior trajectory data of a target user within a preset time period, wherein the user behavior trajectory data includes behavior operation types and corresponding context attributes at multiple discrete time points; Based on the temporal correlation of each behavior operation type in the user behavior trajectory data, a state transition probability matrix corresponding to the target user is constructed, wherein each element in the state transition probability matrix represents a transition probability of the target user migrating from a first behavior operation type to a second behavior operation type; Determine, according to the steady-state convergence of each state node in the state transition probability matrix, a set of absorbing state behavior operation types corresponding to the target user, each behavior operation type in the set of absorbing state behavior operation types satisfies a state transition probability threshold condition within a continuous time period; Based on the behavior operation types in the absorbing state behavior operation type set and the corresponding context attributes, a portrait tag set of the target user is generated, wherein each portrait tag in the portrait tag set is used to describe the stable characteristics of the target user in a specified behavior dimension; After matching and verifying the portrait tag set with a preset portrait tag knowledge base, a dynamically updated portrait tag sequence corresponding to the target user is output.

2. The user portrait label generation method based on the time-aligned transfer model according to claim 1 is characterized in that: The step of obtaining the user behavior trajectory data of the target user within a preset time period includes: Extracting an original behavior event sequence from the target user's interaction log, wherein the original behavior event sequence includes an event triggering timestamp, an event operation type identifier, and an event context parameter; Performing data cleaning on the original behavior event sequence, removing duplicate event records and event records with abnormal timestamps, and obtaining a cleaned behavior event sequence; Segment the cleaned behavior event sequence according to a preset time window granularity to generate behavior subsequences corresponding to multiple time windows, wherein the event triggering timestamp in each behavior subsequence satisfies the time range constraint of the time window; Performing cluster analysis on the event operation type identifiers in each of the behavior subsequences to obtain a set of main behavior operation types and auxiliary behavior operation types corresponding to the time window; Based on the event context parameters and the main behavior operation type and the corresponding auxiliary behavior operation type set, the behavior operation type and context attributes corresponding to each discrete time point in the user behavior trajectory data are generated.

3. The user portrait label generation method based on the time-aligned transfer model according to claim 2 is characterized in that: The data cleaning of the original behavior event sequence, removing duplicate event records and event records with abnormal timestamps, and obtaining a cleaned behavior event sequence includes: Detect whether there is an event record with a timestamp earlier than the system record start time or later than the current system time in the original behavior event sequence, and mark it as an abnormal timestamp event; Identify event records in the original behavior event sequence whose event operation type identifier is empty or does not conform to a preset coding specification, and mark them as invalid operation type events; Traversing every two consecutive event records in the original behavior event sequence, if the time stamp interval between the two consecutive event records is less than a preset minimum operation interval, marking them as repeated triggering events; Remove the abnormal timestamp events, invalid operation type events and repeated triggering events from the original behavior event sequence to generate an intermediate cleaning event sequence; An integrity check is performed on the remaining event records in the intermediate cleaning event sequence, missing default values ​​of context parameters are supplemented, and the post-cleaning behavior event sequence is generated.

4. The method for generating user portrait labels based on a time-aligned transfer model according to claim 2, characterized in that: The step of segmenting the post-cleaning behavior event sequence according to a preset time window granularity to generate behavior subsequences corresponding to multiple time windows includes: Dynamically adjusting the length of the time window granularity according to the distribution density of event triggering timestamps in the post-cleaning behavior event sequence; If the number of event triggering times exceeds a density threshold within a preset first time period, the time window granularity is reduced to a first sub-granularity; If the number of event triggering times in the first time period is lower than a sparse threshold, expanding the time window granularity to a second sub-granularity; Based on the adjusted time window granularity, the post-cleaning behavior event sequence is divided into a plurality of continuous or partially overlapping time window intervals; The frequency of the event operation type identifier within each of the time window intervals is counted, and the operation type with the highest frequency is used as the main behavior operation type.

5. The method for generating user portrait labels based on a time-aligned transfer model according to claim 2, characterized in that: The cluster analysis is performed on the event operation type identifier in each of the behavior subsequences to obtain a set of main behavior operation types and auxiliary behavior operation types corresponding to the time window, including: Converting the behavior subsequence corresponding to each of the time windows into a sequence vector of event operation type identifiers; Using a preset clustering model to perform unsupervised clustering on the sequence vector to obtain multiple candidate behavior clusters; Calculating the similarity between the center vector of each candidate behavior cluster and a preset benchmark behavior pattern vector; Determine the candidate behavior cluster with the highest similarity as the main behavior cluster, and use the operation type corresponding to the main behavior cluster as the main behavior operation type; The operation types in other candidate behavior clusters whose occurrence frequencies are greater than the auxiliary frequency threshold are added to the auxiliary behavior operation type set.

6. The method for generating user portrait labels based on a time-aligned transfer model according to claim 1, characterized in that: The constructing a state transition probability matrix corresponding to the target user based on the temporal correlation of each behavior operation type in the user behavior trajectory data includes: Defining the state space dimension of the state transition probability matrix according to a unique identification set of all behavior operation types in the user behavior trajectory data; Aligning discrete time points in the user behavior trajectory data in chronological order to generate a behavior state transition chain of the target user; Counting the number of migrations between every two adjacent behavior operation types in the behavior state migration chain to generate an initial state transition number matrix; Normalizing the initial state transition frequency matrix to obtain the transition probability value between each pair of behavior operation types, and filling the corresponding position of the state transition probability matrix; The elements with transition probability values ​​of zero in the state transition probability matrix are replaced with a preset minimum probability value to generate a target state transition probability matrix after smoothing.

7. The method for generating user portrait labels based on a time-aligned transfer model according to claim 6, characterized in that: The normalizing process of the initial state transition times matrix to obtain the transition probability value between each pair of behavior operation types and fill the corresponding position of the state transition probability matrix includes: Traverse each row of elements in the initial state transition count matrix and calculate the sum of all elements in the row; Compare the sum value with a preset minimum row sum threshold, and if the sum value is less than the preset minimum row sum threshold, add a pseudo count compensation value to the row to obtain compensated elements in each row; Normalize the probability of each row of elements after compensation so that the sum of all elements in the row is 1, and obtain the normalized probability value; The normalized probability values ​​are rounded to a preset decimal precision to generate a standardized transition probability matrix; The elements in the standardized transition probability matrix that are lower than a preset probability lower limit are replaced with the lower limit value to generate a final state transition probability matrix.

8. The method for generating user portrait labels based on a time-aligned transfer model according to claim 1, characterized in that: The step of determining a set of absorbing state behavior operation types corresponding to the target user according to the steady-state convergence of each state node in the state transition probability matrix includes: Performing power iteration calculation on the state transition probability matrix until the maximum difference of each row element in the state transition probability matrix is ​​less than a preset convergence threshold, thereby obtaining a steady-state distribution vector; Filter out state nodes whose probability values ​​are greater than a preset absorbing state threshold from the steady-state distribution vector, and generate a set of candidate absorbing state behavior operation types; For each behavior operation type in the candidate absorbing state behavior operation type set, verify whether the state self-transition probability of the behavior operation type within a preset number of consecutive time periods is continuously greater than a preset stability threshold; Merge the behavior operation types that are continuously greater than the preset stability threshold into the absorbing state behavior operation type set, and delete the behavior operation types that have not passed the verification in the candidate absorbing state behavior operation type set; The self-transition probability of the state node corresponding to each behavior operation type in the absorbing state behavior operation type set is increased by a preset reinforcement coefficient, the transition probability value from the absorbing state behavior operation type to the non-absorbing state behavior operation type is reduced, and the difference is evenly distributed to the migration paths of other absorbing state behavior operation types, to obtain an adjusted state transition probability matrix; The adjusted state transition probability matrix is ​​re-normalized so that the sum of the elements in each column is 1, thereby obtaining a re-normalized state transition probability matrix; The renormalized state transition probability matrix is ​​used as the updated state transition probability matrix and input into the portrait label generation process of the next time period; Recording the update history of the absorbing state behavior operation type set for dynamic modification of subsequent portrait label sets; The step of increasing the self-transition probability of the state node corresponding to each behavior operation type in the absorbing state behavior operation type set by a preset reinforcement coefficient, reducing the transition probability value from the absorbing state behavior operation type to the non-absorbing state behavior operation type, and evenly distributing the difference to the migration paths of other absorbing state behavior operation types to obtain an adjusted state transition probability matrix includes: Traversing each behavior operation type in the absorbing state behavior operation type set, and identifying the row index of the state node corresponding to the behavior operation type in the state transition probability matrix; In the state transition probability matrix, for the row vector corresponding to the row index, extract the self-transition probability value of the state node corresponding to the absorbing state behavior operation type in the row vector; Multiplying the self-transition probability value by a preset enhancement coefficient to obtain an enhanced self-transition probability value, and replacing the original self-transition probability value in the row vector with the enhanced self-transition probability value; In the row vector, identifying the column indexes corresponding to all non-absorbing state behavior operation types, extracting the transition probability values ​​corresponding to the column indexes, and reducing the transition probability values ​​corresponding to the column indexes to the product of the original transition probability values ​​and a preset attenuation coefficient to obtain the attenuated non-absorbing state transition probability values; Calculating the difference between the original transition probability value corresponding to the column index and the attenuated non-absorbing state transition probability value, generating a difference component for each non-absorbing state transition path, and summing the difference components of all non-absorbing state transfer paths to obtain a total difference; Identify the column indexes corresponding to other absorbing state behavior operation types in the absorbing state behavior operation type set except the current behavior operation type, count the number of the other absorbing state behavior operation types, and divide the total difference by the number of the other absorbing state behavior operation types to obtain an average distribution difference; In the row vector, increasing the transition probability value corresponding to each of the other absorbing state behavior operation types by the average distribution difference amount; Repeat the above steps until the row vectors corresponding to all behavior operation types in the absorbing state behavior operation type set have completed the adjustment of the transfer probability values, recombining all the adjusted row vectors into an intermediate state transfer probability matrix, and performing column normalization processing on each column of the intermediate state transfer probability matrix so that the sum of all elements in each column is equal to 1, thereby generating an adjusted state transfer probability matrix; Manually calibrate the elements in the adjusted state transition probability matrix whose floating-point errors due to column normalization processing exceed the preset error threshold to ensure that the error value of each column is less than the preset error threshold, and output the calibrated state transition probability matrix as the final adjusted state transition probability matrix to the next processing node.

9. The method for generating user portrait labels based on a time-aligned transfer model according to claim 1, characterized in that: The generating the target user's portrait tag set based on the behavior operation types in the absorbing state behavior operation type set and the corresponding context attributes includes: Extracting a context attribute feature vector corresponding to each behavior operation type from the absorbing state behavior operation type set, wherein the context attribute feature vector includes an operation frequency, an operation duration, and an associated resource identifier; Matching the context attribute feature vector with the label definition rules in a preset portrait label knowledge base to determine a candidate portrait label list corresponding to each behavior operation type; Performing semantic aggregation on the tags in the candidate portrait tag list, merging multiple sub-tags describing the same behavior dimension, and generating an aggregated tag set; Calculating the confidence weight of each portrait tag based on the occurrence frequency and time distribution density of each portrait tag in the aggregated tag set; The aggregated tag set is sorted and screened according to the confidence weight, and tags with weight values ​​greater than a preset threshold are retained to generate the portrait tag set.

10. A user portrait label generation system based on a time-aligned transfer model, characterized in that: The user portrait label generation system based on the time-aligned transfer model includes a processor and a memory, the memory is connected to the processor, the memory is used to store programs, instructions or codes, and the processor is used to execute the programs, instructions or codes in the memory to implement the user portrait label generation method based on the time-aligned transfer model described in any one of claims 1 to 9.

Citation Information

Patent Citations

  • User portrait label value analysis method and device

    CN111522828A

  • Solution matching method and system based on user portrait

    CN113076405A

  • Precise pushing method and system based on user portrait

    CN118886980A

  • User profile method and apparatus for credit card client, device, and medium

    WO2022105177A1

Cited By

  • Label list automatic generation method and system for user portraits

    CN121051282A

  • User portrait label value analysis method and system

    CN121093025A

  • Intelligent assessment method and system for special nursing demand of neurosurgical patient

    CN122266776A