AI-based internet of vehicles user layering operation method and system
By jointly analyzing RFM metrics and multi-channel behavioral data, a dynamic correction mechanism for user value status is constructed, which solves the problem of static rule dependence in the user segmentation system of the vehicle network platform, realizes dynamic correction of user value and churn warning, and improves the accuracy and efficiency of operational resource allocation.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- GUANGDONG LEGEND COMM CO LTD
- Filing Date
- 2026-06-22
- Publication Date
- 2026-07-21
Smart Images

Figure CN122434582A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of user operation management technology, and in particular to an AI-based method and system for tiered user operation in the Internet of Vehicles. Background Technology
[0002] Vehicle-to-everything (V2X) platforms have accumulated a wealth of user consumption behavior data, but existing user segmentation systems largely rely on static rules or single spending amount indicators to classify user levels, making it difficult to reflect the dynamic changes in user value over time. The reach to the same user varies significantly across different channels, and the current operational system lacks a systematic verification mechanism for cross-channel behavior consistency. This results in some users' value status being judged on outdated labels for extended periods, and the discrepancy between operational decisions and actual user behavior continues to widen over time.
[0003] On the other hand, users typically exhibit identifiable behavioral precursors before churn, such as prolonged periods of silence and proactive blocking of push notifications. However, existing churn warning methods focus more on the absolute decline in transaction frequency, limiting their ability to monitor these behavioral signals. The intertwined nature of high-value users' incentive dependence and the evolution of channel blocking behavior makes it difficult for operational outreach strategies to accurately match the actual intervention needs of users with different risk levels, thus constraining the efficiency of operational resource allocation. Summary of the Invention
[0004] This invention discloses an AI-based method and system for segmented operation of connected vehicle users. It aims to achieve dynamic correction of user value status through joint analysis of RFM indicators and multi-channel behavioral data. It constructs an accurate user value map by combining incentive dependency risk identification and clustering correction mechanisms. Furthermore, it forms a graded churn early warning and reach failure gradient matching mechanism based on tiered exit quantification, silent duration distribution, and channel blocking feature monitoring. This provides systematic methodological support for refined segmented operation and churn intervention of connected vehicle platform users.
[0005] The first aspect of this invention proposes an AI-based method for tiered user operation in the Internet of Vehicles (IoV), comprising the following steps: Collect user behavior data and consumption record data, and perform RFM index coupling analysis on the user behavior data and consumption record data to generate a user baseline feature set; Hierarchical offset positioning is performed on the user baseline feature set to extract hierarchical offset features. Cross-channel label cross-verification is performed on the hierarchical offset features to generate a contradiction verification set. Based on the contradiction verification set, value dimension decomposition is performed to extract hierarchical features. Based on the hierarchical features, a hierarchical affiliation label is determined. Incentive dependency risk features are extracted from the user behavior data to generate an incentive dependency risk identifier. Based on the incentive dependency risk identifier and the hierarchical affiliation label, incentive non-response clustering is performed to generate a user value map. Based on the user value graph and the user baseline feature set, hierarchical exit quantification is performed to generate exit amplitude value. The distribution of silence duration of high-value users is extracted from the user value graph to generate silence warning interval. The exit amplitude value is used to perform churn trigger node positioning to generate churn prediction coefficient. The churn prediction coefficient and the hierarchical affiliation label are used to perform joint level matching to generate an abnormal churn pattern. The silent warning interval is monitored for push channel blocking characteristics to generate an operational decline identifier. Based on the operational decline identifier and the abnormal churn pattern, a reach failure gradient matching is performed to generate a user operation monitoring report.
[0006] A second aspect of this invention proposes an AI-based vehicle-to-everything (V2X) user tiered operation system, comprising: The data acquisition module is used to collect user behavior data and consumption record data, and to perform RFM index coupling analysis on the user behavior data and consumption record data to generate a user baseline feature set; The feature verification module is used to perform hierarchical offset positioning to extract hierarchical offset features from the user baseline feature set, perform cross-channel label cross-verification on the hierarchical offset features to generate a contradiction verification set, and decompose the value dimension to extract hierarchical features based on the contradiction verification set. The hierarchical clustering module is used to determine the hierarchical affiliation label based on the hierarchical features, extract incentive dependency risk features from the user behavior data to generate incentive dependency risk identifiers, and perform incentive non-response clustering hierarchically based on the incentive dependency risk identifiers and the hierarchical affiliation labels to generate a user value map. The churn prediction module is used to generate a churn magnitude value by performing hierarchical exit quantification based on the user value graph and the user baseline feature set, extract the silence duration distribution of high-value users from the user value graph to generate a silence warning interval, and perform churn trigger node positioning on the exit magnitude value to generate a churn prediction coefficient. The report output module is used to perform joint level matching between the churn prediction coefficient and the hierarchical affiliation label to generate an abnormal churn pattern, perform push channel blocking feature monitoring on the silent warning interval to generate an operational decline identifier, and perform reach failure gradient matching between the operational decline identifier and the abnormal churn pattern to generate a user operation monitoring report.
[0007] The beneficial effects of this invention are reflected in the following points: 1. This invention jointly constructs a user baseline feature set by combining historical RFM transaction characteristics with multi-channel behavioral activity correction items. By comparing channel response directions, it identifies contradictory signals between tags and actual behavior. Then, through user hierarchical flow trajectory mapping and accelerated return area positioning, it completes the dynamic correction of value status. This solves the problem that user stratification on the Internet of Vehicles platform has long relied on static rules and that cross-channel behavioral contradictions are masked by averaging. This allows the stratification results to continuously track changes in users' real consumption behavior rather than being fixed on historical tags. 2. The difference between the proportion of transactions during incentive periods and those during non-incentive periods is quantified as the degree of incentive dependence risk and introduced into a clustering correction process. Users with high incentive dependence risk and repurchase deviations are subject to targeted hierarchical shifts. This avoids the systemic shift of cluster centers caused by short-term surges in transaction behavior during promotions. This makes the level boundaries of high-value user clusters more accurately reflect users' natural consumption capacity rather than occasional transaction performance driven by incentives. 3. Before a significant decline in transaction frequency, churn behavior often leaves identifiable warning signs through prolonged periods of silence and proactive blocking of push notifications. This invention establishes a pre-judgment mechanism by detecting a sudden increase in net outflow at different levels and locating critical nodes in the silence duration of high-value users. Simultaneously, it monitors the blocking behavior of silent user groups by implementing slope fitting, jointly mapping the degree of decline in channel reachability and churn risk level to a reach failure gradient. This drives differentiated configuration of intervention strategies in operational monitoring reports, allowing intervention to be initiated before obvious churn signals appear, reducing the continuous consumption of reach resources on already ineffective channels. Attached Figure Description
[0008] Figure 1 This is a flowchart illustrating the AI-based user tiered operation method for the Internet of Vehicles (IoV) of this invention.
[0009] Figure 2 This is a structural block diagram of the AI-based vehicle network user tiered operation system of the present invention. Detailed Implementation
[0010] In the following description, specific details such as particular system architectures and techniques are set forth for illustrative purposes and not for limitation, in order to provide a thorough understanding of the embodiments of this application. However, those skilled in the art will understand that this application may also be implemented in other embodiments without these specific details. In other instances, detailed descriptions of well-known systems, apparatuses, circuits, and methods have been omitted so as not to obscure the description of this application with unnecessary detail.
[0011] References to "one embodiment" or "some embodiments" as described in this specification mean that one or more embodiments of this application include a specific feature, structure, or characteristic described in connection with that embodiment. Therefore, the phrases "in one embodiment," "in some embodiments," "in other embodiments," "in still other embodiments," etc., appearing in different parts of this specification do not necessarily refer to the same embodiment, but rather mean "one or more, but not all, embodiments," unless otherwise specifically emphasized. The terms "comprising," "including," "having," and variations thereof mean "including but not limited to," unless otherwise specifically emphasized.
[0012] The technical solutions of the embodiments of this application will be described below.
[0013] like Figure 1 As shown, this embodiment of the invention provides an AI-based method for tiered operation of connected vehicle users, including the following steps S11-S15: Step S11: Collect user behavior data and consumption record data, and perform RFM index coupling analysis on the user behavior data and consumption record data to generate a user baseline feature set.
[0014] Specifically, user behavior data and consumption record data are collected. User interaction logs on the connected vehicle platform are the direct source of user behavior data. Four types of events—product browsing trajectory, favorites operation, push notification click response, and session start and end records—are extracted and written to the data layer in real time. Each event record is accompanied by an occurrence timestamp and source channel label, with timestamp accuracy at the second level. Consumption record data is aggregated from the transaction settlement system by user dimension, including three fields for each order: transaction amount, order placement time, and payment status. Cancelled or refunded orders are accompanied by status labels to distinguish them from valid transactions. There is a structural difference in the collection density of user behavior data and consumption record data. User behavior data is continuously generated during each user session, while consumption record data is only generated upon transaction completion. The temporal distribution of these two types of data varies under the same user dimension. It is common for users on the connected vehicle platform to generate dozens of user behavior data records but only one consumption record within the same week. This density difference necessitates dedicated handling of the node alignment relationship between the two types of data in the subsequent coupling and parsing phase. Both user behavior data and consumption record data use the user's unique identifier as the association key. Records with missing or unmatched identifiers are not included in the RFM indicator coupling and parsing process. The time span of both types of data uniformly covers the most recent 180 days to ensure that the RFM indicator extraction has sufficient behavioral historical depth.
[0015] User behavior data and consumption record data are coupled and analyzed using RFM metrics to generate a user baseline feature set. The mapping logic of the three original metrics is based on the aggregated results of consumption record data—the most recent order time corresponds to the R value, the effective transaction record count corresponds to the F value, and the cumulative sum of transaction amount corresponds to the M value. After normalization for all users, the three metrics constitute the RFM vector for each user. The RFM vector extracted solely from consumption record data only reflects historical transaction results and cannot capture the user's current behavioral activity and changes in purchase intention. The introduction of user behavior data fills this gap. The normalized value of the user's most recent browsing time is superimposed on the R value to form the R correction value, and the normalized values of recent browsing frequency and collection behavior density are superimposed on the F value to form the F correction value. The normalized value of push click response rate is added as a fourth component to the vector. After correction, each component is compressed to the 0 to 1 range after normalization for all users. Users with high click-through rates but low transaction frequency exhibit typical characteristics of high reach but low conversion. Users who have clicked on promotional push notifications multiple times in the past 30 days but have no transaction records in their consumption data have low F and M values in their RFM correction vectors, while their behavioral activity correction component is relatively high. This combination of characteristics forms a unique vector distribution pattern in the user baseline feature set. The RFM correction vectors of all users are arranged by user identifier to form the user baseline feature set. Each user entry carries four normalized values: R correction value, F correction value, M normalized value, and push click-through rate. The vector structure of these four values allows for the quantitative comparison of different types of user value states within the same coordinate system.
[0016] Step S12: Perform hierarchical offset positioning on the user baseline feature set to extract hierarchical offset features, perform cross-channel label cross-verification on the hierarchical offset features to generate a contradiction verification set, and extract hierarchical features by decomposing the value dimension based on the contradiction verification set.
[0017] Specifically, hierarchical offset localization is performed on the user baseline feature set to extract hierarchical offset features. The division of adjacent time periods within the statistical period is the basic operational unit of hierarchical offset localization. By comparing the relative hierarchical positions of the RFM vectors of the same user in two consecutive time periods within the user baseline feature set, user groups whose hierarchical positions have changed are identified. The RFM vectors of each user in the user baseline feature set are first divided into four levels—high, medium-high, medium-low, and low—by the total user quantiles. When the hierarchical assignment changes between adjacent statistical periods, it is recorded as a hierarchical offset event. The direction of hierarchical offset is divided into two categories: upward shift and downward shift. An upward shift indicates an increase in user value, while a downward shift indicates a decrease in user value. The hierarchical offset feature records the direction, magnitude, and timing of each offset event. The offset magnitude is the normalized Euclidean distance between the RFM vectors in the user's baseline feature set for two consecutive statistical periods. A larger magnitude indicates a more drastic hierarchical change. Users whose consumption frequency drops sharply within a single statistical period and who have no recent browsing activity, causing their RFM vectors to rapidly slide from high to low levels, have a downward offset direction in their hierarchical offset feature, and their offset magnitude is at the high end of the overall user distribution, forming a clear contrast with users whose RFM vectors fluctuate slightly during the same period. Users whose RFM vectors in the user's baseline feature set do not change hierarchically during adjacent statistical weeks have an offset magnitude of zero and are not triggered to record hierarchical offset features. All hierarchical offset feature entries for users who have experienced hierarchical offsets are arranged by user identifier. The density distribution of offset events on the timeline is often related to the rhythm of operational activities. High-density offset periods usually correspond to a concentrated outbreak of value sinking after the incentive period ends. In these periods, the proportion of downward shifts in offset entries is significantly higher, and the offset magnitude distribution is stretched towards the tail. There is a strong positive correlation between the severity of hierarchical changes for users with larger offset magnitudes and the depth of contradictory signals in channel responses.
[0018] In some embodiments, performing cross-channel label cross-verification on the hierarchical offset features to generate a contradiction verification set includes: classifying the hierarchical offset features by channel affiliation to generate each channel affiliation subset; performing cross-channel consumption response direction comparison on each channel affiliation subset to generate a channel response difference table; locating the RFM and label reverse response intervals based on the channel response difference table to generate a contradiction signal candidate set; and performing contradiction degree critical calibration on the contradiction signal candidate set to generate a contradiction verification set.
[0019] Channel affiliation subsets are generated by classifying the hierarchical offset features. App push notifications, SMS messages, activity pages, and social sharing constitute the four channel affiliation categories. The core basis for affiliation confirmation is the channel type through which each user received and responded to the push notification during the time period of the hierarchical offset event. The same user may generate response records on multiple channels within the same statistical period. The user's hierarchical offset feature entries are copied and assigned to the corresponding subsets according to channel type. Each subset retains the user's consumption response direction label for the corresponding channel. The consumption response direction label is recorded by the positive or negative sign of the change in the user's RFM vector within the statistical window after the channel was reached. A positive sign indicates improved consumption behavior after the reach, while a negative sign indicates further decline in consumption behavior. User offset entries with no channel response records in the hierarchical offset features are grouped into the unreached subset for separate processing. The unreached subset reflects that the hierarchical changes of these users occurred under completely natural behavior rather than operational intervention. It is not included in the comparison of channel response directions during cross-channel comparisons but is retained as a benchmark. The differences in the size of the subsets belonging to each channel type reflect the current operational system's allocation of resources to different channels. The APP push subset is typically the largest, while the social sharing subset is usually the sparsest. When, within a certain statistical period, the number of users moving down from higher levels in the APP push subset is significantly concentrated, while the number of similar users in the SMS reach subset is relatively small, the size difference between the two subsets suggests that the APP push channel has a higher correlation with the downward movement of high-value users within that period. The pairing relationship between user identifiers and consumption response direction labels within each channel subset records the channel background of user level changes under different reach scenarios. The distribution of consumption response direction labels can reveal which type of channel's reach behavior has a stronger temporal correlation with level shift events.
[0020] A channel response difference table is generated by comparing the cross-channel consumption response directions for each channel's subset. Consumption response direction is the core variable for comparison. The consumption response direction annotations of each user in each channel subset are directly extracted for cross-channel comparison. Users with high direction consistency exhibit stable cross-channel behavior, while users with inconsistent directions show contradictory response characteristics between channels. When the same user has response records in two or more channels within each channel subset, the consumption response directions corresponding to each channel are compared pairwise. If the product of the direction signs of each channel is positive, it is considered to be in the same direction; if the product is negative, it is considered to be in the opposite direction. The reverse combination record is a cross-channel response difference entry for that user, containing three fields: user identifier, involved channel pair, and reverse response period. The cross-channel response discrepancy entries for all users are summarized and arranged by user identifier to form a channel response discrepancy table. The more reverse combinations of cross-channel responses for the same user in the table, the lower the consistency of the user's behavior across different channels. For users whose consumption frequency increases after being reached through APP push notifications but continues to decrease after being reached through SMS notifications, a reverse response entry for APP and SMS channels is formed in the channel response discrepancy table, indicating a significant difference in channel preference for this user. The effectiveness of a general outreach strategy in guiding their behavior is significantly differentiated. The entry density in the channel response discrepancy table reflects the prevalence of cross-channel response inconsistency among the current user group. Users with dense entries will experience a mutually offsetting activation and inhibition effect when maintaining a consistent multi-channel push strategy across different channels.
[0021] Based on the channel response difference table, the RFM and tag-based reverse response intervals are located to generate a candidate set of contradictory signals. The RFM and tag-based reverse response interval refers to the period in the channel response difference table where the same user exhibits cross-channel reverse responses, and the user status described by the operational tag during the same period contradicts the actual direction of change in the RFM vector. A contradiction exists when the tag describes increased activity while the actual RFM vector shifts downward; another contradiction exists when the tag describes a risk of decline while the actual RFM vector shifts upward. For each reverse response entry in the channel response difference table, the user's operational tag and the direction of change in the corresponding RFM vector during the reverse response period are extracted. If the two directions contradict each other, the entry is identified as a contradictory signal entry; if they are consistent, the entry is determined to be a channel preference difference rather than a tag contradiction and is not included in the candidates. The contradiction type of contradictory signal entries is labeled according to the deviation dimension between the label and the RFM direction. The deviation dimension can be any of the R, F, or M values. The more deviation dimensions and the larger the deviation amplitude in the corresponding level offset feature, the deeper the contradiction. For example, users whose operational labels are marked as highly active repeat purchase users, but whose channel response difference table shows a recent decrease in both purchase frequency and amount after being reached through major channels, whose R, F, and M values all show negative changes, and whose level offset feature has a high deviation amplitude among all users, are labeled as having deep contradictions in the contradictory signal candidate set. The entire set of contradictory signal entries constitutes the contradictory signal candidate set. Each entry in the contradictory signal candidate set carries a contradiction type label and the number of deviation dimensions. The more deviation dimensions, the more value dimensions the user's operational label deviates from their actual behavior.
[0022] A contradiction verification set is generated by critically calibrating the contradiction severity of the candidate contradiction signal set. The contradiction intensity score M_score = 0.5 × D_norm + 0.5 × C_norm integrates the number of deviation dimensions (D_norm) and the number of channel reverse response combinations (C_norm) into a single evaluation value. The equal weight of the two dimensions reflects the equal importance of the breadth of label deviation and the depth of channel contradiction in judging the degree of contradiction. The M_score ranges from 0 to 1, with a higher value indicating a deeper overall contradiction between the user label and RFM behavior. Items in the candidate contradiction signal set whose M_score exceeds the mean plus one standard deviation of all candidate items are judged as high contradiction items, those between the mean and the mean plus one standard deviation are judged as medium contradiction items, and those below the mean have insufficient contradiction severity to affect the hierarchical attribution judgment and are not included in the contradiction verification set. For high-contradiction and medium-contradiction entries, four pieces of information are extracted: user identifier, M_score value, contradiction type label, and deviation dimension list, forming a contradiction verification set. All entries are arranged in descending order of M_score to form the contradiction verification set. Users with high M_scores tend to continuously generate cross-channel contradiction signals over multiple operational cycles. The discrepancy between their label status and actual behavior has systematic characteristics rather than being a single, occasional occurrence. Accurately identifying their true value level is of great practical significance for the targeted allocation of operational resources. Users whose deviation dimensions cover both R and F values and have a large number of cross-channel reverse response combinations have high D_norm and C_norm, and their M_score ranks high among all candidate entries, indicating a deep discrepancy between operational labels and actual behavior. Users with only a large number of deviation dimensions but few reverse response combinations have a low C_norm, which lowers the overall M_score value, and are identified as medium-contradiction entries. The difference in M_score between the two types of users in the contradiction verification set truly reflects the difference in the completeness of the contradiction evidence.
[0023] In some embodiments, the step of extracting hierarchical features by decomposing the value dimension based on the contradiction verification set includes: mapping the user-level flow trajectory of the contradiction verification set to generate a hierarchical backflow distribution map; locating high-value layer backflow concentration segments through the hierarchical backflow distribution map to generate a backflow direction feature set; performing backflow rate gradient processing on the backflow direction feature set to generate a backflow acceleration distribution; and locating accelerated backflow regions based on the backflow acceleration distribution to generate hierarchical features.
[0024] A hierarchical return flow distribution map is generated by mapping user hierarchical flow trajectories on the contradiction verification set. The list of deviation dimensions and contradiction type annotations carried by each user entry in the contradiction verification set reveal the direction of deviation between the user's hierarchical affiliation and actual behavior. The hierarchical flow trajectory mapping transforms this deviation direction into a dynamic displacement record of users within the hierarchical system. Using the statistical period as the time unit, the hierarchical affiliation change sequence of each user over multiple consecutive periods is extracted. The path of a user moving down a level and then rising back to a high-value level in the change sequence is defined as the return flow trajectory. The return flow trajectories of each user in the contradiction verification set are statistically analyzed according to the occurrence time and the starting hierarchical level. The number of users returning from the same starting hierarchical level to a high-value level within the same time period forms the return flow density value for that starting hierarchical level during that time period. The return flow density values of all starting hierarchical levels across all time periods are arranged to form the hierarchical return flow distribution map. The vertical axis of the hierarchical return flow distribution map corresponds to the initial hierarchical level, and the horizontal axis corresponds to the statistical period number. Locations with high density values form hotspots in the map. The temporal concentration of these hotspots reflects the degree of aggregation of return flow behavior within a specific period. When the density of users from low-to-medium value levels returning to high-value levels is significantly higher than in other periods over three consecutive statistical periods, a prominent hotspot forms on the hierarchical return flow distribution map for that period, indicating a concentrated phenomenon of batch user value recovery. Periods with generally low return flow density values suggest that user hierarchical flow is dominated by a continuous downward shift during this stage, with rare return flow phenomena. The low density values in these periods serve as a background benchmark to highlight the relative anomaly of periods with dense return flow. The density difference between the two types of periods on the hierarchical return flow distribution map visually quantifies the intensity of the return flow behavior.
[0025] For example, the step of locating high-value backflow concentration sections and generating a backflow direction feature set by using the hierarchical backflow distribution map includes: statistically analyzing the number of boundary crossings in adjacent time periods to generate a boundary crossing frequency distribution; locating the peak crossing frequency of each time period based on the boundary crossing frequency distribution to generate a crossing stability distribution; capturing the peak drift trend through the crossing stability distribution to generate a crossing boundary candidate set; and determining the concentrated location of backflow directions based on the crossing boundary candidate set to generate a backflow direction feature set.
[0026] A boundary crossing frequency distribution is generated by statistically analyzing the number of boundary crossings between adjacent time periods in the hierarchical return flow distribution map. A boundary crossing is defined as a change in a user's hierarchical affiliation across the hierarchical boundary line between two adjacent statistical periods. Crossing the boundary line upwards is recorded as a positive crossing, and crossing downwards is recorded as a negative crossing. The number of boundary crossings in each time period in the hierarchical return flow distribution map is calculated by separately counting the number of positive and negative crossings for all users within that time period. Time periods with a high number of positive crossings indicate that the overall user hierarchy is flowing towards higher value, while time periods with a high number of negative crossings indicate a significant decline in overall value. The positive and negative crossing counts for each time period are arranged chronologically to form the boundary crossing frequency distribution. The horizontal axis of the boundary crossing frequency distribution corresponds to the statistical time period number, and the vertical axis records the number of positive and negative crossings respectively. The intersection of the two frequency sequences on the time axis represents the moment when the direction of hierarchical flow changes. Before the intersection, the positive direction is higher than the negative direction, indicating that return flow is dominant; after the intersection, the negative direction exceeds the positive direction, indicating that decline is dominant. When the peak value of the positive frequency distribution of boundary crossings is significantly prominent within a certain statistical period, and the return flow density of high-value layers is synchronously high in the corresponding time period's stratified return flow distribution map, the peak periods of the two distributions highly coincide, confirming the authenticity of the concentrated return flow phenomenon during that period. The sequence of positive and negative frequency differences in the boundary crossing frequency distribution for each time period reflects the intensity change of the net stratified flow direction; periods where the difference is consistently positive and the absolute value increases indicate intervals where the return flow momentum continues to strengthen.
[0027] The peak crossing frequency in each time period is located based on the boundary crossing frequency distribution to generate a crossing stability distribution. Local maxima are determined from the positive crossing frequency sequence of the boundary crossing frequency distribution; nodes with frequencies higher than the two preceding and following time periods are identified as peak crossing frequency nodes. The absolute value of the crossing frequency at each peak node and the net difference in frequency during the same period jointly measure the intensity of the peak. When multiple peak nodes exist in the boundary crossing frequency distribution, the interval between adjacent peak nodes reflects the periodic stability of crossing behavior. Evenly spaced peaks indicate a regular rhythm in crossing behavior, while uneven intervals suggest that crossing behavior is driven by sporadic factors rather than a stable user-level recovery pattern. The crossing stability S_stable is determined by the formula S_stable = 1 / (1 + σ_interval), where σ_interval is the standard deviation of the interval length sequence between adjacent peak nodes after normalization to the mean of the sequence. It is a dimensionless value. The smaller the σ_interval, the more uniform the interval. The closer S_stable is to 1, the more stable the crossing rhythm. When σ_interval approaches zero, S_stable approaches 1. When σ_interval increases infinitely, S_stable approaches zero. The crossing frequency intensity of each peak node and the corresponding S_stable value together constitute the stability description of that node. The stability descriptions of all peak nodes are arranged in chronological order to form the crossing stability distribution. For periods where the number of positive traversals forms multiple peaks with evenly spaced intervals, σ_interval is close to zero and S_stable is high. The stability distribution of traversal is high at these peak positions, indicating that the recovery of user levels during this period exhibits a regular rhythm. For periods with extremely uneven peak intervals, σ_interval is large and S_stable is low. The traversal stability distribution fully records the stability differences between the two types of periods. The S_stable value is consistently high in segments with significant regular rhythms, while the S_stable value is consistently low in segments dominated by occasional fluctuations. The stability of the two types of periods forms a clear numerical stratification in the traversal stability distribution.
[0028] Crossing the stability distribution captures peak drift trends to generate a candidate set of crossing boundaries. Systematic shift is the core of peak drift trend identification. If the interval between adjacent peak nodes shows a monotonically increasing or decreasing trend in three or more consecutive adjacent pairs, a peak drift trend is identified. A monotonically increasing trend indicates an extending crossing behavior cycle, while a monotonically decreasing trend indicates a compressing crossing behavior cycle. Peak node groups exhibiting peak drift trends in the stability distribution are extracted as candidate set entries. Each entry includes three descriptions: drift direction, drift amplitude, and the sequence of peak nodes involved. The drift direction distinguishes between cycle extension and cycle compression. Cycle compression corresponds to an accelerated recovery pace at the user level, a signal of accelerated return flow at high-value levels. Sections with generally low stability in the stability distribution indicate irregular peak distribution, making drift trend detection invalid, and no candidate set entries for crossing boundaries are generated for these sections. All peak node groups exhibiting drift trends are aggregated to form a candidate set for crossing boundaries. Peak node groups with continuously shortening intervals between five consecutive adjacent pairs clearly record a periodic compression trend in their stability values within the crossing stability distribution. The corresponding candidate set entries for crossing boundaries are labeled as periodic compression type, indicating that the pace of users crossing from low-value to high-value layers is accelerating within the time period of this node group. The number of candidate set entries reflects the prevalence of peak drift in the crossing stability distribution. A dense number of entries indicates significant changes in crossing rhythm across multiple time periods within the statistical period. The drift amplitude parameter within each entry indicates the strength of drift differences among node groups; entries with larger drift amplitudes correspond to node groups with more drastic changes in crossing cycles.
[0029] Based on the candidate set of boundary crossings, a feature set of return flow directions is generated to determine the concentrated location of return flow directions. The drift amplitude of each periodically compressed entry in the candidate set is jointly extracted with the sequence of peak nodes involved. Entries with larger drift amplitudes show a more significant acceleration in the return flow rhythm within their corresponding time period, thus gaining higher confirmation weight in the overlay verification. The return flow density values of the hierarchical return flow distribution map within the corresponding time period of each entry are simultaneously extracted for overlay verification. When the periodically compressed time period identified by the candidate set of boundary crossings highly overlaps with the high-density section of the hierarchical return flow distribution map, both criteria jointly confirm the location as the concentrated location of return flow directions; if the overlap is low, the entry in the candidate set of boundary crossings is downgraded and not considered strong evidence of a concentrated location. After confirming the concentrated location of return flow directions, the main directions of user traversal from each starting level to the high-value level within the corresponding time period are extracted. The main directions are determined by the path from the starting level to the high-value level with the most traversals. The path description includes the starting level label and the target level label of the high-value level; both are jointly used to determine the return flow direction at this concentrated location. The time range, main direction of return flow, and corresponding peak return flow density of each confirmed location of return flow direction concentration constitute the return flow direction feature set entries. All entries are arranged in time sequence to form the return flow direction feature set. When the path with the most crossings within a certain concentration location is a crossing from low-value layer to high-value layer, and the density value of the layer return flow distribution map is high during that time period, after the superposition of the two types of signals, this entry is marked in the return flow direction feature set as a large-scale concentrated return flow from low-value layer to high-value layer. This is significantly different in scale and starting layer from another concentration location during the same period, where the low-value layer crosses slightly into the high-value layer. The differences between the two types of entries in terms of starting layer, return flow scale, and crossing stability are fully recorded in the return flow direction feature set. The drift amplitude parameter carried by each entry further indicates the strength of the change in return flow rhythm at each concentration location.
[0030] A return flow rate gradient processing method is applied to the return flow direction feature set to generate a return flow acceleration distribution. The return flow rate is defined as the change in the number of users returning from a certain starting level to a high-value level within a single statistical period. The change is taken as the difference between the return flow density values of two adjacent statistical periods; a positive difference indicates an increase in the return flow rate during that period, while a negative difference indicates a decrease. The larger the absolute value, the more drastic the rate change. The return flow rate change sequence of each concentrated segment in the return flow direction feature set is further calculated using the difference between adjacent rates to obtain acceleration values. Positive acceleration indicates that the return flow is accelerating, while negative acceleration indicates that the return flow is decelerating. The sequence of absolute acceleration values reflects the rhythm of the change in the strength of the return flow momentum; periods with consistently high absolute acceleration values indicate that the user level has the strongest recovery momentum during those periods. The acceleration value sequences of each segment in the backflow direction feature set are arranged sequentially according to the start and end times of the segment. Significant differences in acceleration distribution characteristics exist within each segment. Some segments exhibit a single-peak pattern of initial acceleration followed by deceleration, indicating a clear peak and decline point in the backflow behavior. Other segments show stable acceleration fluctuations, suggesting a steady backflow without a significant acceleration period. Segments where the acceleration sequence shows a clear positive peak in the middle followed by a rapid decline form an acceleration peak in the backflow acceleration distribution. In contrast to segments with uniform acceleration distribution, locations with persistently high acceleration indicate the periods when user value recovery momentum is most concentrated. The sum of all segment acceleration value sequences, arranged by segment number, constitutes the backflow acceleration distribution. The densely distributed locations of high positive acceleration values in the backflow acceleration distribution identify the spatiotemporal regions where user value recovery momentum is most concentrated.
[0031] Based on the reflux acceleration distribution, accelerated reflux regions are located to generate stratified features. The key to determining accelerated reflux regions lies in two simultaneously occurring conditions: the acceleration value is consistently positive and its absolute value exceeds one standard deviation of the mean of the acceleration sequence for that segment. A continuous period satisfying these two conditions is confirmed as an accelerated reflux region. Within this region, the user level recovery momentum is at a peak, and the corresponding user group's level affiliation is most representative during this period. After extracting the start and end times and acceleration peak values of each accelerated reflux region from the reflux acceleration distribution, and combining this with the reflux starting level distribution of the corresponding segment in the reflux direction feature set, the actual value level affiliation of the user group within this accelerated reflux region is jointly determined. The actual value level affiliation considers the magnitude and rate of user acceleration from lower levels, correcting for user value states that operational tags fail to accurately reflect. The actual value level affiliation label for each accelerated reflux region and the corresponding user identifier set constitute stratified feature entries. In stratified features, the same user may have multiple level affiliation labels due to crossing multiple accelerated reflux regions in different statistical periods; the label of the most recent accelerated reflux region is taken as the representative stratified feature for that user. Users in the accelerated return flow region, identified by the return flow acceleration distribution, who rapidly recover from the low-to-medium value layer to the high-value layer, have their stratified feature attribution corrected to the mid-to-high value layer. This clearly corrects their original low-activity status description in the contradictory validation set. The set of stratified feature entries for all users in the accelerated return flow region forms a systematic correction for the users with the deepest stratified description deviation in the contradictory validation set. The difference in rank between these users after stratified feature correction and their original RFM quantile attribution intuitively demonstrates the supplementary role of behavioral data in portraying the true value status of users.
[0032] Step S13: Determine the hierarchical affiliation label based on the hierarchical features, extract incentive dependency risk features from user behavior data to generate incentive dependency risk identifiers, and perform incentive non-response clustering hierarchically based on incentive dependency risk identifiers and hierarchical affiliation labels to generate user value maps.
[0033] Specifically, tiered affiliation labels are determined based on tiered features. The core principle of prioritizing correction is to determine tiered affiliation labels—users covered by tiered features are assigned tiers based on their actual value, while users not covered by tiered features are directly assigned tier labels based on the current quantile of the RFM vector in the user baseline feature set. When there is a conflict between the tiered feature correction result and the RFM quantile assignment, the tiered feature correction result prevails. Conflicting users have correction annotations added to their tiered affiliation label entries to distinguish them from directly affiliated users who have not undergone correction. The correction annotations record the direction of change in tiered affiliation before and after correction, along with the time coordinates of the corresponding accelerated return flow region in the tiered features, preserving traceable information of the correction basis. Tiered affiliation labels cover all users in four tiers: high value, medium-high value, medium-low value, and low value. The proportion of users in each tier within the tiered affiliation label distribution directly reflects the current value structure of the user pool. The ratio of corrected-annotated users to directly affiliated users within the same tier further reveals the support strength of behavioral data and historical transaction data in the tier determination basis. A higher proportion of corrected-annotated users indicates that the tier determination relies more on the behavioral correction results of the accelerated return flow region. When the stratified feature correction results within a certain statistical period reclassify a group of users originally labeled as low-activity users into the medium-high value stratum, the stratification attribution label is corrected in the level allocation of this group of users. This group of users jumps from the low-value label corresponding to the original quantile of the user baseline feature set to the medium-high value stratum. The corrected stratification attribution label is closer to the current state of the actual value recovery of this group of users. The time coordinates of the accelerated return flow area recorded in the correction annotation provide a verifiable source of behavioral evidence for the verification of the discrepancy between cluster weight allocation and repeat purchase.
[0034] Incentive dependency risk features are extracted from user behavior data to generate incentive dependency risk labels. These features describe the behavioral response patterns of users to promotional activities and incentives. If a user's historical transaction records are highly concentrated during promotional periods rather than in daily life, their willingness to consume is naturally low, and their consumption behavior is easily interrupted once the incentive ends. The transaction records of each user in the consumption record data are divided into two categories based on whether they occurred during promotional periods: transactions during incentive periods and transactions during non-incentive periods. The ratio of transactions during incentive periods to those during non-incentive periods constitutes the quantitative basis for incentive dependency risk features. The incentive dependency rate R_dep is determined by the formula R_dep=N_promo / (N_promo+N_natural), where N_promo is the number of transactions during incentive periods, N_natural is the number of transactions during non-incentive periods, and R_dep ranges from 0 to 1. Push click response records in user behavior data serve as an auxiliary corroborating dimension for incentive dependency risk features. Users with high click response rates but who only convert during incentive periods, along with a high R_dep, provide double corroboration, making the degree of incentive dependency more credible. Users with an R_dep value exceeding the user mean plus one standard deviation are classified as having high incentive dependency risk; those between the mean and the mean plus one standard deviation are classified as having medium incentive dependency risk; and those below the mean are classified as having low incentive dependency risk. The R_dep value of each user is paired with their risk level to form an incentive dependency risk identifier. This identifier records the R_dep value and corresponding risk level using the user identifier as an index. All user entries are sorted in descending order of R_dep. Users whose entire transaction record within the past 180 days occurred during platform promotional periods and who have no natural transaction records outside of promotional periods have an R_dep of 1.0, the highest among all users. This user is classified as having a high risk level in the incentive dependency risk identifier, indicating that their value maintenance is highly dependent on continuous incentive investment, and the risk of churn is extremely high if the incentive cycle is interrupted. High-risk users with R_dep values at the tail end of the user distribution often appear in the statistical period after the incentive activity ends. During this period, the consumption record data of high-incentive-dependent users shows a sharp contraction, forming a clear stratification with the normal fluctuations of low-R_dep users.
[0035] In some embodiments, the step of generating a user value map by implementing incentive-inactive clustering based on the incentive dependency risk identifier and the hierarchical affiliation label includes: assigning label weights to the hierarchical affiliation label to generate a weighted affiliation label set; performing initial clustering based on the weighted affiliation label set to generate an initial clustering result; locating repeat purchase deviation user samples from the initial clustering result to generate a deviation frequency distribution; and using the deviation frequency distribution and the incentive dependency risk identifier to perform hierarchical correction on the initial clustering result to generate a user value map.
[0036] A weighted attribution label set is generated by assigning label weights to hierarchical attribution labels. The core starting point for weight allocation is the difference in credibility—user labels corrected by hierarchical features provide behavioral evidence to accelerate the return flow region, and their credibility is significantly higher than user labels directly attributed based solely on RFM quantiles. This difference is quantified through weight values, giving high-credibility labels stronger pull in shaping cluster centers. The base weight value for the hierarchical attribution label of corrected-label users is set to 0.8, and the base weight value for uncorrected directly attributed users is set to 0.5. These two base values are added to the normalized M-value in the user's RFM vector as a consumption contribution correction term. Before addition, the normalized M-value is compressed to the 0-1 range; after addition, the value is normalized to the 0-1 range for all users, forming the final weight value for each user. The weight distribution within each level of hierarchical attribution labels reflects the proportion of high-credibility labels in the user group of that level. When the proportion of corrected-label users is high in the high-value layer, the overall weight of that level is relatively high, indicating sufficient basis for judging the high-value layer; when directly attributed users dominate in the low-value layer, the overall weight is relatively low. All users' hierarchical level labels and corresponding weight values are paired to form a weighted attribution label set. High-value users with higher M-value normalized values after hierarchical feature correction have higher weights in the weighted attribution label set, and their corresponding cluster centers have stronger pull during the initial clustering stage. High-value users who are directly assigned based solely on RFM quantiles and have lower M-value normalized values have significantly lower weights, and their labels have a correspondingly limited influence on cluster centers during the initial clustering.
[0037] Initial clustering results are generated based on a weighted attribution label set. The prior constraint on the number of clusters comes from the total number of hierarchical attribution label levels, which is fixed at four. The user group labeled at each level directly serves as the initial member set of the corresponding cluster. The weighted mean of the feature vectors of all users at each level is used as the initial cluster center. The weights of the weighted mean are directly taken from the weight values of each user in the weighted attribution label set; higher-weighted users exert greater pull on the initial cluster center. After the initial cluster centers are determined, a weighted Euclidean distance is calculated between the feature vectors of all users and each center. Each user is assigned to the cluster corresponding to the nearest center. After assignment, each cluster center is updated to the weighted mean of the current member's feature vector. This assignment and update process is repeated until the displacement of each cluster center is less than 0.1% of the global maximum inter-cluster distance for two consecutive rounds, at which point convergence occurs. The weight differences in the weighted attribution label set allow users with high confidence labels to dominate the convergence direction of the cluster centers, while low-weight users are directly assigned and passively adapt to the center positions shaped by high-weight users. This ensures that the category boundaries of the initial clustering results are closer to the distribution of user groups supported by behavioral evidence. In the initial clustering results, each user's cluster affiliation number and corresponding intra-cluster distance are recorded as entries. Users with large intra-cluster distances indicate that their feature vectors have a low degree of matching with their respective cluster centers. During the repurchase deviation test, the reliability of the affiliation of such users needs to be verified first. Users who are classified into high-value clusters but whose intra-cluster distances are among the top 10% of the largest in that cluster have a significant deviation between their feature vectors and the high-value cluster centers. During the repurchase deviation test, the verification of such users has the highest priority.
[0038] This study identifies users exhibiting repurchase divergence from the initial clustering results and generates a divergence frequency distribution. Repurchase divergence is defined as a user initially classified into a high-value cluster, but whose recent consecutive statistical periods show a continuously increasing repurchase interval and a systematic decline in transaction amount, indicating that their actual consumption behavior is contrary to the typical characteristics of the high-value cluster. For each member of the high-value and mid-to-high-value clusters in the initial clustering results, the repurchase interval trend is examined. Users whose repurchase intervals monotonically increase for three or more consecutive statistical periods and whose cumulative transaction amount decreases by more than 20% of the mean at the time of clustering are identified as repurchase divergence users. Users with larger intra-cluster distances are prioritized in the repurchase divergence test. Users whose intra-cluster distance is among the top 10% and who also meet the condition of monotonically increasing repurchase intervals have a higher confidence level in the divergence determination. The proportion of repeat purchase users within each cluster level is arranged chronologically by statistical period to form a deviation frequency distribution. The horizontal axis of the deviation frequency distribution corresponds to the statistical period number, and the vertical axis corresponds to the deviation proportion of each cluster level. Periods with a continuously rising deviation proportion indicate that the deviation between users' actual consumption behavior and cluster affiliation within that cluster is intensifying, and the accuracy of the cluster boundaries in describing these users is decreasing. Periods where the deviation proportion changes abruptly within a single statistical period corroborate the changes in the intra-cluster distance distribution in the initial clustering results during the same period. The reliability of the deviation judgment is highest when both signals deviate simultaneously. When the deviation proportion of a high-value cluster suddenly increases and a large number of users with large intra-cluster distances are added to that cluster during the same period, the superposition of the two signals indicates that the high-value cluster during that period is being interfered with by users whose repeat purchase behavior has declined, and the cluster center is being pulled in an atypical direction.
[0039] The initial clustering results are hierarchically corrected using deviation frequency distribution and incentive dependence risk indicators to generate a user value map. Correction trigger nodes are anchored to the peak period of deviation ratio between high-value and medium-high-value clusters in the deviation frequency distribution. A higher deviation ratio indicates a larger proportion of users with inaccurate clustering in the initial clustering results during that period. The intersection of the confirmed repeat purchase deviation user samples and the high-risk users in the incentive dependence risk indicators during the peak period is taken. These users possess both repeat purchase behavior decline and incentive dependence risk indicators, and are corrected with the highest priority. Users who only meet the repeat purchase deviation condition but have a low incentive dependence risk indicator risk level are corrected with the next highest priority. Users with a high incentive dependence risk indicator but who have not triggered a repeat purchase deviation judgment are not corrected, are included in the observation list, and are returned to the deviation frequency distribution process for verification in the next statistical period. High-priority users are moved down two levels from their current cluster level to reflect the superposition of dual risks, while medium-priority users are moved down one level to correspond to the magnitude of a single risk. After both types of corrections are completed, the mean of the feature vectors of each cluster center as the corrected members is recalculated. The updated cluster centers are then used to re-assign all users in the initial clustering results, forming a stable cluster distribution after correction. The user value map uses the final cluster assignment level as the main index, with additional risk level and correction status labels for each user's corresponding incentive dependency risk. All user entries are arranged from highest to lowest cluster level to form the user value map. Users in the original high-value cluster that simultaneously have evidence of both repurchase deviation and high incentive dependency risk are moved down two levels to the medium-low value layer in the user value map and labeled with high incentive dependency risk, indicating that the user's ability to maintain value is questionable. Users in the medium-high value cluster that only trigger repurchase deviation and have low incentive dependency risk are moved down one level to the medium-low value layer in the user value map. The difference in the entry structure of the two types of users in the user value map reflects the difference in the source of the correction basis.
[0040] Step S14: Based on the user value map and user baseline feature set, perform hierarchical exit quantification to generate exit magnitude value, extract the distribution of silence duration of high-value users from the user value map to generate silence warning interval, and perform churn trigger node positioning on exit magnitude value to generate churn prediction coefficient.
[0041] Specifically, exit magnitude values are generated through hierarchical exit quantization based on the user value map and the user baseline feature set. The core of hierarchical exit quantization is to compare the difference between the current cluster affiliation level of each user in the user value map and the corresponding hierarchical position of the initial statistical period RFM vector quantile in the user baseline feature set. This difference reflects the net change in hierarchical level for that user during the complete observation period. The absolute number of levels downgraded multiplied by the average distance between feature vectors between each level yields the user's hierarchical exit quantization value. Users who downgrade by more levels and have larger distances between levels have higher exit quantization values. Users with added correction labels in the user value map participate in exit quantization with the corrected level as their current hierarchical level, ensuring that the exit magnitude reflects the true change after behavioral correction. Directly affiliated users without correction directly participate in exit quantization with the level corresponding to the RFM quantile in the user baseline feature set. The quantization methods for these two types of users are distinguished by source labels in the exit magnitude value entries. The exit quantification values for each user's tier are normalized to form an exit magnitude value. Users with high exit magnitude values indicate a significant decline from a higher tier during the observation period, while users with low values indicate that their tiers have remained relatively stable or have shifted upwards. Former high-value users who continuously shifted three tiers down to a lower-value tier during the observation period had higher exit quantification values and correspondingly larger exit magnitude values among all users. For former mid-to-low-value users who only shifted one tier down, their exit magnitude values were at a moderate level. The difference between these two types of users is directly related to their starting tier and the magnitude of their shift. Users with high starting tiers and large shifts have the most significant impact on operational value loss.
[0042] In some embodiments, the step of extracting the distribution of silence duration of high-value users from the user value graph to generate a silence warning interval includes: classifying the user activity of the user value graph to generate an activity stratification sequence; selecting high-value silent users from the activity stratification sequence to generate a silent user set; performing cumulative silence duration analysis on the silent user set to generate a silence cluster distribution; and identifying critical silence nodes based on the silence cluster distribution to generate a silence warning interval.
[0043] The user value graph is used to classify user activity levels and generate a hierarchical activity sequence. The activity level of users at each level of the user value graph is quantified by a three-dimensional weighted model A_score = 0.4 × B_norm + 0.35 × P_norm + 0.25 × T_norm. Browsing frequency (B_norm), push notification clicks (P_norm), and platform session duration (T_norm) are integrated into a single activity score. These three behavioral indicators characterize the user's platform participation from three dimensions: initiation frequency, reach response, and deep engagement. The A_score ranges from 0 to 1, with browsing frequency having the highest weight, reflecting the dominant role of behavior initiation frequency in judging activity level. The A_score is divided into four levels based on the quantile of all users: high activity, medium activity, low activity, and inactive. A score above the 75th percentile is high activity, 50th to 75th percentile is medium activity, 25th to 50th percentile is low activity, and below the 25th percentile is inactive. In the user value graph, each user's cluster affiliation level is paired with their corresponding A_score activity level. The activity level changes of all users in each statistical period are recorded sequentially to form an activity stratification sequence. If, in a certain statistical period, the proportion of inactive users in the high-value stratification sequence suddenly increases while the proportion of inactive users in the low-value stratification remains relatively stable, the significant difference in the inactive proportions between the two strata suggests that the high-value strata are facing more concentrated activity decline pressure during that period, and some high-value users may have entered a state of inactivity and observation before churn. In the activity stratification sequence, if the same user maintains an inactive level for multiple consecutive statistical periods but remains in the high-value stratification, they are a high-priority target for identifying inactive users. While their consumption records have not yet triggered a stratification shift, the decline in their behavior has already given an early warning signal. The stability of their cluster affiliation depends on historical transaction accumulation rather than the continued support of current activity.
[0044] A set of silent users is generated by selecting high-value silent users from the activity stratification sequence. In this paper, high-value silent users refer to those whose user value graph clustering level is in the high-value or medium-high-value layer and whose activity level in the current statistical period of the activity stratification sequence is silent. Both conditions must be met simultaneously. Users who are only silent in activity but whose clustering level is in the low-value layer are not included in the silent user set, as the silence of low-value users has a limited impact on the overall operational value loss. User identifiers that meet both conditions in each statistical period of the activity stratification sequence are extracted and summarized by period. Each member of the silent user set is labeled with an incentive dependency risk level. High-value silent users with high incentive dependency risk receive higher attention weight in the cumulative analysis of silent duration, and the temporal relationship between their silent behavior and the incentive termination point is tracked in the cumulative analysis. When a high-value user remains in a silent state for multiple consecutive statistical periods in an activity stratification sequence, the silent user set includes the user's identifier during these periods. The persistence of the silent state forms a continuous record in the user set, indicating that the user has entered a stage of persistent activity loss rather than a brief session interruption. The span of the continuous record directly determines the user's position in the subsequent cumulative analysis of silent duration. When a medium-to-high-value user triggers the silent judgment only once in a single statistical period and then immediately resumes a medium-activity level, the single trigger forms an isolated record in the silent user set, indicating a brief decline in activity. The difference in the continuity of records between the two types of users in the silent user set directly affects the duration calculation results in the cumulative analysis of silent duration. The cumulative duration of continuously silent users continues to increase as the statistical period progresses, while the silent duration of isolated record users is recalculated after an interruption.
[0045] A cumulative silence duration analysis is performed on the set of silent users to generate a silence cluster distribution. The current silence duration for each user is calculated by multiplying the number of consecutive statistical periods by the number of days in the statistical period. When a consecutive period is interrupted and then resumes, the silence duration is recalculated, and the cumulative silence duration is the duration of the most recent consecutive silent period. The cumulative silence duration for each user in the set is segmented into 14-day intervals. The number of users within each interval forms the silence cluster density for that interval. The silence cluster densities of all intervals are arranged from shortest to longest silence duration to form a silence cluster distribution. This distribution reveals the concentration of high-value silent users across different silence duration stages. Intervals with significantly high silence density represent a large concentration of silent users within that silence duration and are priority targets for identifying critical silence nodes. The distribution pattern of density values across intervals reflects the natural clustering pattern of high-value silent users across different silence duration stages. When the density of a specific silent duration interval in a certain silent cluster distribution is significantly higher than that of adjacent intervals, and the density drops sharply after that interval, it indicates that there is a clear differentiation node in the silent state of a large number of users within that duration. Some users recover their activity after receiving effective intervention before this duration, while the recovery probability of other users continues to decline after the silent duration crosses this interval. The inflection point between the high-density interval before the sharp drop and the low-density interval after the sharp drop is the core signal for locating the critical silent node. The larger the density difference on both sides of the inflection point, the more significant the segmentation effect of this critical duration on the probability of activity recovery.
[0046] Based on the silent clustering distribution, critical silent nodes are identified to generate silent warning intervals. A critical silent node is defined as an inflection point in the silent clustering distribution where the silent density begins to decline significantly from a high level. Before this inflection point, the user group within the silent duration interval still has high potential for activity recovery; after the inflection point, the probability of activity recovery decreases continuously with the extension of the silent duration. The absolute value sequence of the density difference between adjacent intervals in the silent clustering distribution is extracted. The position where the difference changes from positive to negative and the absolute value exceeds one standard deviation of the mean of the entire interval density sequence is determined as the critical inflection point. The lower bound of the silent duration interval corresponding to the critical inflection point is used as the critical silent node. The specific number of days for the critical silent node is dynamically revised as the silent clustering distribution is updated. Changes in the behavioral patterns of high-value silent users in different statistical periods can cause the critical node position to drift. Extending the critical silent node by one interval length in the direction of shorter silent duration forms a warning pre-window, and extending it by one interval length in the direction of longer silent duration forms a warning post-window. The pre-window and post-window together constitute the silent warning interval. When high-value users enter the pre-window of the silent warning interval, a highest priority warning is triggered; when mid-to-high-value users enter the pre-window, a secondary warning is triggered. When a high-value user's cumulative silence time has entered the pre-emptive window and the incentive dependency risk indicator simultaneously displays a high incentive dependency risk, the combination of these two conditions ensures that the user receives the highest priority for intervention during the operational decline indicator generation phase. The dual superposition of silence time and incentive dependency risk indicates that once the user exceeds the critical point, the difficulty of recovery will be further amplified due to the inertia of incentive dependency. When a medium-to-high-value user's silence time has just entered the pre-emptive window and the incentive dependency risk is low, a secondary warning is issued. The boundary of the silence warning interval is revised synchronously with the silence cluster distribution update cycle.
[0047] In some embodiments, the step of performing churn trigger node location and generating churn prediction coefficient on the exit amplitude value includes: performing hierarchical net outflow statistics on the exit amplitude value to generate a net outflow time series; performing sudden increase moment detection on the net outflow time series to generate a sudden increase node set; performing churn trigger interval location based on the sudden increase node set to generate a node distribution map; and determining the churn risk level and generating a churn prediction coefficient according to the node distribution map.
[0048] The exit magnitude values are statistically analyzed to generate a time series of net outflows. Within each statistical period, the exit magnitude values of all users who moved down a level are summed, minus the sum of the recovery magnitude values of users who moved up a level during the same period. The recovery magnitude value is calculated symmetrically to the exit magnitude value, i.e., the number of upward moves is multiplied by the average distance between the characteristic vectors of each level, and then normalized to all users. The net outflow L_net is determined by the formula L_net=Σ(V_down_i)-Σ(V_up_j), where V_down_i is the exit magnitude value of the i-th downward-moving user, V_up_j is the recovery magnitude value of the j-th upward-moving user, and L_net is a positive or negative real number. Its absolute value dynamically changes with the size of the user group, reflecting the overall magnitude of net tiered flow during that period. A positive L_net indicates a net downward flow of the entire tier during that period, while a negative L_net indicates a net upward recovery. The L_net values for each statistical period are arranged chronologically to form a net outflow time series. Periods where L_net is consistently positive and increasing indicate an accelerating loss of user value. The significant contribution of V_down_i from users leaving the high-value tier makes the net outflow time series more sensitive to high-value tier churn events. Periods with concentrated downward movement of high-value tier users will produce a prominent peak in the net outflow time series. In one statistical period, when a large number of high-value tier users move downward in the user value map, the cumulative value of V_down_i is significantly higher than the cumulative value of V_up_j. During this period, L_net shows a prominent positive deviation, indicating a significant net outflow from the high-value tier. In another statistical period, when the number of users moving downward and upward in the high-value tier is roughly equal, L_net approaches zero, and the net outflow time series shows a stable trend during this period. The comparison of the two time series characteristics intuitively reflects the differences in the intensity of user tier flow in different periods.
[0049] A sudden increase moment detection is performed on the net outflow time series to generate a sudden increase node set. The rate of change of net outflow is obtained by dividing the difference of net outflow in adjacent statistical periods in the net outflow time series by the time interval. Periods with a positive rate of change and an absolute value exceeding twice the standard deviation of the global mean are identified as sudden increase moments, representing an accelerated and concentrated outbreak of user value churn. When two or more consecutive periods in the net outflow time series meet the sudden increase condition, these periods are grouped into a sudden increase interval. The start time, end time, and peak net outflow within the interval constitute the sudden increase description entry for that interval. Isolated sudden increase moments are retained as single-point entries, indicating that the net outflow in that period experienced a one-time surge followed by a decline, belonging to an occasional concentrated churn. The set of all sudden increase intervals and isolated sudden increase moments constitutes the sudden increase node set. The ratio of sudden increase intervals to isolated sudden increase moments in the sudden increase node set intuitively reflects whether the churn event is dominated by a continuous acceleration or an occasional concentrated type. When multiple surge intervals appear consecutively in the node set of a certain statistical period, and the peak value of net outflow in each interval shows an increasing trend, it indicates that the damage to the user pool caused by each wave of churn is accumulating and amplifying. The larger the increase in the peak value of the surge interval, the more likely that the churn momentum is continuously accumulating between each surge rather than fluctuating randomly. In contrast, the concentration and persistence of churn events are significantly weaker in statistical periods with only a few isolated surge moments. The difference in the entry structure of the surge node set between the two types of cases fully records the essential difference in the nature of churn in different statistical periods.
[0050] A node distribution map is generated based on the location of churn trigger intervals using a set of rapidly increasing nodes. After spreading out the churn intervals and isolated churn moments in the churn node set along the time axis, an extension boundary is formed by extending one statistical period to each side of the start and end times of each churn interval. This extension operation preserves the warning period before the churn and the continuation period after the churn. When the interval between two adjacent extension boundaries is less than one statistical period, they are merged into a single continuous churn trigger interval. The position and duration of each churn trigger interval on the time axis, along with the peak net outflow within the corresponding interval of the churn node set, together form a node marker in the node distribution map. The node marker formed by isolated churn moments is typically smaller than the marker corresponding to churn intervals. The two types of churn trigger intervals are distinguished by different labels in the node distribution map, and the size of the node markers visually reflects the intensity differences of each churn concentration event. All node markers are arranged in chronological order to form a node distribution map. Periods with dense node markers indicate frequent and concentrated churn events, while periods with sparse node markers indicate relatively stable user tier flow. Within a six-month statistical period, node markers are concentrated in the last two months of the quarter. This regular distribution characteristic reveals the seasonal concentration pattern of churn risk in this business scenario. The density of node markers at the end of the quarter and the distribution pattern of the peak net outflow corresponding to each marker together depict the systematic boosting effect of the end-of-quarter incentive rhythm on the net outflow of high-value tiers.
[0051] Churn prediction coefficients are generated based on the node distribution map to determine the churn risk level. The node level in the node distribution map is jointly determined by two indicators: peak net outflow and node duration. A node is marked as high-risk if its peak exceeds twice the global mean standard deviation and its duration exceeds three statistical periods; a node meeting only one of these criteria is marked as medium-risk; and a node not meeting either criterion is marked as low-risk. The churn prediction coefficient P_churn for all users within the corresponding statistical period for each risk level node is determined by the formula P_churn=w_n×G_node+w_layer×G_layer+w_d×G_dep, where G_node is the normalized value of the node risk level, G_layer is the normalized value of the user value map hierarchy affiliation level, G_dep is the normalized value of the incentive dependency risk identifier risk level, and the weight coefficients of w_n, w_layer, and w_d are 0.4, 0.35, and 0.25, respectively. The value of P_churn ranges from 0 to 1. In the node distribution diagram, during high-churn-risk periods, users in the high-value tier with high incentive dependence risk have all three components at high levels, with P_churn at its highest among all users. Users in the low-value tier with low incentive dependence risk have low P_churn. The intermediate value is determined by natural interpolation between these two extreme combinations. The P_churn value of intermediate users exhibits a continuous distribution depending on the specific combination of the normalized values of the three components. Different tiers and different degrees of incentive dependence form a clear gradient stratification in the P_churn distribution. The difference in P_churn coefficient between high-value, high-incentive-dependency users and low-value, low-incentive-dependency users during a high-churn-risk period clearly quantifies the combined amplification effect of tier affiliation and incentive dependence on churn risk. All user churn prediction coefficients are arranged in descending order of P_churn.
[0052] Step S15: Perform joint level matching between the churn prediction coefficient and the hierarchical affiliation label to generate an abnormal churn pattern; perform push channel blocking feature monitoring on the silent warning interval to generate an operational decline indicator; and perform reach failure gradient matching based on the operational decline indicator and the abnormal churn pattern to generate a user operation monitoring report.
[0053] Specifically, the churn prediction coefficient and hierarchical affiliation label are used for joint level matching to generate abnormal churn patterns. The combination of both high churn prediction coefficients and high hierarchical affiliation labels results in the greatest operational value loss—users with both high churn prediction coefficients and high hierarchical affiliation labels are the preferred abnormal churn type identified by joint level matching. The 4×4 combination matrix is formed by the intersection of the quartile range of the churn prediction coefficient distribution across all users and the four levels of hierarchical affiliation labels. Combinations with churn prediction coefficients in the upper quartile and hierarchical affiliation labels in the high-value or mid-to-high-value tier are classified as abnormal churn pattern A. Combinations with churn prediction coefficients in the upper quartile but hierarchical affiliation labels in the low-value tier are classified as abnormal churn pattern B. Although B-type users have low hierarchical affiliation, their high prediction coefficients indicate a recent abnormal and rapid decline. Combinations with churn prediction coefficients in the middle half are classified as either observed abnormal or steady-state churn patterns based on their hierarchical affiliation label levels. Neither of these two types triggers active abnormal churn pattern labeling. In the abnormal churn pattern, the identifiers of Class A and Class B users, their corresponding churn prediction coefficient values, and matching types together constitute each user entry. All entries are arranged in descending order of churn prediction coefficient. High-value users are matched as Class A when their churn prediction coefficient is in the highest range among all users, and their operational outreach priority is marked as the highest urgency level in the user operation monitoring report. Low-value users are matched as Class B when their churn prediction coefficient is also high. Their previously high consumption behavior has recently shown an abnormal decline, and it is necessary to combine the incentive dependence risk indicator to determine whether it is a natural decay after the incentive stops. Class A and Class B entries are arranged side by side in the abnormal churn pattern. The difference in churn prediction coefficient values between the two types of users and the combination of their hierarchical affiliation level jointly determine the intervention priority they receive in the outreach failure gradient matching stage.
[0054] In some embodiments, the step of monitoring the push channel shielding characteristics of the silent warning interval to generate an operational decline identifier includes: performing time-series sampling on the silent warning interval to generate a reachable quantity sequence; calculating the active shielding frequency distribution based on the reachable quantity sequence to generate a shielding time-series distribution; extracting the shielding duration slope after a sudden increase through the shielding time-series distribution to generate a shielding recovery distribution; and determining the shielding concentration interval based on the shielding recovery distribution to generate an operational decline identifier.
[0055] A time-series sampling method is used to generate a reach availability sequence for the silence warning interval. The platform push system logs serve as the daily data source for channel availability. "Available" indicates that the channel's push notifications have not been blocked by users or the system has limited their frequency; "unavailable" indicates that reach is impossible due to user-initiated blocking or system frequency control. These two types of unavailability are distinguished by the "Source of Status Change" field in the logs. The total number of available push channels for each user each day is summed and arranged chronologically to form a daily reach availability sequence for that user. The average of the daily sequences for all users within the silence warning interval is used to form a group reach availability sequence. Periods of continuous decline in the group sequence indicate an increase in active blocking behavior within the group while it is in a silent state. If, within a statistical period, the reach availability sequence for the user group within the silence warning interval shows a continuous decline for two consecutive weeks, with the reduction in APP push channel availability being most significant, it suggests that a large number of users in this group have actively disabled APP notification permissions during the silence period. When the availability of SMS channels does not change significantly during the same period, the divergence in the trends of APP and SMS channel availability often reveals differences in blocking motives. APP blocking is usually directly related to excessive push frequency or high content repetition, while changes in SMS channel availability are more influenced by system frequency control strategies. The available volume of each push channel in the reach availability sequence is recorded independently. The trend differences between channels are fully preserved in the sequence in the form of channel-specific numerical pairs. When the available volume of a single channel drops sharply while other channels remain stable, it is presented in the sequence as a clear differentiation in the numerical values between channels.
[0056] The proactive blocking frequency distribution is generated based on the reach availability sequence to produce a blocking time series distribution. The proactive blocking frequency is obtained by dividing the number of unavailable channels caused by user proactive blocking in the push logs on which the reach availability sequence is based by the total number of all channels on that day. Unavailable records caused by system frequency control are excluded based on the status change source field to ensure that the blocking frequency only reflects user proactive defense behavior rather than platform policy restrictions. The proactive blocking frequency of all users is summarized by statistical day to form the proactive blocking frequency distribution. The values of each statistical day in the proactive blocking frequency distribution are arranged in chronological order to form the blocking time series distribution. The horizontal axis of the blocking time series distribution corresponds to the statistical day number, and the vertical axis corresponds to the proactive blocking frequency value of that day. Periods with a continuously rising blocking frequency indicate that the user group's resistance to push notifications is deepening, while periods with a stable blocking frequency indicate that the current outreach strategy has not triggered a significant blocking response. In the available reach sequence, the available reach for APP push and SMS channels corresponds to the blocking frequency calculation for each channel. The blocking frequencies for the two channels are recorded independently in the blocking time series distribution. The daily sequences of APP channel blocking frequency and SMS channel blocking frequency are arranged side by side in the blocking time series distribution. The comparison of the values of the two sequences during the same period reveals the differences in user resistance to different channels within the same time period. During the peak period of push activities, when the blocking frequency of APP channel continues to rise while the blocking frequency of SMS channel remains low during the same period, the divergence in the trends of the two sequences is shown in the blocking time series distribution as an increasing trend of the frequency difference between channels. The larger the difference during the same period, the more significant the difference in user acceptance of the two types of channels. The period when blocking behavior is concentrated on a specific channel forms an asymmetrical distribution pattern in the blocking time series distribution, with a prominent peak on one channel and a stable peak on the other.
[0057] For example, the step of extracting the shielding duration slope after a sudden increase from the shielding time series distribution to generate a shielding recovery distribution includes: performing sudden increase event location on the shielding time series distribution to generate a shielding sudden increase time sequence; extracting the duration interval after the sudden increase from the shielding sudden increase time sequence to generate a shielding duration subset; performing slope fitting calculation on the shielding duration subset to generate a shielding slope set; and performing distribution statistics based on the shielding slope set to generate a shielding recovery distribution.
[0058] A surge event localization method is implemented for the masking time series distribution to generate a masking surge time sequence. A surge event is identified when the absolute value of the difference between the active masking frequencies of adjacent statistical days exceeds twice the global mean and the direction is positive. The positive direction ensures that only nodes where the masking frequency accelerates are captured, avoiding misinterpreting the frequency recovery process as a surge. When two consecutive days meet the condition, they are merged into a single surge event, and the maximum masking frequency of the two days is used. Each surge event is arranged according to its occurrence date and corresponding masking frequency value to form a masking surge time sequence. Periods with densely occurring surge events in the sequence indicate that operational outreach has caused continuous harassment to users, while periods with sparsely occurring surge events indicate that user masking behavior has not experienced a concentrated outbreak during that phase. When a user group actively blocks content for two consecutive days after a push notification campaign ends, the two days are merged into a single surge event. This event is accurately presented as a single record in the surge event time sequence to avoid deviations caused by repeatedly including the same event in the subsequent slope fitting stage. Periods with a smooth distribution of blocking time without any dates triggering a surge judgment are not recorded in the surge event time sequence. The surge event time sequence entries are dense during the peak period of the push notification campaign and sparse during the stable period. The date range of the concentrated surge events and the time interval between the push notification campaign nodes are completely preserved in the date distribution of the sequence entries. The time interval between the surge event that immediately follows the push notification campaign node and the time interval between the surge event that occurs later are different in the entry time interval. The difference in the interval between the two types of situations reflects the degree of lag in the user's blocking response under different push notification intensities.
[0059] A sustained interval after the sudden increase in shielding intensity is extracted from the time series of shielding surge events to generate a sustained shielding subset. The sustained interval after the sudden increase for each surge event in the shielding surge time series is defined as the continuous period during which the active shielding frequency of the shielding time series distribution remains higher than the baseline level of the 7-day mean before the surge. The start date of the sustained interval is the day after the surge event, and the end date is the day before the first drop back to the baseline level. The length of the sustained interval is determined by the difference in the number of days between the end date and the start date. Cases where the frequency drops back to the baseline level the day after the surge event have a sustained interval length of zero and are marked as brief surges in the sustained shielding subset, distinguishing them from longer-duration true shielding spread events. These two types of entries are differentiated during the slope fitting stage. The duration intervals and lengths of all surge events in the time series of sudden increases in blocking are summarized to form the blocking duration subset. When the frequency of active blocking remains higher than the baseline for several consecutive days after a surge event, the duration interval of that item in the blocking duration subset is relatively long, reflecting that the surge triggered the spread of continuous user blocking behavior. When another surge event only lasts for two days and then the frequency drops, the duration interval of the corresponding item is relatively short. The surge events corresponding to items with long duration intervals often occur during periods of consistently high operational reach, and user blocking reactions do not naturally subside after the surge but continue to accumulate. Items with short duration intervals are more likely to correspond to situations where a single surge in push density is followed by a return to normal rhythm, and blocking behavior quickly converges after the push density drops. The difference in length between the two types of duration intervals is directly reflected in the number of days in the blocking duration subset.
[0060] A shielding slope set is generated by slope fitting calculation on the shielding duration subset. The input data for linear fitting is the daily active shielding frequency sequence within the duration interval of each surge event in the shielding duration subset. The shielding recovery slope K_rec = (f_end - f_start) / T_dur, where f_end is the shielding frequency on the last day of the duration interval, f_start is the shielding frequency on the first day, T_dur is the number of days in the duration interval, and K_rec represents the daily frequency change rate. The sign indicates the direction of shielding frequency development after the surge, and the absolute value quantifies the rate of change. A negative K_rec indicates that the shielding frequency is falling back after the surge, and a positive K_rec indicates that the shielding frequency is still rising after the surge. For entries in the shielding duration subset with duration intervals of less than three days, there are too few data points. Therefore, a low-reliability label is added to the K_rec of these entries in the shielding slope set, and they are given a lower weight in the distribution statistics. All surge events, along with their corresponding K_rec values and low-reliability labeling states, constitute a set of blocking slope records. When a surge event has a long duration and a large absolute value of negative K_rec, it indicates a rapid decline in blocking frequency, suggesting that user blocking behavior is quickly and naturally alleviated after the surge. Conversely, when another surge event has a positive K_rec, it indicates that the blocking frequency continues to rise after the surge, and the operational decline deepens further within that duration. Two entries with similar absolute K_rec values but opposite signs correspond to symmetrical blocking behaviors for their respective surge events: one recovers rapidly with the same slope, while the other continues to deteriorate at the same rate. This symmetrical distribution in the blocking slope set visually represents two extreme trends in user blocking behavior under different surge event triggering scenarios. Entries with smaller absolute K_rec values, regardless of whether they are positive or negative, indicate a gradual change in blocking frequency with no obvious trend after the surge. These entries contribute little to the density of both positive and negative groups during the distribution statistics phase.
[0061] The shielding recovery distribution is generated based on the distribution statistics of the shielding slope set. All K_rec values in the shielding slope set are grouped according to their positive and negative directions, forming two sets of numerical basis for the shielding recovery distribution. The negative slope group represents a recovery trend after a sudden increase in shielding frequency, while the positive slope group represents a worsening trend of continued increase in shielding frequency after a sudden increase. The frequency distribution of the absolute values of K_rec within each group is statistically analyzed in intervals with a width of 0.01, calculating the proportion of each interval. When the overall density proportion of the positive slope group exceeds that of the negative slope group, it is determined that the user group in the current silence warning interval is in a state dominated by shielding spread, and the concentrated shielding interval is determined by the duration of the interval corresponding to the range of K_rec values with the highest positive slope density. When the proportion of the negative slope group is higher, it is determined that the shielding is in a state dominated by natural recovery. Low-reliability markers in the shielding slope set participate in the density statistics of each group with a weight of 0.3 to avoid the unstable slope estimation of short duration intervals excessively affecting the distribution pattern. When the overall proportion of the positive slope group is significantly higher than that of the negative slope group, and the absolute value of the positive K_rec is concentrated in a large range, it indicates that the blocking frequency did not drop after multiple surge events within the statistical period. The continuous push notifications accumulated a systematic blocking inertia among users in the silent warning range. Users' defensive attitude towards push channels continued to deepen after each surge rather than naturally subsiding over time. When the negative slope group is dominant and the absolute value of K_rec is relatively large, the blocking frequency quickly dropped back to the baseline after each surge event. Users' acceptance of push channels could recover on its own after a brief period of resistance. The two types of statistical periods show completely opposite distribution patterns in the density proportion of the positive and negative groups in the blocking recovery distribution.
[0062] Based on the shielding recovery distribution, the concentrated shielding intervals are determined to generate operational decline indicators. Operational decline indicators mark the concentrated periods during which push messages are unable to effectively reach users in the silent warning interval. The core range is anchored to the continuous interval corresponding to the highest density slope interval in the shielding recovery distribution, extending five days to both sides to cover the leading and continuation segments of shielding spread, forming the decline coverage period of the operational decline indicator. When the shielding recovery distribution is dominated by a positive slope, the available push channel volume continuously narrows during the decline coverage period, and a large number of push messages cannot be delivered to users; the severity of the operational decline indicator is judged as high decline. When the shielding recovery distribution is dominated by a negative slope but the absolute value is relatively small, shielding recovery is slow, and the delivery rate is difficult to recover in the short term; this is judged as medium decline. When the absolute value of the negative slope is relatively large, shielding recovers quickly, and the delivery rate can return to normal levels within a few days; this is judged as low decline. The operational decline indicator marks the severity level of each decline coverage period and the number of users in the corresponding silent warning interval within that period. When the number of users in the silent warning interval is large within a certain decline coverage period and the masking recovery distribution is dominated by a positive slope, the operational decline indicator is classified as high decline, indicating that the push notification reach to a large number of high-value silent users has been systematically narrowed during that period. When the number of users in the decline coverage period is small and the negative slope is dominant, the severity is classified as low decline. When the number of users in the silent warning interval remains high and the masking frequency increases sharply during a high decline period, the actual reach loss during the decline coverage period is amplified by both the number of affected users and the depth of masking spread. High decline entries with a large number of users and a large absolute value of the masking recovery slope, and medium decline entries with a small number of users and slow masking recovery are distinguished by severity level in the operational decline indicator. The two types of entries correspond to different intensity gradient judgment results during the reach failure gradient matching stage.
[0063] User operation monitoring reports are generated by matching reach failure gradients based on operational decline indicators and abnormal churn patterns. The reach failure gradient measures the depth of failure of regular push notifications to target users and determines the intervention intensity accordingly. The severity level of each decline indicator's coverage period corresponds one-to-one with the matching type of each user in the abnormal churn pattern. The combination of high decline severity and users in the A-type abnormal churn pattern is determined as the highest reach failure gradient. These users face the dual challenges of a systemic narrowing of push channels and the risk of high churn due to their high value; regular push notifications are essentially ineffective, requiring the initiation of backup reach paths such as manual outbound calls or dedicated offline services. Combinations of high decline with B-type and medium decline with A-type are determined as the next highest reach failure gradients. Combinations of low decline or steady-state churn patterns are determined as low reach failure gradients. Higher reach failure gradient levels indicate more timely intervention and a more differentiated reach approach. The user set corresponding to each gradient level and the suggested reach strategy together form the gradient configuration entries. The highest gradient recommends one-on-one reach with a dedicated account manager, the next highest gradient recommends small-batch, precise content push notifications, and the low gradient maintains the regular push notification rhythm. All tier configuration items are arranged from highest to lowest tier level. Combined with the time coordinates of each decline coverage period and the user identifier list, a user operation monitoring report is generated. When the number of users with the highest tier in a certain statistical period is significantly higher than that of the previous period and the severity of the operation decline indicator increases in the same period, the strategy recommendation in the user operation monitoring report should be adjusted to a higher intensity personalized intervention configuration. When the number of users with the highest tier continues to decline in another statistical period and the severity of the operation decline indicator tends to decline overall, it indicates that the intervention measures in the previous period have produced a positive effect.
[0064] To implement the AI-based vehicle-to-everything (V2X) user tiered operation method corresponding to the above method embodiments, and to achieve the corresponding functional and technical effects, see [link to documentation]. Figure 2 , Figure 2 This application provides a structural block diagram of an AI-based vehicle-to-everything (V2X) user tiered operation system, which includes: Data acquisition module 201 is used to collect user behavior data and consumption record data, and to perform RFM index coupling analysis on the user behavior data and the consumption record data to generate a user baseline feature set; Feature verification module 202 is used to perform hierarchical offset positioning to extract hierarchical offset features from the user baseline feature set, perform cross-channel label cross-verification on the hierarchical offset features to generate a contradiction verification set, and extract hierarchical features by value dimension decomposition based on the contradiction verification set. The hierarchical clustering module 203 is used to determine the hierarchical affiliation label based on the hierarchical features, extract incentive dependency risk features from the user behavior data to generate an incentive dependency risk identifier, and perform incentive non-response clustering hierarchically based on the incentive dependency risk identifier and the hierarchical affiliation label to generate a user value map. The churn prediction module 204 is used to perform hierarchical exit quantification based on the user value map and the user baseline feature set to generate an exit amplitude value, extract the distribution of silence duration of high-value users from the user value map to generate a silence warning interval, and perform churn trigger node positioning on the exit amplitude value to generate a churn prediction coefficient. The report output module 205 is used to perform joint level matching between the churn prediction coefficient and the hierarchical affiliation label to generate an abnormal churn pattern, perform push channel blocking feature monitoring on the silent warning interval to generate an operational decline identifier, and perform reach failure gradient matching between the operational decline identifier and the abnormal churn pattern to generate a user operation monitoring report.
[0065] The aforementioned AI-based vehicle-to-everything (V2X) user tiered operation system can implement the AI-based V2X user tiered operation method described in the above method embodiments. The options in the above method embodiments are also applicable to this embodiment and will not be detailed here. The remaining content of this application's embodiments can be referred to the content of the above method embodiments, and will not be repeated in this embodiment.
[0066] The purpose of the above embodiments is to reproduce and derive the technical solution of the present invention by way of example, and to fully describe the technical solution, purpose and effect of the present invention. The purpose is to enable the public to have a more thorough and comprehensive understanding of the disclosure of the present invention, and not to limit the scope of protection of the present invention.
Claims
1. An AI-based method for segmented operation of connected vehicle users, characterized in that, include: Collect user behavior data and consumption record data, and perform RFM index coupling analysis on the user behavior data and consumption record data to generate a user baseline feature set; Hierarchical offset positioning is performed on the user baseline feature set to extract hierarchical offset features. Cross-channel label cross-verification is performed on the hierarchical offset features to generate a contradiction verification set. Based on the contradiction verification set, value dimension decomposition is performed to extract hierarchical features. Based on the hierarchical features, a hierarchical affiliation label is determined. Incentive dependency risk features are extracted from the user behavior data to generate an incentive dependency risk identifier. Based on the incentive dependency risk identifier and the hierarchical affiliation label, incentive non-response clustering is performed to generate a user value map. Based on the user value graph and the user baseline feature set, hierarchical exit quantification is performed to generate exit amplitude value. The distribution of silence duration of high-value users is extracted from the user value graph to generate silence warning interval. The exit amplitude value is used to perform churn trigger node positioning to generate churn prediction coefficient. The churn prediction coefficient and the hierarchical affiliation label are used to perform joint level matching to generate an abnormal churn pattern. The silent warning interval is monitored for push channel blocking characteristics to generate an operational decline identifier. Based on the operational decline identifier and the abnormal churn pattern, a reach failure gradient matching is performed to generate a user operation monitoring report.
2. The method according to claim 1, characterized in that, The step of performing cross-channel label cross-verification on the hierarchical offset features to generate a contradiction verification set includes: The hierarchical offset features are classified by channel affiliation to generate each channel affiliation subset; A cross-channel consumption response direction comparison is performed on each of the aforementioned channel subsets to generate a channel response difference table; Based on the channel response difference table, the RFM and tag reverse response intervals are located to generate a set of contradictory signal candidates; The contradiction degree critical calibration is performed on the contradictory signal candidate set to generate a contradiction verification set.
3. The method according to claim 1, characterized in that, The step of extracting hierarchical features by decomposing the value dimension based on the contradiction verification set includes: The contradictory verification set is mapped to user-level flow trajectories to generate a hierarchical backflow distribution map; The high-value layer backflow concentration area is located and a backflow direction feature set is generated by using the hierarchical backflow distribution map. A return flow rate gradient processing is performed on the return flow direction feature set to generate a return flow acceleration distribution; Based on the aforementioned reflux acceleration distribution, the accelerated reflux region is located to generate layered features.
4. The method according to claim 1, characterized in that, The step of generating a user value map by implementing incentive-free clustering and hierarchical generation based on the incentive dependency risk identifier and the hierarchical affiliation label includes: The hierarchical attribution labels are weighted to generate a weighted attribution label set; Initial clustering results are generated based on the weighted affiliation label set. From the initial clustering results, locate the repeat purchase deviation user samples and generate the deviation frequency distribution; The user value map is generated by hierarchically correcting the initial clustering results using the deviation frequency distribution and the incentive dependency risk identifier.
5. The method according to claim 1, characterized in that, The step of performing churn trigger node location and generating churn prediction coefficient on the exit amplitude value includes: The exit amplitude value is statistically analyzed at different levels to generate a time series sequence of net outflow. The net outflow time series is subjected to sudden increase moment detection to generate a sudden increase node set; Based on the rapidly increasing node set, a node distribution map is generated by locating the churn trigger interval. Based on the node distribution map, the churn risk level is determined and a churn prediction coefficient is generated.
6. The method according to claim 1, characterized in that, The step of extracting the distribution of inactivity duration of high-value users from the user value graph to generate inactivity warning intervals includes: The user value graph is then used to generate an activity level sequence by classifying user activity. A set of silent users is generated by filtering high-value silent users from the activity hierarchy sequence; Perform cumulative silence duration analysis on the silent user set to generate a silence cluster distribution; Based on the aforementioned silent clustering distribution, critical silent nodes are identified, and silent early warning intervals are generated.
7. The method according to claim 1, characterized in that, The step of generating an operational decline identifier by monitoring the channel blocking features of the silent warning interval includes: The silence warning interval is time-series sampled to generate a sequence of available reach quantities; Based on the available reach sequence, the active shielding frequency distribution is calculated to generate the shielding time sequence distribution; The shielding recovery distribution is generated by extracting the shielding duration slope after a sudden increase from the shielding temporal distribution. Based on the shield recovery distribution, the shield concentration interval is determined to generate an operational decline indicator.
8. The method according to claim 3, characterized in that, The step of locating high-value backflow concentration areas and generating a backflow direction feature set through the hierarchical backflow distribution map includes: The boundary crossing frequency distribution is generated by statistically analyzing the number of boundary crossings in adjacent time periods on the hierarchical backflow distribution map. Based on the boundary crossing frequency distribution, the peak crossing frequency of each time period is located to generate the crossing stability distribution; The peak drift trend is captured by the crossing stability distribution to generate a candidate set for crossing the boundary. Based on the candidate set of boundary crossings, the concentrated locations of the return flow direction are determined, and a return flow direction feature set is generated.
9. The method according to claim 7, characterized in that, The step of extracting the shielding duration slope after a sudden increase from the shielding time-series distribution to generate a shielding recovery distribution includes: The sudden increase event is located in the shielding time sequence distribution to generate a shielding sudden increase time sequence; The duration interval after the sudden increase in the shielding time sequence is extracted to generate a shielding duration subset; The shielding slope set is generated by performing slope fitting calculation on the shielding persistence subset; Based on the shielding slope set, distribution statistics are performed to generate the shielding recovery distribution.
10. An AI-based vehicle-to-everything (V2X) user tiered operation system, characterized in that: include: The data acquisition module is used to collect user behavior data and consumption record data, and to perform RFM index coupling analysis on the user behavior data and consumption record data to generate a user baseline feature set; The feature verification module is used to perform hierarchical offset positioning to extract hierarchical offset features from the user baseline feature set, perform cross-channel label cross-verification on the hierarchical offset features to generate a contradiction verification set, and decompose the value dimension to extract hierarchical features based on the contradiction verification set. The hierarchical clustering module is used to determine the hierarchical affiliation label based on the hierarchical features, extract incentive dependency risk features from the user behavior data to generate incentive dependency risk identifiers, and perform incentive non-response clustering hierarchically based on the incentive dependency risk identifiers and the hierarchical affiliation labels to generate a user value map. The churn prediction module is used to generate a churn magnitude value by performing hierarchical exit quantification based on the user value graph and the user baseline feature set, extract the silence duration distribution of high-value users from the user value graph to generate a silence warning interval, and perform churn trigger node positioning on the exit magnitude value to generate a churn prediction coefficient. The report output module is used to perform joint level matching between the churn prediction coefficient and the hierarchical affiliation label to generate an abnormal churn pattern, perform push channel blocking feature monitoring on the silent warning interval to generate an operational decline identifier, and perform reach failure gradient matching between the operational decline identifier and the abnormal churn pattern to generate a user operation monitoring report.