Attribution analysis method and device for cloud disk APP user active causes

By transforming the user activity attribution analysis of cloud storage apps into a supervised learning problem, and combining data cleansing, time series modeling, and fusion prediction, the shortcomings of existing methods in user activity motivation analysis are addressed, achieving high-precision and interpretable attribution analysis, and supporting precise operation and product optimization.

CN121744209APending Publication Date: 2026-03-27CHINA MOBILE INTERNET CO LTD +1
View PDF 0 Cites 1 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-22
Publication Date
2026-03-27

AI Technical Summary

Technical Problem

Existing attribution analysis methods for user activity in cloud storage apps fail to fully consider the temporal dynamics of user behavior, the interactive effects between features, and the differences among different user groups. This results in one-sided attribution results with insufficient interpretability, making it difficult to accurately capture the true motivations behind user activity.

Method used

The user activity attribution analysis is transformed into a supervised learning problem based on short-term user behavior sequences. Through a progressive technology chain of data cleansing, time series modeling, fusion prediction, and interpretable attribution, including data noise reduction and anomaly cleaning, training the fusion prediction model, and calculating the contribution of behavioral features to the prediction results, the attribution factors are determined.

Benefits of technology

It enables in-depth, accurate, and interpretable analysis of the motivations for user activity in cloud storage apps, improving the accuracy and interpretability of attribution analysis and providing reliable data-driven decision-making support for precise operations and product optimization.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121744209A_ABST
    Figure CN121744209A_ABST
Patent Text Reader

Abstract

The invention relates to an attribution analysis method and device for active causes of a cloud disk APP user. The method comprises the following steps: acquiring behavior characteristic data of the cloud disk APP user in a current statistical period; preprocessing the behavior characteristic data, wherein preprocessing is used for noise detection and filtering and abnormal user detection; based on the preprocessed behavior characteristic data, training a fusion prediction model used for predicting whether the user is retained in the next statistical period; calculating the contribution degree of each behavior feature to a prediction result based on the prediction result of the fusion prediction model; and determining an attribution factor which causes the user to be active in the current statistical period based on the contribution degree. According to the scheme, through training the fusion prediction model, time sequence dependence and complex feature interaction in a behavior sequence are accurately captured, so that high-precision prediction of user retention is realized; and the contribution degree of each behavior feature to a prediction result is calculated, so that the attribution factors causing the user activity are reversely deduced and determined, and the real motivation of the user activity can be accurately captured.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of artificial intelligence technology, and in particular to an attribution analysis method and apparatus for analyzing the motivations of cloud storage app users' activity. Background Technology

[0002] In recent years, with the widespread adoption of cloud storage services, cloud storage apps have accumulated massive amounts of user behavior data. Accurately identifying the core drivers of user activity from this data has become crucial for businesses to optimize product experience and implement precise operations. Effective activity attribution analysis can help platforms understand the motivations behind user behavior, providing data support for feature optimization, user retention, and marketing strategy development.

[0003] Currently, most common user activity attribution methods are based on statistical summaries or linear models. They determine the motivations for activity by weighting features or using scorecards to model user behavior data such as app usage. However, these methods often fail to fully consider the temporal dynamics of user behavior, the interactive effects between features, and the differences between different user groups. This results in one-sided attribution results with insufficient interpretability, making it difficult to accurately capture the true motivations behind user activity. Summary of the Invention

[0004] This application provides an attribution analysis method and apparatus for the motivations of cloud storage app users.

[0005] According to a first aspect of the embodiments of this application, an attribution analysis method for the motivations of user activity in a cloud storage app is provided, the method comprising: Obtain behavioral characteristic data of cloud storage app users within the current statistical period; The behavioral feature data is preprocessed for noise detection and filtering, as well as abnormal user detection. Based on the preprocessed behavioral feature data, a fusion prediction model is trained to predict whether users will remain in the next statistical period. Based on the prediction results of the fusion prediction model, the contribution of each behavioral feature to the prediction results is calculated. Based on the contribution level, the attribution factors that led to the user's activity during the current statistical period are determined.

[0006] According to a second aspect of the embodiments of this application, an attribution analysis device for the activity motivation of cloud storage APP users is provided, the device comprising: The feature data acquisition module is used to acquire behavioral feature data of cloud disk APP users within the current statistical period; A preprocessing module is used to preprocess the behavioral feature data, and the preprocessing is used for noise detection and filtering, as well as abnormal user detection. The model training module is used to train a fusion prediction model based on preprocessed behavioral feature data to predict whether a user will remain in the next statistical period. The contribution calculation module is used to calculate the contribution of each behavioral feature to the prediction result based on the prediction result of the fusion prediction model through an interpretive model. The attribution factor determination module is used to determine, based on the contribution level, the attribution factors that led to the user's activity during the current statistical period.

[0007] According to a third aspect of the embodiments of this application, an electronic device is provided. The electronic device includes a memory and a processor, wherein the memory stores a computer program, and the processor executes the program to implement the method described above.

[0008] According to a fourth aspect of the embodiments of this application, a computer-readable storage medium is provided having a computer program stored thereon, which, when executed by a processor, implements the methods described above in this application.

[0009] According to a fifth aspect of the embodiments of this application, a computer program product is provided, including a computer program that, when executed by a processor, implements the methods described above in this application.

[0010] The attribution analysis method and apparatus for user activity motivation of cloud storage apps provided in this application transform the attribution problem of user activity motivation analysis into a supervised learning problem of predicting user retention based on short-term user behavior sequences. This is achieved through a progressive technical chain of "data purification - temporal modeling - fusion prediction - interpretable attribution." Specifically, it acquires multi-dimensional user behavior feature data and performs data denoising and anomaly cleaning to improve data quality. It trains a fusion prediction model to accurately capture temporal dependencies and complex feature interactions in the behavior sequence, thereby achieving high-precision prediction of user retention. Furthermore, it calculates the contribution of each behavioral feature to the prediction result, thereby inferring and determining the attribution factors leading to user activity and accurately capturing the true motivations for user activity. This improves the accuracy, interpretability, and business relevance of attribution analysis, providing a reliable data-driven decision-making basis for precise operation and product optimization. Attached Figure Description

[0011] Further details, features, and advantages of this application are disclosed in the following description of exemplary embodiments in conjunction with the accompanying drawings, in which: Figure 1 A flowchart illustrating an exemplary embodiment of this application of an attribution analysis method for user activity motivation in a cloud storage app; Figure 2A flowchart of an attribution analysis method for user activity drivers in a cloud storage app, provided as another exemplary embodiment of this application; Figure 3 A schematic block diagram of the functional modules of an attribution analysis device for the motivation of user activity in a cloud storage APP provided as an exemplary embodiment of this application; Figure 4 A structural block diagram of an electronic device provided in an exemplary embodiment of this application; Figure 5 A structural block diagram of a computer system provided for an exemplary embodiment of this application. Detailed Implementation

[0012] Embodiments of this application will now be described in more detail with reference to the accompanying drawings. While some embodiments of this application are shown in the drawings, it should be understood that this application can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of this application. It should be understood that the drawings and embodiments of this application are for illustrative purposes only and are not intended to limit the scope of protection of this application.

[0013] It should be understood that the steps described in the method embodiments of this application may be performed in different orders and / or in parallel. Furthermore, the method embodiments may include additional steps and / or omit the steps shown. The scope of this application is not limited in this respect.

[0014] The term "comprising" and its variations as used herein are open-ended, meaning "including but not limited to". The term "based on" means "at least partially based on". The term "one embodiment" means "at least one embodiment"; the term "another embodiment" means "at least one additional embodiment"; the term "some embodiments" means "at least some embodiments". Definitions of other terms will be given in the following description. It should be noted that the concepts of "first", "second", etc., mentioned in this application are used only to distinguish different devices, modules, or units, and are not intended to limit the order of functions performed by these devices, modules, or units or their interdependencies.

[0015] It should be noted that the terms "a" and "a plurality of" used in this application are illustrative rather than restrictive, and those skilled in the art should understand that, unless otherwise expressly indicated in the context, they should be understood as "one or more". The names of the messages or information exchanged between multiple devices in the embodiments of this application are for illustrative purposes only and are not intended to limit the scope of these messages or information.

[0016] It is understood that before using the technical solutions disclosed in the various embodiments of this application, users should be informed of the types, scope of use, and usage scenarios of the personal information involved in this application in an appropriate manner in accordance with relevant laws and regulations, and user authorization should be obtained.

[0017] For example, upon receiving a user's active request, a prompt message is sent to the user to explicitly inform them that the requested operation will require the acquisition and use of the user's personal information. This allows the user to independently choose whether to provide personal information to the software or hardware, such as the electronic device, application, server, or storage medium performing the operations of this application's technical solution, based on the prompt message.

[0018] As an optional but non-limiting implementation, in response to a user's active request, sending a prompt message to the user can be done via a pop-up window, where the prompt message can be presented in text format. Furthermore, the pop-up window can also include a selection control allowing the user to choose "agree" or "disagree" to provide personal information to the electronic device. It is understood that the above notification and user authorization process is merely illustrative and does not limit the implementation of this application; other methods that comply with relevant laws and regulations may also be applied to the implementation of this application.

[0019] This application first provides an attribution analysis method for the motivations of user activity in cloud storage apps. It transforms the attribution question of "why a user is active in the current period" into a supervised learning problem that predicts user retention in the next period based on short-term behavior patterns. Finally, it uses model interpretability technology to infer the core activity motivations. This method achieves more accurate and insightful user behavior attribution by introducing innovative data preprocessing, model fusion, and interpretability analysis. Figure 1 As shown, the method may include the following steps: In step S110, behavioral characteristic data of cloud disk APP users within the current statistical period are obtained.

[0020] In this step, multi-dimensional behavioral data generated by the target user group within a specified statistical period (e.g., the last 30 days) can be extracted from the cloud storage app's backend database or data warehouse. This data constitutes the basic feature set for subsequent analysis, typically including but not limited to: user function usage characteristics (such as the frequency of operations like uploading, downloading, backing up, sharing, and viewing), consumption behavior characteristics (such as package type and spending amount), activity participation information, network performance indicators, and terminal device information. These features aim to comprehensively depict users' product usage habits and status.

[0021] In step S120, the behavioral feature data is preprocessed for noise detection and filtering, as well as abnormal user detection.

[0022] This step aims to improve data quality and lay a solid foundation for model training. Preprocessing mainly includes the following two parts: Noise Detection and Filtering: Given the time-series nature of user behavior data, this implementation introduces frequency domain analysis techniques for noise reduction. Specifically, the user behavior time-series data is converted from the time domain to the frequency domain using a discrete Fourier transform. High-frequency noise components caused by abnormal operations or system jitter are identified and filtered out. Then, an inverse Fourier transform is used to restore the clean time-domain data, thus preserving the true behavior patterns.

[0023] Abnormal user detection: By constructing feature vectors that integrate multi-dimensional features such as device, time, and user attributes, and performing detection based on preset rules (such as "one device with multiple numbers" and abnormally high login frequency), non-real user or malicious behavior data can be identified and excluded to prevent them from interfering with model training.

[0024] In addition, preprocessing may optionally include steps such as data cleaning (e.g., missing value imputation, standardization) and feature selection (e.g., based on IV values).

[0025] In step S130, based on the preprocessed behavioral feature data, a fusion prediction model is trained to predict whether a user will remain in the next statistical period.

[0026] This step utilizes preprocessed, high-quality data to train a predictive model that combines the advantages of deep learning and ensemble learning. The model aims to learn the complex mapping between user behavior patterns in the current cycle and whether they remain active in the next cycle (i.e., "retention").

[0027] The model architecture in this embodiment can be a fusion model combining Long Short-Term Memory (LSTM) and XGBoost. LSTM excels at capturing long-term and short-term temporal dependencies in user behavior sequences, while XGBoost is adept at handling high-dimensional static features and capturing complex nonlinear interactions.

[0028] Training optimization: During the training process, optimization strategies tailored to business needs can be introduced. For example, different weights can be assigned to training samples based on the user's lifecycle stage (introduction, growth, maturity, etc.) to make the model focus more on high-value users; or the penalty for "false negatives" (i.e., real retained users being misclassified as churn) can be increased in the loss function to align with the core business concern of reducing user churn.

[0029] In step S140, based on the prediction results of the fusion prediction model, the contribution of each behavioral feature to the prediction results is calculated.

[0030] After obtaining a fusion model with good predictive performance, the influence of each input feature on the model to make a specific prediction (such as predicting that a user will be retained) is quantified.

[0031] The implementation example can calculate the contribution of each behavioral feature to the prediction result using an interpretation method. This interpretation method can employ Explainable Artificial Intelligence (XAI) methods based on Shapley values ​​from cooperative game theory, such as SHAP. By calculating the Shapley value for each feature, it is used as the contribution of that feature to the model's final prediction result. This value not only has a clear mathematical meaning (the marginal contribution of the feature in an average sense) but also indicates the direction of the influence (positively promoting retention or negatively leading to churn).

[0032] In step S150, attribution factors that led to a user's activity in the current statistical period are determined based on contribution.

[0033] The example demonstrates how the calculated feature contribution is analyzed and interpreted to transform the model output into actionable business insights.

[0034] The features can be sorted according to their contribution to identify key features that play a decisive role in predicting user retention. Furthermore, this embodiment can also calculate the joint interaction contribution value between features to identify feature combinations that work together and synergistically affect user retention (e.g., the combined effect of "frequent use of sharing functions" and "participation in specific operational activities").

[0035] These identified key features or combinations of features are determined to be the core attributable factors that cause users to remain active during the current statistical period. For example, the analysis results may show that the main driver of a user group's activity is the combined effect of "frequent use of the phone backup function" and "holding a large-capacity storage plan".

[0036] This embodiment achieves in-depth, accurate, and interpretable analysis of the motivations for user activity in cloud storage apps through a full-chain innovation from data cleaning and model building to result interpretation, providing strong data-driven decision support for subsequent precise operations, product optimization, and user experience improvement.

[0037] The attribution analysis method for user activity drivers in cloud storage apps provided in this application transforms the attribution problem of user activity driver analysis into a supervised learning problem of predicting user retention based on short-term user behavior sequences. This is achieved through a progressive technical chain of "data cleansing - temporal modeling - fusion prediction - interpretable attribution." Specifically, it acquires multi-dimensional user behavior feature data and performs data denoising and anomaly cleaning to improve data quality. It trains a fusion prediction model to accurately capture temporal dependencies and complex feature interactions in the behavior sequence, thereby achieving high-precision prediction of user retention. Furthermore, it calculates the contribution of each behavioral feature to the prediction result, thereby inferring and determining the attribution factors leading to user activity. This improves the accuracy, interpretability, and business relevance of attribution analysis, providing a reliable data-driven decision-making basis for precise operation and product optimization.

[0038] In this embodiment, the preprocessing of the above-mentioned behavioral feature data also includes introducing a dynamic decay function, which assigns a weight to each behavioral feature that decays over time based on the time of occurrence of the behavior and the importance of the feature.

[0039] In user behavior analysis, behaviors occurring at different times have varying degrees of impact on the current state, with recent behaviors often being more valuable than those from further back. To address the issue of insufficient consideration of the time-dependent decay of behavioral data in existing technologies, this embodiment innovatively introduces a dynamic decay function during the data preprocessing stage.

[0040] The dynamic decay function is used to dynamically calculate a decay weight for each behavioral feature based on the time of its occurrence and its own importance, so that the model can more sensitively capture recent changes in behavior and key features.

[0041] The implementation example achieves dual regulation by introducing this dynamic decay function: first, it emphasizes recent behavior based on the time factor; second, it highlights key behaviors based on feature importance. This enables the preprocessed feature data to more accurately reflect the core, time-sensitive behavioral patterns driving the user's current and recent active state, laying a more scientific data foundation for building a high-precision prediction model and effectively improving the accuracy and timeliness of the final attribution results.

[0042] In this embodiment, the preprocessing of the behavioral feature data also includes user grouping, which divides users into different groups based on user features, and performs subsequent model training and attribution analysis for each user group.

[0043] Among the massive user base of cloud storage apps, different user groups often exhibit drastically different behavioral patterns and motivations due to differences in their usage purposes, habits, and attributes. If a single, all-user model is used for analysis, it is prone to introducing noise due to significant differences in the distribution of characteristics between groups, making it difficult for the model to capture the deep-seated patterns of specific groups and resulting in "averaged" attribution results lacking specificity. To address this issue, this embodiment introduces a user segmentation process in the preprocessing stage.

[0044] Specifically, this step first utilizes unsupervised clustering algorithms (such as K-means, hierarchical clustering, etc.) to perform cluster analysis on all users to be analyzed based on preprocessed user characteristics (e.g., feature usage preferences, consumption levels, terminal device types, active time patterns, etc.), automatically grouping users with similar characteristics and behavioral patterns into the same group. For example, it may identify typical groups such as "high-frequency storage backup users," "social sharing active users," "light tool users," and "users at potential churn risk."

[0045] After user segmentation, this embodiment does not mix all data for unified modeling. Instead, it independently executes the subsequent model training and attribution analysis process for each segmented user subgroup. That is, it builds a dedicated fusion prediction model for each group and trains the model parameters based on the data of that group; then, it uses the trained group-specific model to perform prediction and attribution analysis on users within that group.

[0046] By implementing user segmentation, this embodiment achieves significant technical benefits: First, it improves the accuracy of model prediction and attribution because each sub-model only needs to learn the behavioral patterns of a relatively homogeneous user group, avoiding interference from differences in characteristics between different groups. Second, it enhances the business interpretability and relevance of attribution results, allowing operators to clearly identify the core activity drivers of different user groups (such as "sharing users") and formulate highly differentiated operational strategies (such as designing exclusive sharing incentive activities for this group). Third, it optimizes the computational efficiency and generalization ability of the model. The model structure trained on segmented groups can be more streamlined, and because its learning objectives are more specific, its generalization performance within its respective group is better.

[0047] Therefore, user segmentation is a key step in achieving precise and refined user behavior analysis in this solution. It ensures that the final attribution conclusions can truly reflect the core needs and driving factors of different user segments, laying a solid data analysis foundation for personalized product operation.

[0048] Based on the above embodiments, in another embodiment provided in this application, the noise detection and filtering includes: (1) Convert user behavior feature data from the time domain to the frequency domain.

[0049] Because user actions within the product (such as logging in, uploading, sharing, etc.) are recorded chronologically, they naturally form time-series data. To identify the "noise" generated by accidental misoperations, brief system malfunctions, or unintended abnormal interactions, this embodiment innovatively introduces frequency domain analysis techniques from the field of signal processing into user behavior data analysis. Specifically, for a single user behavior time-series data... x ( n ),in n =0,1,…,N 1 represents a time point, N is the sequence length, and the sequence is transformed from the time domain to the frequency domain using the Discrete Fourier Transform.

[0050] The frequency domain components obtained after transformation reveal the amplitude and phase information of different frequency components in the original time series. Normal and stable user behavior patterns usually correspond to lower frequency components, while sudden and sporadic noise often manifests as abnormal high-frequency components.

[0051] (2) Identify and filter high-frequency components that are above a preset threshold in the frequency domain.

[0052] After conversion to the frequency domain, by analyzing the amplitude spectrum of each frequency component, high-frequency components with amplitudes significantly higher than the background can be clearly identified. These components correspond to sharp noise in the time-domain data. This embodiment dynamically sets a filtering threshold, which can be determined based on the statistical distribution of frequency domain amplitudes (e.g., selecting a high percentile as the threshold). All high-frequency components with amplitudes higher than this threshold will be identified as noise and filtered out (e.g., their amplitudes will be zeroed or significantly attenuated). This embodiment directly removes noise components from a frequency domain perspective, retaining the low-frequency and main frequency components that represent the user's actual behavior patterns.

[0053] The preset threshold is dynamically determined based on the local fraction of the frequency domain component amplitude.

[0054] In noise detection using frequency domain analysis, setting a universally applicable and accurate threshold to distinguish between signal and noise is crucial. To address the issue that fixed thresholds may not adapt to the varying energy distribution of behavioral data from different users or time periods, this embodiment optimizes the method for determining the "preset threshold" by employing a dynamic determination strategy. Specifically, the threshold is set by analyzing the local distribution of the amplitude sequences of each frequency component obtained after frequency domain transformation. For example, a sliding window approach can be used to calculate specific local quantiles (such as the 95th or 98th quantile) of the frequency domain amplitude within the window, serving as the dynamic threshold for that window's data. Frequency components with amplitudes higher than these quantiles are considered abnormal high-frequency noise and filtered out. This method adapts to the fluctuation range of the data itself, ensuring effective noise identification during both stable and fluctuating periods. It improves the adaptability and accuracy of noise filtering, avoids signal distortion or noise residue caused by improper threshold settings, and provides a higher-quality data foundation for subsequent modeling.

[0055] (3) Convert the filtered frequency domain data back to the time domain to obtain the denoised behavioral feature data.

[0056] After filtering out high-frequency noise components, the processed frequency domain data is transformed back into the time domain using an inverse discrete Fourier transform, resulting in purified user behavior time series data. Compared to the original data, this method minimizes the interference of random and abnormal fluctuations, providing a smoother and more stable reflection of users' true behavioral habits and trends.

[0057] In the embodiment, when converting user behavior feature data from the time domain to the frequency domain, a sliding window technique can be used to segment the time series data and dynamically adjust the statistical thresholds used to detect outliers within each window.

[0058] In this embodiment, to accommodate streaming data arrivals or long sequence processing, a sliding window technique can be incorporated. The long-duration behavioral sequence is divided into multiple overlapping or continuous sliding windows, and the aforementioned time-frequency conversion, noise filtering, and inverse conversion operations are performed independently on the data within each window. Simultaneously, the statistical threshold for noise determination can be dynamically adjusted based on the local statistical characteristics (such as mean and variance) of the data within the window, thereby achieving more refined and adaptive noise removal for non-stationary behavioral sequences.

[0059] By implementing this embodiment, frequency domain analysis was successfully applied to the non-traditional signal processing field of user behavior data cleaning. This effectively distinguishes and filters out high-frequency noise hidden in time-domain data, significantly improving the signal-to-noise ratio and quality of behavioral feature data subsequently used for model training. This lays a crucial data foundation for building a more robust and accurate prediction model, fundamentally improving the reliability of the final user activity attribution results.

[0060] Based on the above embodiments, in another embodiment provided in this application, when detecting abnormal users, step S120 may further include the following steps: In step S121, a high-dimensional feature vector is constructed that integrates device features, time series features, user attribute features, and usage scenario features.

[0061] Since abnormal or non-genuine user behavior patterns often manifest in multiple dimensions of abnormal features and their combinations, single-dimensional detection is prone to missed detections. Therefore, this embodiment designs a multi-dimensional feature fusion detection mechanism. Specifically, it extracts and constructs a four-dimensional sub-feature set from the original behavioral feature data: Equipment characteristics (f(x)) device Examples include the device's unique identifier (ID), model, brand, and operating system.

[0062] Time series features (f(x)) time Examples include login time distribution, operation frequency timing patterns, and session duration sequences.

[0063] User attribute features (f(x)) user Examples include account registration duration, historical spending levels, and user tags.

[0064] Use case features (f(x)) usage Examples include commonly used functional modules, access network environment (Wi-Fi / mobile network), and geolocation mode.

[0065] These sub-features are concatenated to form a unified high-dimensional feature vector: FeatureVector=[f(x) device ,f(x) time ,f(x) user ,f(x) usage This vector comprehensively represents the integrated patterns of user behavior across devices, times, attributes, and scenarios, providing a rich information foundation for accurate anomaly identification.

[0066] In step S122, the high-dimensional feature vector is detected based on preset rules, which include determining that the number of accounts logged in by the same device within a preset time period exceeds a first threshold, or the user login frequency exceeds a second threshold.

[0067] Based on the obtained fused feature vector, this embodiment performs anomaly detection using pre-defined rules based on business common sense and data statistics. These rules directly apply to specific dimensions of the feature vector, for example: Multiple Accounts on One Device Detection Rule: Count the number of different accounts logged in by a specific device (such as a mobile phone) within a preset time period (e.g., the last 30 days). If this number exceeds a set first threshold (e.g., 4), the device is deemed to have abnormal association behavior, which may include account sharing, machine registration, or malicious attacks.

[0068] Abnormal login frequency detection rules: Calculate the login frequency of users. If a user's login frequency is significantly higher than the average level of the normal user group (i.e., exceeds the second threshold determined based on statistical distribution), it is marked as abnormal behavior, which may correspond to crawler scripts, traffic fraud, or account theft.

[0069] The rule base in this embodiment can be continuously expanded based on actual operational experience and data distribution, for example, by adding rules such as "login from multiple locations within a short period of time" and "extremely abnormal function usage sequences". By applying these rules to scan and judge high-dimensional feature vectors, abnormal user samples can be screened out efficiently and accurately.

[0070] By implementing this embodiment, an upgrade in anomaly detection has been achieved, moving from single-dimensional judgment to multi-dimensional collaborative perception. The fusion of feature vectors provides a more comprehensive characterization of user behavior, while the detection mechanism based on clear rules offers interpretable and controllable anomaly filtering methods. This effectively ensures the purity of the training data input to subsequent prediction models, eliminating interference from non-real user behavior on the model's learning of real user activity patterns, thereby improving the robustness and reliability of the final activity attribution model.

[0071] In the embodiments, the above-mentioned fusion prediction model can be a fusion model of deep learning model and machine ensemble learning model.

[0072] This deep learning model can be a long short-term memory network model. When training the long short-term memory network model, different importance weights can be assigned to the training samples according to the user's life cycle stage.

[0073] To accurately predict user retention in future cycles and thus infer their current activity drivers, this embodiment constructs a fusion prediction model. This model does not employ a single algorithm but creatively combines the advantages of deep learning models with machine learning ensemble models to simultaneously capture the complex temporal dependencies and nonlinear interactions between high-dimensional static features in user behavior patterns.

[0074] Specifically, the preferred deep learning model is LSTM. The core reason for choosing LSTM is that user behavior data (such as login and operation records over multiple consecutive days) is essentially a time series, and its current state is often influenced by both long-term behavioral patterns and short-term fluctuations. LSTM, through its unique forget gate, input gate, output gate, and cell state mechanism, can effectively learn and remember long-term dependencies in time series and flexibly update the remembered content, making it very suitable for modeling the dynamic evolution of user behavior.

[0075] In training this LSTM model, this embodiment performed a key optimization: assigning different importance weights to training samples based on the user's lifecycle stage. w t The user lifecycle is typically divided into stages such as introduction, growth, maturity, dormancy, and churn. The stability of user behavior, value contribution, and predictive significance vary at each stage. For example: Mature users exhibit stable and high-value behavior, forming the core of product revenue, and are therefore given high weight (e.g., ...). ω =1.5), enabling the model to focus on learning its stable retention behavior patterns.

[0076] Users in the growth stage are at a critical stage of value conversion and are therefore given higher weight (e.g.) ω =1.2), to accurately capture the characteristics of its transition to maturity.

[0077] Although the user data in the introductory phase is limited, it has great potential and is given a baseline weight (e.g.) ω =1.0).

[0078] The behavioral patterns of users in the dormant and churn periods may contain churn warning signals, but because their current activity level is low, they are given lower weight (e.g., ...). ω =0.8 or 0.5), moderately reduce its influence during training to avoid the model overfitting to inactive modes.

[0079] This optimization is achieved by modifying the loss function, for example, by introducing sample weights into the binary cross-entropy loss function:

[0080] in, w t This corresponds to the weights of different stages in the user lifecycle. This makes model training no longer "one-size-fits-all," but rather focuses on learning the retention patterns of high-value user groups, making the final predictions and attributions more aligned with the core business objective of maintaining and improving the activity and retention of core users.

[0081] By adopting a fusion architecture of LSTM and ensemble learning, and implementing lifecycle-based sample weighted training, the fusion prediction model constructed in this embodiment not only has the ability to process complex time-series data, but also its learning process is deeply aligned with business logic, thus providing a strong model foundation for achieving high-precision user retention prediction and subsequent accurate attribution.

[0082] In this embodiment, when training the long short-term memory network model, a penalty factor is introduced to impose an additional loss penalty on false negative prediction errors. This penalty factor is determined based on the ratio of positive to negative samples in the training data.

[0083] In the specific business scenario of user retention prediction, the cost of prediction errors is not asymmetrical. The harm of "false negative" errors (i.e., users will actually stay, but the model incorrectly predicts they will churn) is particularly significant. This is because it means the model fails to identify users with lasting value to the product, potentially leading to ineffective operational measures, missed opportunities for user retention, and substantial business losses. Conversely, the cost of "false positive" errors (predicting retention but actually churning) is relatively controllable. To address this business concern and deeply align model training with business objectives, this embodiment introduces a penalty factor for false negative errors during the training phase of the LSTM model. α ).

[0084] Specifically, this penalty factor α The value is not set arbitrarily, but is dynamically calculated and determined based on the ratio of positive samples (retained users) to negative samples (churned users) in the training dataset.

[0085] The principle behind this design is that when the number of positive samples is far less than the number of negative samples (i.e., the proportion of retained users is small, making them a more valuable minority class that needs to be correctly identified), the calculated... α The value will be less than 1. This factor will be introduced into the loss function specifically to amplify the loss caused by false negative errors.

[0086] Based on the weighted binary cross-entropy loss function, this embodiment constructs a loss function that includes a penalty factor. Lpenalty Its formal expression is as follows:

[0087] in, y t The labels are real (1 indicates retention, 0 indicates churn). This represents the retention probability predicted by the model. This corresponds to a false negative error: the actual situation is that the data is retained ( y t =1), but the model predicts no retention (probability). Low, therefore (High). This loss is multiplied by a penalty factor. α .item The weight for false positive errors remains at 1.

[0088] By introducing a penalty factor based on sample proportion α This embodiment implements loss function recalibration. During training, the model significantly experiences a heightened "pain" from misclassifying positive samples as negative samples (false negatives), driving it to optimize parameters and more cautiously and proactively learn feature patterns that correctly identify retained users. This effectively reduces the incidence of critical business errors, enabling the trained LSTM model to maintain high overall accuracy while specifically improving its sensitivity and accuracy in identifying high-value retained users. This provides a more reliable and truthful model foundation for subsequent activity attribution based on prediction results.

[0089] To fully leverage the advantages of different model types and build a more powerful prediction system, this embodiment provides a first specific model fusion strategy—feature-level fusion. The core idea of ​​this strategy is to utilize the powerful feature extraction capabilities of deep learning models to enhance the quality of input information in traditional ensemble learning models. Therefore, the fusion method of the aforementioned deep learning model and machine ensemble learning model is feature-level fusion, and its specific implementation includes the following steps: (1) Use deep learning models to extract temporal features of user behavior sequences.

[0090] First, preprocessed user behavior sequence data (e.g., frequency of function use, login interval, etc. arranged in chronological order) is input into the trained LSTM model. The LSTM model processes the sequence data step by step through its internal gated recurrent units and outputs a comprehensive, high-dimensional hidden state vector at the final time step. ht This vector ht It encapsulates the dynamic evolution pattern, long-term dependencies, and short-term fluctuation characteristics of user behavior throughout the entire observation window. It is a more representative abstraction of time-series features obtained by deep nonlinear transformation of the original sequence data.

[0091] (2) The temporal features are combined with the user's static features to form a fused feature.

[0092] Simultaneously, static user characteristics are acquired; these characteristics do not change or change slowly over time, such as user demographics (age, region), plan type, terminal device brand, and historical spending level. Then, the temporal feature vector output by the LSTM model is... htThe vector is concatenated with all relevant static feature vectors to form a new, higher-dimensional fused feature vector. This step achieves an organic combination of dynamic temporal patterns and static attribute states at the information level.

[0093] (3) Input the fused features into the machine ensemble learning model for training and prediction.

[0094] Finally, this feature vector, incorporating deep temporal information, is used as input data and fed into a powerful machine learning ensemble model (e.g., XGBoost, LightGBM, or Random Forest) for training or prediction. Ensemble learning models excel at handling high-dimensional, structured features and can effectively capture complex nonlinear interactions between features and handle feature importance filtering. At this point, the input received by the model is no longer the original, relatively independent time-series and static data, but rather a fused feature containing deep temporal patterns pre-mined by LSTM, enabling it to make more accurate judgments.

[0095] This embodiment creatively combines the advantages of LSTM in sequence modeling and feature abstraction with the advantages of ensemble learning models in fitting complex relationships and making efficient predictions through this feature-level fusion approach. This allows the final prediction model to understand both the "storyline" (temporal evolution) of user behavior and its "contextual setting" (static attributes), thus achieving a more comprehensive and accurate prediction of user retention. This architecture lays the foundation for inferring high-quality activity attribution from high-quality predictions.

[0096] To provide an alternative model integration solution, this embodiment offers a second specific model fusion strategy: model-level fusion (also known as stacking). This strategy allows different types of models to learn and predict independently, and then integrates them at the decision-making level to combine the "wisdom" of multiple models. The specific implementation includes the following steps: (1) Deep learning sub-model and machine ensemble learning sub-model are trained respectively.

[0097] In this approach, the two models are trained independently, with no data flow between them during the training phase. Specifically: Deep learning sub-models (such as LSTM networks) are trained independently using user behavior sequence data as input to learn how to predict user retention.

[0098] Machine ensemble learning sub-models (such as the XGBoost model) are trained independently using the user’s feature data (which may include static features and statistical features extracted from the sequence) as input to learn the same prediction task.

[0099] The two sub-models use the same or their own optimized training and validation sets, and converge to their respective best states, forming two "experts" with different perspectives and areas of expertise.

[0100] (2) Obtain the first prediction result of the deep learning sub-model for the target user and the second prediction result of the machine ensemble learning sub-model for the target user.

[0101] When a prediction needs to be made for a target user, the corresponding data is input into the two pre-trained sub-models mentioned above: By inputting their behavioral sequence data into a deep learning sub-model, we obtain the model's first prediction result regarding user retention. This result implies a deep understanding of the time series patterns by the model.

[0102] The feature data is input into the machine ensemble learning sub-model to obtain the model's second prediction result regarding user retention. This result implies the model's comprehensive judgment on the complex relationships between features.

[0103] (3) The first prediction result and the second prediction result are weighted and summed to obtain the final prediction result.

[0104] Finally, the independent predictions from the two sub-models are fused using a linear weighting method to form the final integrated prediction result. The formula is expressed as:

[0105] in, w 1 and w 2 represents the weight coefficients of the two models (satisfying...) w 1+ w (2=1). These weights can be dynamically determined by optimizing performance metrics (such as AUC, F1 score) on the validation set to ensure that models that perform more reliably under a particular data distribution are given higher decision weights.

[0106] This embodiment constructs a robust "committee decision-making" system through this model-level fusion approach. It allows diverse models to make independent judgments based on their respective data perspectives and algorithmic principles, which are then weighted and aggregated, effectively reducing the risk of overfitting or systematic bias that may exist with a single model. This fusion strategy typically achieves more stable and accurate final prediction performance than any single sub-model, providing a more reliable and robust predictive foundation for subsequent attribution analysis and is one of the key design features for improving the generalization ability of the entire method.

[0107] In this embodiment, the calculation of the contribution of each behavioral feature to the prediction result may specifically include: An additive feature attribution method based on Shapley values ​​is adopted to calculate the Shapley value of each input feature to the predicted value of the fusion prediction model output, and this value is used as the contribution of each behavioral feature to the prediction result.

[0108] After obtaining high-precision user retention predictions through a fusion prediction model, understanding "why the model made this prediction" and "which user behavior features dominated the prediction" is crucial for transforming the model output into actionable attribution conclusions. To this end, this embodiment introduces the Shapley Value theory based on cooperative game theory and concretizes it as an additive feature attribution method. This method quantifies the marginal contribution of each input feature to the model's final prediction, thereby obtaining a clear, fair, and theoretically sound measure of feature contribution.

[0109] Specifically, for a pre-trained fusion prediction model f (Regardless of whether its internal structure is feature-level fusion or model-level fusion), and a specific target user sample and its corresponding set of all features. N This embodiment uses the model's predicted values. f ( N This is interpreted as the sum of the individual contributions of each feature plus a baseline value. Its explanatory model... g The mathematical form is as follows:

[0110] in, M The total number of features. z ′∈{0,1} M It is a simplified vector. =1 indicates a feature k It exists (is observed). =0 indicates a feature k Missing. 0 is the baseline value, which typically represents the expected prediction of the model when all features are missing (i.e., the global average prediction). k That is, the desired feature k The Shapley value characterizes the effect of this feature on the current predicted value. f ( N ) contribution.

[0111] Shapley value k The calculation strictly follows its definition in cooperative game theory: it equals the feature kThe average marginal contribution across all possible combinations of feature subsets. The specific calculation formula is:

[0112] in, N It is the set of all features.

[0113] S yes N The middle does not contain features i any subset (i.e. S N / { i}).

[0114] f ( S ) indicates that only a subset is used. S The predicted value of the model is the feature that is not in the model (usually by using the feature that is not in the model). S The feature values ​​in the model are set to a "missing" state to simulate this.

[0115] It is a weighting factor used to ensure uniform weighting across all permutations.

[0116] That is, characteristics i Add to subset S The marginal contribution brought about by time.

[0117] By calculating the above formula (in practice, efficient approximation algorithms based on model characteristics are often used, such as TreeSHAP for tree models), each feature can be obtained. i corresponding i This value has a clear interpretation: symbol: i >0 indicates that the feature has a positive effect on predicting user retention; i A value less than 0 indicates a negative inhibitory effect (which may be related to the risk of churn).

[0118] Absolute value size: | i The size of | directly reflects the strength of the influence of this feature.

[0119] This embodiment achieves the following by employing an additive feature attribution method based on Shapley values: Precise and fair contribution allocation: Shapley values ​​satisfy fairness axioms (such as symmetry, additivity, etc.), ensuring that contribution allocation considers both the individual effects of a feature and its interactions with other features, resulting in mathematically unique and convincing results.

[0120] Model independence and global-local consistency: This method is theoretically applicable to any model ("black box" or "white box"), and the computed local explanations (explanations for individual predictions) are consistent with the global behavior of the model.

[0121] Interpretable output: Ultimately, each prediction outputs a clear list of contributions, allowing operations or product personnel to intuitively see "which behaviors" (e.g., "number of shares in the past 7 days") and "how significant the impact" (…). k =+0.15) together contributed to the prediction that "this user is likely to remain active." This provides a direct, quantitative, and credible basis for the next step of identifying core attribution factors, and is a key bridge connecting high-precision predictions with actionable business insights.

[0122] Based on the above embodiments, in another embodiment provided in this application, the method may further include: Calculate the joint interaction contribution value between at least two features, which is used to quantify the synergistic effect of the feature combination on the prediction result.

[0123] In this embodiment, although single-feature contribution analysis based on Shapley values ​​can reveal the independent impact of each feature, in real-world scenarios, user activity motivations are often not the result of a single feature, but rather the product of multiple features interacting and working together. For example, when the features "frequent use of sharing functions" and "recent participation in community activities" coexist, their positive impact on user retention may far exceed the simple sum of their independent contributions, resulting in a synergistic effect of "1+1>2". To capture and quantify this complex interaction between features, this embodiment further introduces the calculation of joint interaction contribution values ​​based on single-feature Shapley value analysis.

[0124] Specifically, for any two features i and j Their joint interaction contribution value ij Defined as: the combined effect of these two features on the model's predictions, after deducting their individual contributions, when both features are present. Its calculation is based on an extension of the Shapley interaction value, and is formalized as follows:

[0125] in: N It is the set of all features.

[0126] S For features not included i and jAny feature subset ( S N / { i , j}).

[0127] f ( ) represents the predicted value of the model given a subset of features.

[0128] Core items Features were measured i and j In subset S The pure interaction effect that occurs when they occur simultaneously in a context is the portion of their combined effect that exceeds the sum of their individual effects.

[0129] Joint interaction contribution value ij Interpretation: ij >0: Indicates a feature i and j There is a positive synergy (complementary effect), meaning that their combined effect on the predicted outcome (such as retention) is greater than the sum of their individual contributions.

[0130] ij <0: indicates a feature i and j There is a negative synergy (offsetting effect), meaning that when they occur simultaneously, they weaken each other's influence, and the combined effect is weaker than the sum of their individual contributions.

[0131] ij ≈0: This indicates that there is basically no interaction between the two, and their joint effect can be approximated as the superposition of independent contributions.

[0132] By calculating and analyzing the joint interaction contribution values ​​between features, this embodiment achieves a leap from "individual insight" to "combined insight" in attribution analysis: Unveiling the underlying motivational mechanisms: This involves identifying key feature combinations that drive user activity, rather than just isolated key features. This more closely aligns with the true psychological and behavioral patterns of user decision-making, resulting in more in-depth and accurate attribution conclusions.

[0133] Guiding precise strategy formulation: Providing more refined leverage points for product operations. For example, if a strong positive interaction is found between "using AI features" and "owning a flagship device," then AI features can be promoted to high-end device users in a focused manner, and combined benefits can be designed to maximize utility.

[0134] Enhancing model interpretability: Making the decision-making logic of complex models more transparent. By showcasing the interaction networks between important features, it helps technical and business personnel jointly understand model behavior and build a deeper level of trust in the prediction results.

[0135] This embodiment enables the final attribution analysis to not only answer "which factors are important", but also "which factors combined are more important and how important", thus providing unprecedented insights for data-driven advanced product optimization and refined operation.

[0136] After obtaining the contribution value (SHAP value) of each feature and the joint interaction contribution value between features, this embodiment provides the final step of transforming these quantitative indicators into clear and actionable business conclusions. This method aims to accurately identify the most decisive drivers from a multitude of features. Therefore, in another embodiment provided in this application, the above-mentioned determination of the attribution factors leading to user activity in the current statistical period may specifically include: (1) Sort according to the contribution of each feature and dynamically determine the set of important features based on the preset cumulative contribution ratio threshold.

[0137] First, all input features are sorted in descending order of their individual feature contribution (absolute value of SHAP). Then, instead of subjectively or fixedly selecting the top N features, a preset cumulative contribution ratio threshold P (e.g., P=80%) is introduced to dynamically define the boundaries of important features. Specifically, starting with the feature with the highest contribution, its contribution is accumulated until the cumulative contribution reaches or exceeds the preset proportion P (e.g., 80%) of the total contribution. At this point, all features involved in the accumulation constitute the set of important features. This step can be formally expressed as:

[0138] in, τ This refers to the dynamically determined number of important features, corresponding to the top ranking. τ The features are incorporated into the important feature set. This method follows the "Pareto principle" to ensure that the selected feature set can explain most of the variation in the prediction results, while avoiding noise caused by including a large number of features that contribute little.

[0139] (2) Select the feature combination with the highest joint interaction contribution value from the set of important features as the core attribution factor.

[0140] After obtaining the set of important features, further analysis is conducted on the interactions between these key features. Specifically, the joint interaction contribution value of all possible feature pairs (or specified higher-order combinations) within this set is calculated. ijThen, the joint interaction contribution values ​​of these feature combinations are sorted, and one or more feature combinations with the highest joint interaction contribution values ​​are selected.

[0141] The selected feature combination represents a group of factors that not only contribute highly on their own but also generate the strongest synergistic effect (whether positively promoting or negatively inhibiting) when combined. This embodiment identifies such feature combinations as the core attribution factors driving user activity in the current period. For example, the analysis results might show that the feature combination of "average daily uploads in the past week" and "used storage space exceeds 70%" has the highest positive joint interaction contribution value, indicating that "high-frequency upload behavior when storage space is nearing saturation" is the core composite driver for this user's high activity.

[0142] By implementing this step, this embodiment achieves the final implementation of attribution analysis: Dynamic Focus: By dynamically determining the set of important features through cumulative contribution thresholds, the analysis can adapt to the set of key driving factors behind different users or different models' predictions, thus improving the adaptability and accuracy of the method.

[0143] Deep insights: Beyond single-feature analysis, by identifying and determining the feature combination that contributes the most to joint interactions as the core attribution, we can reveal more complex and realistic user behavior mechanisms, namely that user activity is often triggered by a set of interrelated conditions or behaviors.

[0144] Precise guidance: The final output of "core attribution factors" is a clear and quantifiable combination of features, which provides extremely specific and powerful data-driven decision-making basis for product optimization (such as optimizing the linkage between storage prompts and upload processes) and precise operation (such as pushing upload incentive and expansion solution packages to users who are close to their storage limits), greatly improving the business value and application effectiveness of attribution analysis.

[0145] In yet another embodiment provided in this application, the method may further include: Based on the identified attribution characteristics, perform at least one of the following application operations: generate and push personalized content that matches the attribution factors, trigger intervention mechanisms for users at risk of churn, or optimize cloud drive APP functional modules related to core attribution factors.

[0146] This embodiment aims to seamlessly transform the profound and interpretable attribution conclusions obtained from the aforementioned analysis steps into actionable and measurable business actions, thereby achieving a closed loop from "data insights" to "business value" and enhancing user activity and product competitiveness. Specific application operations include, but are not limited to, the following three categories: (1) Generate and push personalized content that matches the attribution factors.

[0147] The system automatically generates or matches relevant personalized information from the content library based on the core attribution factors determined for a specific user or user group, and delivers it accurately through channels such as APP push, SMS or in-site messages.

[0148] Implementation: For example, if analysis reveals that the core activity attribution for a user group is a combination of "frequent use of photo backup function" and "sensitivity to storage space," then content such as "photo intelligent classification tips" tutorials or "exclusive storage expansion discounts" can be automatically generated and pushed to that group. If the attribution shows that "participation in limited-time events" is the main motivation, then invitations can be prioritized for that user when similar new events are launched.

[0149] Technical effects: It enables personalized marketing and services, significantly improves the relevance of push content and click-through rates, avoids invalid harassment, and strengthens users' perception of the core value of the product.

[0150] (2) Trigger intervention mechanisms for users at risk of churn.

[0151] The system will use negative features or feature combinations that are strongly correlated with churn risk and found in the attribution analysis (e.g., "sudden drop in login frequency" and "long time since the last use of core functions") as early warning signals to automatically trigger preset intervention processes.

[0152] Implementation: For example, when the model identifies a user with a high risk of churn and the attribution factor points to "poor experience due to insufficient storage space causing upload failure", the system can automatically trigger intervention: immediately push free temporary storage space or cleanup tool guide to the user; at the same time, mark the user and assign them to customer service or operations personnel for proactive care.

[0153] Technical Results: The system shifted from a passive to a proactive approach, enabling precise retention based on user insights. By implementing targeted interventions before user churn, it effectively improved the retention rate of high-risk users and extended their user lifecycle.

[0154] (3) Optimize the cloud disk APP functional modules related to core attribution factors.

[0155] Core attribution factors (especially feature usage characteristics) that recur and make a significant positive contribution to user activity will be used as important inputs for product iteration and optimization, guiding the improvement of functional modules or the development of new features.

[0156] Implementation: For example, if global attribution analysis generally indicates that "ease of file sharing" and "post-sharing interactive feedback" are key drivers of user activity, the product team can prioritize optimizing the interaction design of the sharing process and adding a notification function when the shared link is accessed. Conversely, if an important function is identified as a negative contributor or an irrelevant factor, it can be considered for redesign or de-prioritization.

[0157] Technical impact: It shifts product optimization decisions from experience-driven to data-driven. It ensures that R&D resources are invested in features that truly impact user activity and retention, improving the efficiency and effectiveness of product iterations and fundamentally enhancing user experience and product appeal.

[0158] To ensure the model's simplicity, interpretability, and prevent overfitting, thereby improving training efficiency and generalization ability, this embodiment includes a crucial feature selection step before inputting the preprocessed feature data into the model for training. This step aims to eliminate features with weak predictive power or redundancy from a large pool of candidate features, focusing on a subset of features that substantially contribute to predicting user retention. Therefore, in this embodiment, a feature selection step is also included before training the fusion prediction model. This feature selection step includes: (1) Calculate the IV value of each behavioral feature data.

[0159] First, for each feature to be screened (especially those that have been discretized or binned), its Information Value (IV) is calculated. The IV value is a commonly used quantitative indicator to measure the predictive power of a feature on the target variable (in this application, user retention). Its calculation is based on the difference between the distributions of positive (retained users) and negative (non-retained users) samples under different value groups for that feature. The specific calculation typically involves the following steps: Appropriate binning discretization is performed on continuous features.

[0160] Within each bin, calculate the proportion of positive samples in that bin to the total number of positive samples. P good ), and the proportion of negative samples to total negative samples ( P bad ).

[0161] The weight of evidence (WOE) for each bin is calculated using the following formula: WOE =ln( P good / P bad ).

[0162] The IV value of this feature is calculated using the following formula: .

[0163] A higher IV value indicates a stronger predictive ability of the feature for the target variable; a lower IV value indicates a weaker predictive ability.

[0164] (2) Delete features whose IV values ​​are lower than the preset prediction ability threshold.

[0165] Set a preset predictive ability threshold (e.g., IV < 0.02). This threshold, determined based on experience and business scenarios, is used to distinguish between features with predictive value and ineffective features. Compare the calculated IV value of each feature with this threshold, and delete all features with IV values ​​below this threshold. These deleted features are generally considered to have no meaningful correlation between their value distribution and user retention, contributing negligibly to model prediction. Retaining them is not only unhelpful but may also introduce noise, increase model complexity, and lead to overfitting.

[0166] In this embodiment, to further address the multicollinearity problem among features, this step can be combined with correlation analysis. For example, after deleting features with low IV values, the Pearson correlation coefficient among the remaining features can be calculated. For feature pairs with excessively high correlation coefficients (e.g., greater than 0.7), the feature with the higher IV value is retained, and the other is deleted.

[0167] This embodiment removes noisy features, allowing the model to focus more on learning patterns with strong predictive signals, thus contributing to the construction of a more robust and generalizable prediction model. Furthermore, it reduces the dimensionality of input features, lowering the computational and storage overhead of model training and accelerating the training process. The final feature set input to the model consists of features that have been quantitatively verified to contribute to prediction, making subsequent attribution analysis based on methods such as Shapley values ​​clearer and more reliable, avoiding misinterpretations of irrelevant features.

[0168] In the embodiment, in the step of obtaining the behavioral characteristic data of cloud disk APP users within the current statistical period, the current statistical period is the most recent 30 days.

[0169] In this embodiment, to balance the timeliness and stability of behavioral pattern capture, the "current statistical period" is specifically set to the most recent 30 days. This time window is chosen based on the following considerations: Firstly, a 30-day period is sufficient to cover most short-term user behavioral habits (such as weekly activity patterns and participation in monthly activities) and form statistically significant sequence data, ensuring that the features have sufficient representativeness and discriminative power. Secondly, it avoids the lag in historical data caused by excessively long time spans (e.g., behavior from six months ago has significantly diminished its impact on current activity), and also prevents the period from being too short (e.g., only the most recent 7 days) from failing to reflect stable patterns due to occasional fluctuations. This setting aligns with common practices in mobile internet product user behavior analysis, ensuring data timeliness while providing a stable and representative snapshot of recent behavior for subsequent time-series modeling and prediction.

[0170] In this embodiment, the preprocessing step for the above-mentioned behavioral feature data further includes: standardizing the continuous features, wherein the standardization process adopts the Z-score normalization method.

[0171] To eliminate model training bias caused by differences in units and value ranges among different continuous features, and to ensure that each feature has a comparable starting point of importance in model learning, this embodiment adds a standardization step for continuous features in the preprocessing. Specifically, the classic Z-score normalization method is adopted, which is calculated as follows: for any continuous feature, subtract the mean of that feature among all samples from each of its original values, and then divide by its standard deviation. After this processing, the numerical distribution of the feature will be transformed into a standard normal distribution with a mean of 0 and a standard deviation of 1. This processing transforms features at different scales to the same scale, which helps improve the training speed and stability of gradient descent-based models (such as LSTM) and prevents certain features with large numerical ranges from occupying an inappropriate dominant position in the model (such as distance-based algorithm components). This step is a routine but crucial step in data preprocessing, laying an important foundation for the subsequent model to efficiently and fairly learn the complex relationship between features and the target.

[0172] In the embodiment, in the step of training the fusion prediction model for predicting whether a user will remain in the next statistical period, the K-fold cross-validation method is used to evaluate and select the model performance.

[0173] To ensure the trained fusion prediction model possesses robust generalization ability and avoids evaluation bias or overfitting caused by a single training-validation set partition, this embodiment introduces a K-fold cross-validation method during the model training and optimization stages. Specifically, the preprocessed and clustered user sample data is randomly divided into K mutually exclusive subsets of roughly equal size (e.g., K=5 or 10). In each training iteration, one subset is used as the validation set, and the remaining K-1 subsets are combined as the training set. This is used to train the model and evaluate its performance (e.g., accuracy, AUC, F1 score). This process is repeated K times, ensuring that each subset is used as a validation set once. Finally, the average of the K evaluation results is used as a reliable estimate of the model performance. This method fully utilizes limited data, ensures the stability and reliability of the model evaluation results, and provides a scientific and objective basis for comparing different model architectures (e.g., feature-level fusion and model-level fusion), hyperparameter tuning, and the final model selection. Especially in the implementation that incorporates user segmentation, K-fold cross-validation can be performed independently within each user group, thereby accurately selecting the optimal prediction model for different groups.

[0174] In this embodiment, when using the additive feature attribution method based on Shapley values, the TreeExplainer interpreter is specifically used to calculate the feature contribution of the tree model part in the fusion prediction model.

[0175] When using Shapley values ​​to perform interpretable analysis of fusion prediction models, considering that the models constructed in this solution typically include tree-based machine learning ensemble models (such as XGBoost, Random Forest, etc.), this embodiment preferably employs the TreeExplainer interpreter, specifically optimized for tree models, to balance accuracy and computational efficiency. This interpreter is a crucial component of the SHAP (Shapley Additive exPlanations) framework. It leverages the unique structure of tree models (such as branching conditions and leaf node values) to efficiently and accurately calculate the Shapley value of each feature in a specific prediction, without resorting to time-consuming feature permutations or sampling approximations. Its calculation process strictly adheres to the theoretical definition of Shapley values, while algorithm optimization reduces computational complexity from exponential to polynomial levels related to tree depth and the number of leaf nodes, enabling rapid processing of high-dimensional features and large-scale data. Using TreeExplainer not only ensures accurate and consistent feature contribution quantification for the tree model portion but also guarantees the feasibility and efficiency of the entire attribution explanation process in real-world business systems, serving as a key technological implementation connecting high-performance fusion prediction models with actionable business insights.

[0176] In this embodiment, the method further includes a model deployment step: encapsulating the trained fusion prediction model and interpretive model into an application programming interface (API) for use by the background service system of the cloud disk APP; wherein, the interpretive model is used to calculate the contribution of each behavioral feature to the prediction result.

[0177] To enable seamless integration of the attribution analysis method provided in this application into actual business systems, achieving automated and real-time analysis and decision support, this embodiment further describes the model deployment steps. Specifically, the fully trained and validated fusion prediction model (used to predict user retention) and its supporting interpretive model (used to calculate feature contribution) are engineered and encapsulated into a standard application programming interface (API) service. This API service is typically deployed on cloud servers or enterprise internal computing platforms, defining clear input (such as user ID or feature vector) and output (such as retention prediction probability, feature contribution list) specifications. The backend service system of the cloud disk APP (such as user profiling system, personalized recommendation engine, operation management platform) can call this API via the network to obtain prediction results and attribution analysis reports for specific users in real time. This deployment method achieves decoupling and reuse of analytical capabilities and business systems, enabling complex attribution analysis to serve various real-time business scenarios such as dynamic marketing, churn warning, and product optimization in a low-latency and highly available manner. It completes the last mile of implementation from offline models to online intelligent services, greatly enhancing the practical value and commercial efficiency of this technical solution.

[0178] In this embodiment, the method can run in a distributed computing framework, and the preprocessing of the behavioral feature data and the model training process are executed in parallel on multiple computing nodes.

[0179] To address the computational and storage challenges posed by massive amounts of user behavior data and ensure the efficiency and scalability of analysis and processing, this embodiment supports operation within distributed computing frameworks (such as Apache Spark, Hadoop, or cloud-based distributed processing services). Specifically, computationally intensive tasks such as data acquisition, preprocessing (including frequency domain analysis and anomaly detection), and even model training (such as large-scale hyperparameter search and parallel training of multi-user group models) are decomposed into multiple parallelizable subtasks and scheduled to run simultaneously on multiple computing nodes in a distributed cluster. For example, data from different user groups can be allocated to different nodes for independent feature engineering and model training; different parameter combinations in hyperparameter grid search can also be attempted in parallel across nodes. This parallel processing mechanism significantly shortens the entire process time from raw data to final attribution results, enabling this method to efficiently process data from hundreds of millions of users. This meets the requirements of modern internet companies for timely analysis and large-scale data processing capabilities, and is a crucial engineering guarantee for this technical solution to be deployed in actual production environments and realize its commercial value.

[0180] This application's embodiments maintain consistency in user activity motivations within a short period, transforming statistical motivation analysis methods into supervised learning-based retention methods. This allows for the prediction of user behavior in the next cycle based on short-term activity characteristics. The process involves acquiring feature data of active users within the app; introducing frequency domain analysis-based noise detection and multi-dimensional feature fusion methods for handling abnormal data and users, and preprocessing the features; introducing a dynamic decay function that combines time and feature importance for dynamic exponential decay; in the deep learning algorithm, introducing sample importance weights and dynamically adjusting the penalty factor α; and fusing the optimized deep learning algorithm with ensemble learning algorithms in machine learning to obtain a model that meets performance requirements; introducing the SHAP interpretation method, constructing a dynamic adjustment method for attribution thresholds, and a SHAP-based joint feature interaction analysis interpretation method to determine the contribution of various temporal and static feature combinations to the model's predicted values; ultimately determining the main motivations for user activity within a short period and innovating application scenarios based on this.

[0181] Therefore, in another embodiment provided by this application, this application provides an attribution analysis method for the motivations of cloud storage app users' activity, such as... Figure 2 As shown, the method may include the following steps: S210: Obtain user data to be predicted from active users. For the time window selection, choose an appropriate time window based on the modeling requirements, such as data from the most recent 30 days.

[0182] In terms of time dimension, it covers information from multiple dimensions, including basic information features, version model, APP operation behavior features, terminal features, network features, and user consumption behavior features. This data forms a historical training dataset for activity attribution prediction.

[0183] This activity attribution prediction focuses on analyzing the reasons for user activity, specifically exploring the following five dimensions: Functional features: Covers user operations on the APP, including 30 functional features such as backup, upload, photo album, transfer, video viewing, reading, AI functions, audiobook, and sharing.

[0184] User consumption behavior characteristics include user plan type, data usage, and average monthly spending.

[0185] Activity participation information: Records users' participation in activities within the app.

[0186] Network characteristics: Reflects the network speed and quality experienced by users while using the app.

[0187] Terminal characteristics: This includes information such as the user's terminal brand, price, and telecommunications provider.

[0188] For details, please refer to Table 1.

[0189] Table 1:

[0190] S220, Data Preprocessing and User Segmentation. This involves handling missing and outlier values, data cleaning, user segmentation, and cross-validation to divide the dataset and validation set.

[0191] (1) Introduce outlier processing of noise detection method based on frequency domain analysis.

[0192] Missing value imputation. The proportion of missing values ​​for statistical features is used; features with more than 50% missing values ​​are directly deleted, while features with less than 50% missing values ​​are imputed using a random forest model from the scikit-learn library in Python. Data standardization is then performed.

[0193] In outlier handling, frequency domain analysis is innovatively applied to a non-traditional signal processing field: user behavior data. Taking advantage of the time-series nature of user behavior data, a dynamic noise detection method based on frequency domain analysis is introduced. The core idea is to use frequency domain analysis techniques to identify and filter noise features in the behavior data, thereby optimizing subsequent modeling and attribution analysis.

[0194] By using the Discrete Fourier Transform formula, the time-series data of user behavior is transformed from the time domain to the frequency domain and decomposed into multiple frequency components:

[0195] Where X(k) represents the frequency domain component, x(n) is the time series data, and N is the total number of data points. This step maps the behavioral data from the time domain representation to the frequency domain representation, effectively distinguishing the energy of different frequency components.

[0196] Frequency domain analysis identifies anomalous peaks in high-frequency components. These high-frequency components are often associated with data noise, such as high-frequency fluctuations caused by abnormal user operations or system jitter. Anomaly detection involves calculating the amplitude of each frequency component in the frequency domain and dynamically setting thresholds (such as local fractions or fixed energy thresholds) using statistical methods to detect high-frequency anomalies.

[0197] Finally, an inverse Fourier transform is performed to restore the data. After filtering out high-frequency noise, the data is restored to the time domain using the restored time domain data. High-frequency noise is effectively filtered out, and the main behavioral feature signals are preserved.

[0198]

[0199] A sliding window technique is introduced into the frequency domain analysis process. The window length is dynamically adjusted according to the time granularity of user behavior, and the time series is divided into sliding windows. The frequency domain characteristics of the data within each window are calculated in real time. The mean and standard deviation within the window range are dynamically adjusted to remove local outliers and ensure that the method can adapt to the needs of data stream processing.

[0200] (2) Construct an abnormal user detection mechanism based on multi-dimensional feature fusion.

[0201] By integrating device features, time series features, user attribute features, and usage scenario features, a unified high-dimensional feature vector is generated. Each feature represents device usage, time-based behavior, user attributes, and scenario preferences. Abnormal features are defined and extracted for detection.

[0202] Among them, for multiple accounts on one device, the correlation between the device and the account is checked: AnomalousDevice={Device∣AccountCount(Device)>Threshold}, with the threshold set to 4. If the same device logs in to more than 4 accounts within a month, it is judged as abnormal. Among them, more than 90% of users log in to 4 or fewer accounts.

[0203] Analyze login frequency: If the login frequency is much higher than the average of normal users, it will be marked as abnormal.

[0204] (3) Introduction of dynamic decay function.

[0205] This application example introduces a dynamic decay function in the data processing of user activity attribution. Time and feature importance weights are introduced into the features, and the dynamic decay function is used to adjust the importance of the features. By assigning time weights to features, their changes over time are reflected, thereby dynamically adjusting the model's sensitivity to data from different time periods. In this embodiment, the adjustment function uses an exponential decay model, with the basic model as follows:

[0206] in, It is the decay rate, calculated by using a random forest tree model algorithm to predict the importance score of features. , It is a time interval.

[0207]

[0208]

[0209] Based on this, the influence range of different features is considered, and the decay rate is dynamically adjusted according to the importance of the features. This allows for a slower decay of highly important features.

[0210] (4) Feature filtering.

[0211] Preliminary selection based on IV value: Selecting appropriate features can effectively prevent overfitting. In this application, a feature processing method based on IV value is selected. IV value represents the predictive ability of input features to target features. Features with IV values ​​less than 0.02 that have no predictive ability are deleted. The correlation between features is judged by combining the Pearson correlation coefficient. Features with small IV values ​​and correlation coefficients > 0.7 are deleted. The selection results are used as the input features of the model.

[0212] (5) User group processing.

[0213] By using clustering models to group users, and analyzing the characteristics of each cluster based on the clustering results in an unsupervised manner, the model can be trained specifically for each cluster, which may have different behavioral patterns and feature distributions. This allows for more accurate prediction of user behavior across different groups. Clustering reduces generalization error between different groups, thereby improving the overall model's generalization ability and prediction accuracy. Furthermore, it allows for focused analysis of the feature scores and influencing factors of each group, leading to a better understanding of the factors behind the model results and user behavior.

[0214] S230, Model Training and Evaluation. This involves feature extraction from preprocessed data, model training, model fusion, model evaluation, and model iteration.

[0215] 1. Deep learning model construction.

[0216] Given that the learning scenario leans towards analyzing daily user behavior and the time-series characteristics of user behavior data, such as clicks, logins, purchases, and logouts exhibiting temporal correlation (i.e., strong dependencies between actions), and that behavior data may contain long-term patterns (e.g., periodic activity) and short-term patterns (e.g., high-frequency behavior during an event), capturing the long-term and short-term dependencies in the time series is a key consideration when introducing LSTM.

[0217] LSTM is specifically designed for processing time series data and effectively addresses the problem of long-term dependencies. Through its memory units and gating mechanisms (input gate, forget gate, output gate), LSTM can memorize key information over long time spans while forgetting irrelevant historical data. Furthermore, for user behavior data, LSTM can capture short-term active patterns and long-term behavioral trends in time series data.

[0218] The formula for an LSTM cell is expressed as follows: 1) Gate of Oblivion:

[0219] in It is input. It is the hidden state of the previous time step. It determines how much of the past information to forget.

[0220] 2) Input Gate:

[0221]

[0222] Decide how much new input information to retain.

[0223] 3) Cell state update:

[0224] in It is the cell state at the previous time step. 4) Output gate:

[0225]

[0226] Determine the output of the current hidden state.

[0227] The following optimizations are added to this scheme during the LSTM training process: ① Sample weighting: To address the problem of unbalanced sample distribution, dynamic weight adjustment is used to assign differentiated contributions to different samples during the training process, thereby distinguishing the importance of samples.

[0228]

[0229] in, Based on the user lifecycle activity level in the actual business, such as the introduction phase, growth phase, maturity phase, dormancy phase, and churn phase, different weights are assigned, as follows: Introducing users: Data is scarce but has high potential value, ω=1.0.

[0230] Growth users: key stage of conversion, high value, ω=1.2.

[0231] Mature users: exhibit stable behavior, are familiar with the product, and contribute high value, ω=1.5.

[0232] Silent users: Sparse behavior, reduced importance, ω=0.8.

[0233] Churned users: Behavioral patterns may contain key features, with slightly lower but not zero weight, ω=0.5.

[0234] ② False negative penalty factor: Additional penalty is imposed on key categories that are missed in the prediction to improve the model's sensitivity to low-probability events.

[0235]

[0236] The implementation example introduces a penalty factor. This approach imposes a higher penalty on false negatives, causing the model to pay more attention to them during training. This optimization effectively reduces significant errors in retention rate prediction, captures key features relevant to business needs, and thus improves overall performance.

[0237] ③ Dynamic time window: By combining the time dynamics of user behavior data, the time span of the input sequence is adjusted, making the model more flexible in capturing short-term and long-term dependencies.

[0238] 2. Based on the integration of deep learning and XGboost.

[0239] This approach combines the deep learning algorithm LSTM with the traditional machine learning algorithm XGboost for fusion modeling and training, and employs two fusion strategies for performance analysis: model fusion and feature fusion.

[0240] Feature-level fusion: Use LSTM to extract time series features, and output The features are then concatenated with the user's other static features to form a new feature matrix. This new feature matrix is ​​then input into XGBoost for training.

[0241] Model-level fusion: LSTM and XGBoost are trained separately to obtain two independent models. The prediction results of the two models are then fused using a stacking method.

[0242] in, , These are model weights, adjusted based on validation set performance.

[0243] 3. Model Training Based on the above algorithm framework, the training of the model includes data input, feature engineering, and model parameter optimization to obtain the trained model.

[0244] S240, Feature Importance Optimization and Attribution Weight Determination. Feature importance is optimized using the Shap interpretation improvement algorithm introduced in this case. Finally, by dynamically adjusting the attribution weights and joint feature contributions, the final attribution result of the model is output.

[0245] 1. Introduce the Shap interpretation algorithm.

[0246] This scheme introduces the Shap interpretation algorithm and innovatively extends it from single feature contributions to contributions from combined features. Traditional ensemble learning and deep learning models are limited by their ability to assess the importance of only a single feature, failing to clearly define how features influence prediction results. By incorporating the Shap algorithm and extending the attribution interpretation method for joint features into the model algorithm, this approach not only quantifies the specific impact of each feature on the prediction result but also clearly reflects the positive or negative nature of this impact, providing a more intuitive basis for model interpretability and decision optimization.

[0247] The algorithm is based on The value is interpreted as an additive feature attribution method, and the model's predicted value is interpreted as a linear function of two variables, with the function formula as follows:

[0248] in, Representation and explanation model, Indicates the number of input features. Represent a constant, , Does the observed feature exist? = or = , Represent each feature The attribution value.

[0249] in, Represents the feature set, express{ A subset of}, where the fraction represents the probability corresponding to different combinations of features. Representing different combinations of features The prediction results of the model when it is included in the model and when it is not included in the model.

[0250] The formula for the Shap value of joint contributions is:

[0251] The Shap interpretation algorithm is introduced, which can calculate the Shap value based on TreeExplainer in the interpretation library.

[0252] 2. Calculate feature contributions and construct feature interaction contributions.

[0253] Single feature contribution: Calculate the Shap value of the XGBoost model for each input feature.

[0254] Time series feature contribution: output of time steps via LSTM The Integrated Gradients method is used to estimate the feature importance of LSTM in the time dimension.

[0255] Feature interaction contribution and joint contribution: We use Shap interaction values ​​to calculate the interaction contribution between the LSTM output and the XGBoost static features. Feature interaction refers to the joint effect of two or more features on the target variable exceeding the sum of their individual effects. Traditional models only consider the contribution of a single feature, ignoring the complex interaction relationships between features. We propose an extension based on Shap feature interaction, which combines multiple features to form the joint contribution of multiple model outputs, especially for sequence features. and static features Expanding Shap to a joint interpretation:

[0256] in, This represents the pure interaction contribution between two features. and Representing sequence features respectively and static features The independent contribution (without considering interactions with other features).

[0257] The contribution of interactive features measures the boosting or offsetting effect between features. >0, sequence features and static features There is positive interaction, which collaboratively improves prediction; if <0, sequence characteristics and static features There are negative interactions, which cancel each other out.

[0258] By assessing the overall contribution of feature combinations to the model's active attribution through joint contributions, we can guide feature selection or feature combination.

[0259] By combining attribution thresholds to rank individual attribution features and identifying high-contribution features, and based on... Calculate and sort the joint contributions of all pairwise combinations of features to find the feature combination with the largest interaction contribution.

[0260] 3. Construct a dynamic adjustment mechanism for attribution thresholds.

[0261] Dynamic interpretation of attribution threshold: Automatically adjusts the attribution threshold based on the feature contribution distribution. Before dynamic adjustment k Cumulative percentage of Shap values ​​for each feature:

[0262] in, P Set it to 80%, defined according to the Pareto principle to ensure that more important features are covered.

[0263] By combining attribution thresholds to rank individual attribution features and identifying high-contribution features, and based on... The joint contributions of time-series features and static features are calculated and ranked to find the feature combination with the largest interaction contribution as the final result of user attribution.

[0264] S250, Application of attribution methods.

[0265] In cloud storage operations, attribution models are used to analyze and understand the motivations behind user activity and payment behaviors. This allows the operations team to more accurately target user groups and adjust marketing strategies based on attribution results, thereby improving user engagement and conversion rates, and achieving precise operations. This is mainly reflected in the following aspects: 1. Dynamic marketing and targeted push notifications.

[0266] Personalized recommendation strategy: Dynamically adjust recommended content based on user activity attribution results. For example, design a precise push strategy based on user preferences for a specific functional module (such as file sharing or online collaboration).

[0267] Time-sensitive push notifications: Optimize the timing of marketing campaign push notifications based on users' time activity characteristics (such as high-frequency use on weekdays or surges on holidays) to improve conversion rates.

[0268] Scenario-based event design: Design differentiated promotional activities or membership benefits based on user groups (e.g., storage users and sharing users).

[0269] 2. User churn warning and intervention.

[0270] Churn user early warning: By combining attribution analysis to identify low-frequency or potential churn users, early intervention can be carried out through abnormal attribution characteristics (such as gradually decreasing login frequency).

[0271] Behavioral incentive mechanisms: Design incentives that match attribution results, such as providing free storage expansion or special feature experiences, to encourage user return.

[0272] Reactivating dormant users: For users whose attribution results show a reason for inactivity (such as running out of high storage space), push relevant solutions (such as cleanup tools or expansion discounts).

[0273] 3. Function optimization and product improvement.

[0274] Functional module evaluation: Based on the attribution results, evaluate the actual usage effect of certain functional modules (such as which functional modules users use frequently) and optimize low-usage functions.

[0275] Interaction design optimization: After analyzing attribution characteristics, optimize key interaction points in the user's active path (such as reducing the number of steps to upload files).

[0276] New Feature Exploration: Leveraging attribution analysis results, develop features that meet users' deeper needs. For example, for users whose core motivation is sharing frequency, features integrated with social networking can be further explored.

[0277] This application optimizes the ensemble algorithm model by introducing sample weights and a penalty factor for false negative classification errors. This makes the model training more focused on active samples and such errors, improving the accuracy of the model for high-value active customer groups. By training the model in subgroups after user clustering, the prediction accuracy for segmented groups is improved. User features are optimized by introducing a dynamic decay function, thereby enhancing the model's sensitivity to recent behavior. The Shap interpretation method is introduced, constructing a dynamic adjustment method for attribution thresholds and a Shap-based joint feature interaction analysis interpretation method. This provides insights into user activity attribution issues and enables the determination of user attribution characteristics based on attribution scores. Furthermore, innovative methods for handling abnormal data and abnormal users are introduced, including frequency domain analysis-based noise detection and multi-dimensional feature fusion, to preprocess features. In cloud storage operation practices, the attribution model has been applied in multiple scenarios, including dynamic marketing and precise push notifications, user churn warning and intervention, and function optimization and product improvement.

[0278] By dividing each function into corresponding functional modules, this application provides an attribution analysis device for the activity motivation of cloud storage APP users. The attribution analysis device for the activity motivation of cloud storage APP users can be a server, a terminal, or a chip applied to a server. Figure 3 This is a schematic block diagram of the functional modules of an attribution analysis device for the motivations of cloud storage app users, provided as an exemplary embodiment of this application. Figure 3 As shown, the attribution analysis device for the motivations of cloud storage app users' activity includes: Feature data acquisition module 31 is used to acquire behavioral feature data of cloud disk APP users in the current statistical period; Preprocessing module 32 is used to preprocess the behavioral feature data, and the preprocessing is used for noise detection and filtering, as well as abnormal user detection; Model training module 33 is used to train a fusion prediction model based on preprocessed behavioral feature data to predict whether a user will remain in the next statistical period. The contribution calculation module 34 is used to calculate the contribution of each behavioral feature to the prediction result based on the prediction result of the fusion prediction model through an interpretive model. Attribution factor determination module 35 is used to determine, based on the contribution, the attribution factors that led to the user's activity in the current statistical period.

[0279] In another embodiment provided in this application, the preprocessing of the behavioral feature data further includes introducing a dynamic decay function to assign a weight that decays over time to each behavioral feature based on the time of occurrence of the behavior and the importance of the feature.

[0280] In another embodiment provided in this application, the preprocessing of the behavioral feature data further includes user grouping processing, which divides users into different groups based on user features, and performs subsequent model training and attribution analysis for each user group.

[0281] In another embodiment provided in this application, the noise detection and filtering includes: Transform user behavior data from the time domain to the frequency domain; Identify and filter high-frequency components above a preset threshold in the frequency domain; The filtered frequency domain data is inversely converted back to the time domain to obtain the denoised behavioral feature data.

[0282] In another embodiment provided in this application, when converting user behavior feature data from the time domain to the frequency domain, a sliding window technique is used to segment the time series data and dynamically adjust the statistical thresholds used to detect outliers within each window.

[0283] In another embodiment provided in this application, the abnormal user detection includes: Construct a high-dimensional feature vector that integrates device features, time series features, user attribute features, and usage scenario features; The high-dimensional feature vector is detected based on preset rules, including determining that the number of accounts logged in by the same device within a preset time period exceeds a first threshold, or the user login frequency exceeds a second threshold.

[0284] In another embodiment provided in this application, the fusion prediction model is a fusion model of a deep learning model and a machine ensemble learning model.

[0285] In another embodiment provided in this application, the deep learning model is a long short-term memory network model. When training the long short-term memory network model, different importance weights are assigned to the training samples according to the user's life cycle stage.

[0286] In another embodiment provided in this application, when training the long short-term memory network model, a penalty factor is introduced to impose an additional loss penalty on false negative prediction errors. The penalty factor is determined based on the ratio of positive to negative samples in the training data.

[0287] In another embodiment provided in this application, the fusion method of the deep learning model and the machine ensemble learning model is feature-level fusion, including: The deep learning model is used to extract temporal features of user behavior sequences; The temporal features are concatenated with the user's static features to form a fused feature; The fused features are input into the machine ensemble learning model for training and prediction.

[0288] In another embodiment provided in this application, the fusion method of the deep learning model and the machine ensemble learning model is model-level fusion, including: Deep learning sub-models and machine ensemble learning sub-models are trained separately. Obtain the first prediction result of the deep learning sub-model for the target user, and the second prediction result of the machine ensemble learning sub-model for the target user; The first prediction result and the second prediction result are weighted and summed to obtain the final prediction result.

[0289] In another embodiment provided in this application, the calculation of the contribution of each behavioral feature to the prediction result includes: An additive feature attribution method based on Shapley values ​​is used to calculate the Shapley value of each input feature for the predicted value output by the fusion prediction model, and this value is used as the contribution of each behavioral feature to the prediction result.

[0290] In another embodiment provided in this application, the apparatus further includes a joint interaction contribution value calculation module, specifically used for: Calculate the joint interaction contribution value between at least two features, which is used to quantify the synergistic effect of the feature combination on the prediction result.

[0291] In another embodiment provided in this application, determining the attribution factors that cause the user to be active during the current statistical period includes: The features are sorted according to their contribution and the set of important features is dynamically determined based on a preset cumulative contribution ratio threshold. From the set of important features, the feature combination with the highest joint interaction contribution value is selected as the core attribution factor.

[0292] In another embodiment provided in this application, the device further includes an operation execution module, specifically used for: Based on the identified attribution characteristics, perform at least one of the following application operations: generate and push personalized content that matches the attribution factors, trigger intervention mechanisms for users at risk of churn, or optimize cloud drive APP functional modules related to core attribution factors.

[0293] In another embodiment provided in this application, before training the fusion prediction model, a feature selection step is further included, the feature selection step comprising: Calculate the IV value for each behavioral feature data; Features with IV values ​​below a preset predictive ability threshold are deleted.

[0294] In another embodiment provided in this application, in the step of obtaining behavioral characteristic data of cloud disk APP users within the current statistical period, the current statistical period is the most recent 30 days.

[0295] In another embodiment provided in this application, the step of preprocessing the behavioral feature data further includes: standardizing the continuous features, wherein the standardization process adopts the Z-score normalization method.

[0296] In another embodiment provided in this application, in the step of training a fusion prediction model for predicting whether a user will remain in the next statistical period, the K-fold cross-validation method is used to evaluate and select the model performance.

[0297] In another embodiment provided in this application, when using the additive feature attribution method based on Shapley values, the TreeExplainer interpreter is specifically used to calculate the feature contribution of the tree model part in the fusion prediction model.

[0298] In another embodiment provided in this application, the preset threshold is dynamically determined based on the partial number of the frequency domain component amplitude.

[0299] In another embodiment provided in this application, the apparatus further includes a model deployment module, specifically used for: The trained fusion prediction model and interpretive model are encapsulated into an application programming interface (API) for use by the background service system of the cloud storage app; wherein, the interpretive model is used to calculate the contribution of each behavioral feature to the prediction result.

[0300] In another embodiment provided in this application, the device operates in a distributed computing framework, and the preprocessing of the behavioral feature data and the model training process are executed in parallel on multiple computing nodes.

[0301] For details about the device, please refer to the corresponding description of the method described above; it will not be repeated here.

[0302] This application also provides an electronic device, including: at least one processor; a memory for storing executable instructions of the at least one processor; wherein the at least one processor is configured to execute the instructions to implement the method disclosed in the embodiments of this application.

[0303] Figure 4 This is a schematic diagram of the structure of an electronic device provided as an exemplary embodiment of this application. For example... Figure 4 As shown, the electronic device 1800 includes at least one processor 1801 and a memory 1802 coupled to the processor 1801. The processor 1801 can perform the corresponding steps in the methods disclosed in the embodiments of this application.

[0304] The processor 1801 described above can also be called a central processing unit (CPU), which can be an integrated circuit chip with signal processing capabilities. Each step in the method disclosed in this application can be implemented by the integrated logic circuitry in the hardware of the processor 1801 or by instructions in software form. The processor 1801 can be a general-purpose processor, a digital signal processor (DSP), an ASIC (Application Specific Integrated Circuit), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. The general-purpose processor can be a microprocessor or any conventional processor. The steps of the method disclosed in this application can be directly implemented by a hardware decoding processor, or implemented by a combination of hardware and software modules in the decoding processor. The software modules can be located in the memory 1802, such as random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, registers, or other mature storage media in the art. The processor 1801 reads information from the memory 1802 and, in conjunction with its hardware, completes the steps of the above method.

[0305] Furthermore, the various operations / processes according to this application, when implemented via software and / or firmware, can be transmitted from a storage medium or network to a computer system with a dedicated hardware architecture, such as... Figure 5 The computer system 1900 shown is equipped with the programs that constitute the software. When various programs are installed, the computer system is able to perform various functions, including those described above. Figure 5 A structural block diagram of a computer system provided for an exemplary embodiment of this application.

[0306] Computer System 1900 is intended to represent various forms of digital electronic computer devices, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. Electronic devices can also represent various forms of mobile devices, such as cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the present application described and / or claimed herein.

[0307] like Figure 5 As shown, the computer system 1900 includes a computing unit 1901, which can perform various appropriate actions and processes based on a computer program stored in a read-only memory (ROM) 1902 or a computer program loaded from a storage unit 1908 into a random access memory (RAM) 1903. The RAM 1903 may also store various programs and data required for the operation of the computer system 1900. The computing unit 1901, ROM 1902, and RAM 1903 are interconnected via a bus 1904. An input / output (I / O) interface 1905 is also connected to the bus 1904.

[0308] Multiple components in computer system 1900 are connected to I / O interface 1905, including: input unit 1906, output unit 1907, storage unit 1908, and communication unit 1909. Input unit 1906 can be any type of device capable of inputting information into computer system 1900. Input unit 1906 can receive input digital or character information and generate key signal inputs related to user settings and / or function control of the electronic device. Output unit 1907 can be any type of device capable of presenting information and may include, but is not limited to, a monitor, speaker, video / audio output terminal, vibrator, and / or printer. Storage unit 1908 may include, but is not limited to, hard disks and optical disks. Communication unit 1909 allows computer system 1900 to exchange information / data with other devices via a network such as the Internet, and may include, but is not limited to, modems, network cards, infrared communication devices, wireless communication transceivers, and / or chipsets, such as Bluetooth™ devices, WiFi devices, WiMax devices, cellular communication devices, and / or the like.

[0309] The computing unit 1901 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 1901 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 1901 performs the various methods and processes described above. For example, in some embodiments, the methods disclosed in the embodiments of this application can be implemented as a computer software program, which is tangibly contained in a machine-readable medium, such as storage unit 1908. In some embodiments, part or all of the computer program can be loaded and / or installed on an electronic device via ROM 1902 and / or communication unit 1909. In some embodiments, the computing unit 1901 can be configured to perform the methods disclosed in the embodiments of this application by any other suitable means (e.g., by means of firmware).

[0310] This application also provides a computer-readable storage medium, wherein when the instructions in the computer-readable storage medium are executed by a processor of an electronic device, the electronic device is able to perform the methods disclosed in this application.

[0311] The computer-readable storage medium in this application embodiment may be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. The aforementioned computer-readable storage medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specifically, the aforementioned computer-readable storage medium may include an electrical connection based on one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination of the foregoing.

[0312] The aforementioned computer-readable medium may be included in the aforementioned electronic device; or it may exist independently and not assembled into the electronic device.

[0313] This application also provides a computer program product, including a computer program, wherein the computer program, when executed by a processor, implements the methods disclosed in the embodiments of this application.

[0314] In embodiments of this application, computer program code for performing the operations of this application can be written in one or more programming languages ​​or a combination thereof. These programming languages ​​include, but are not limited to, object-oriented programming languages ​​such as Java, Smalltalk, and C++, as well as conventional procedural programming languages ​​such as C or similar languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network (including a local area network (LAN) or a wide area network (WAN)), or it can be connected to an external computer.

[0315] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this application. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.

[0316] The modules, components, or units described in the embodiments of this application can be implemented in software or hardware. The names of the modules, components, or units do not necessarily constitute a limitation on the module, component, or unit itself.

[0317] The functions described above in this document can be performed at least in part by one or more hardware logic components. For example, without limitation, exemplary hardware logic components that can be used include: field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), system-on-a-chip (SoCs), complex programmable logic devices (CPLDs), and so on.

[0318] The above description is merely an embodiment of this application and an explanation of the technical principles employed. Those skilled in the art should understand that the scope of disclosure in this application is not limited to technical solutions formed by specific combinations of the above-described technical features, but should also cover other technical solutions formed by arbitrary combinations of the above-described technical features or their equivalents without departing from the above-described concept. For example, technical solutions formed by substituting the above features with (but not limited to) technical features with similar functions disclosed in this application.

[0319] While specific embodiments of this application have been described in detail by way of examples, those skilled in the art should understand that the above examples are for illustrative purposes only and are not intended to limit the scope of this application. Those skilled in the art should understand that modifications can be made to the above embodiments without departing from the scope and spirit of this application. The scope of this application is defined by the appended claims.

Claims

1. An attribution analysis method for the motivations of user activity in cloud storage apps, characterized in that, The method includes: Obtain behavioral characteristic data of cloud storage app users within the current statistical period; The behavioral feature data is preprocessed for noise detection and filtering, as well as abnormal user detection. Based on the preprocessed behavioral feature data, a fusion prediction model is trained to predict whether users will remain in the next statistical period. Based on the prediction results of the fusion prediction model, the contribution of each behavioral feature to the prediction results is calculated. Based on the contribution level, the attribution factors that led to the user's activity during the current statistical period are determined.

2. The method according to claim 1, characterized in that, The preprocessing of the behavioral feature data also includes introducing a dynamic decay function, which assigns a weight to each behavioral feature based on the occurrence time of the behavior and the importance of the feature, and the weight of the behavior decays over time.

3. The method according to claim 1, characterized in that, The preprocessing of the behavioral feature data also includes user segmentation, which divides users into different groups based on user features, and performs subsequent model training and attribution analysis for each user group.

4. The method according to claim 1, characterized in that, The noise detection and filtering includes: Transform user behavior data from the time domain to the frequency domain; Identify and filter high-frequency components above a preset threshold in the frequency domain; The filtered frequency domain data is inversely converted back to the time domain to obtain the denoised behavioral feature data.

5. The method according to claim 4, characterized in that, When converting user behavior feature data from the time domain to the frequency domain, a sliding window technique is used to segment the time series data and dynamically adjust the statistical thresholds used to detect outliers within each window.

6. The method according to claim 1, characterized in that, The abnormal user detection includes: Construct a high-dimensional feature vector that integrates device features, time series features, user attribute features, and usage scenario features; The high-dimensional feature vector is detected based on preset rules, including determining that the number of accounts logged in by the same device within a preset time period exceeds a first threshold, or the user login frequency exceeds a second threshold.

7. The method according to claim 1, characterized in that, The fusion prediction model is a fusion model of deep learning model and machine ensemble learning model.

8. The method according to claim 7, characterized in that, The deep learning model is a long short-term memory network model. When training the long short-term memory network model, different importance weights are assigned to the training samples according to the user's life cycle stage.

9. The method according to claim 8, characterized in that, When training the Long Short-Term Memory network model, a penalty factor is introduced to impose an additional loss penalty on false negative prediction errors. The penalty factor is determined based on the ratio of positive to negative samples in the training data.

10. The method according to claim 7, characterized in that, The fusion method of the deep learning model and the machine ensemble learning model is feature-level fusion, including: The deep learning model is used to extract temporal features of user behavior sequences; The temporal features are concatenated with the user's static features to form a fused feature; The fused features are input into the machine ensemble learning model for training and prediction.

11. The method according to claim 7, characterized in that, The fusion method of the deep learning model and the machine ensemble learning model is model-level fusion, including: Deep learning sub-models and machine ensemble learning sub-models are trained separately. Obtain the first prediction result of the deep learning sub-model for the target user, and the second prediction result of the machine ensemble learning sub-model for the target user; The first prediction result and the second prediction result are weighted and summed to obtain the final prediction result.

12. The method according to claim 1, characterized in that, The calculation of the contribution of each behavioral feature to the prediction result includes: An additive feature attribution method based on Shapley values ​​is used to calculate the Shapley value of each input feature for the predicted value output by the fusion prediction model, and this value is used as the contribution of each behavioral feature to the prediction result.

13. The method according to claim 12, characterized in that, The method further includes: Calculate the joint interaction contribution value between at least two features, which is used to quantify the synergistic effect of the feature combination on the prediction result.

14. The method according to claim 13, characterized in that, The determination of attribution factors leading to user activity during the current statistical period includes: The features are sorted according to their contribution and the set of important features is dynamically determined based on a preset cumulative contribution ratio threshold. From the set of important features, the feature combination with the highest joint interaction contribution value is selected as the core attribution factor.

15. The method according to claim 1, characterized in that, The method further includes: Based on the identified attribution characteristics, perform at least one of the following application operations: generate and push personalized content that matches the attribution factors, trigger intervention mechanisms for users at risk of churn, or optimize cloud drive APP functional modules related to core attribution factors.

16. The method according to claim 1, characterized in that, Before training the fusion prediction model, a feature selection step is also included, which includes: Calculate the IV value for each behavioral feature data; Features with IV values ​​below a preset predictive ability threshold are deleted.

17. The method according to claim 1, characterized in that, In the step of obtaining behavioral characteristic data of cloud disk APP users within the current statistical period, the current statistical period is the most recent 30 days.

18. The method according to claim 1, characterized in that, The preprocessing steps for the behavioral feature data further include: standardizing continuous features using the Z-score normalization method.

19. The method according to claim 1, characterized in that, In the step of training the fusion prediction model used to predict whether users will remain in the next statistical period, the K-fold cross-validation method is used to evaluate and select the model performance.

20. The method according to claim 12, characterized in that, When using the additive feature attribution method based on Shapley values, the TreeExplainer interpreter is specifically used to calculate the feature contribution of the tree model part in the fusion prediction model.

21. The method according to claim 4, characterized in that, The preset threshold is dynamically determined based on the local fraction of the frequency domain component amplitude.

22. The method according to claim 1, characterized in that, The method also includes a model deployment step: encapsulating the trained fusion prediction model and interpretive model into an application programming interface (API) for use by the background service system of the cloud disk APP; wherein, the interpretive model is used to calculate the contribution of each behavioral feature to the prediction result.

23. An attribution analysis device for the motivations of user activity in a cloud storage app, characterized in that, The device includes: The feature data acquisition module is used to acquire behavioral feature data of cloud disk APP users within the current statistical period; A preprocessing module is used to preprocess the behavioral feature data, and the preprocessing is used for noise detection and filtering, as well as abnormal user detection. The model training module is used to train a fusion prediction model based on preprocessed behavioral feature data to predict whether a user will remain in the next statistical period. The contribution calculation module is used to calculate the contribution of each behavioral feature to the prediction result based on the prediction result of the fusion prediction model through an interpretive model. The attribution factor determination module is used to determine, based on the contribution level, the attribution factors that led to the user's activity during the current statistical period.

24. An electronic device, characterized in that, include: At least one processor; Memory for storing the at least one processor-executable instruction; The at least one processor is configured to execute the instructions to implement the method as described in any one of claims 1-22.

25. A computer-readable storage medium, characterized in that, When the instructions in the computer-readable storage medium are executed by the processor of the electronic device, the electronic device is enabled to perform the method as described in any one of claims 1-22.

26. A computer program product, characterized in that, Includes a computer program that, when executed by a processor, implements the method according to any one of claims 1-22.

Citation Information

Cited By

  • E-commerce user behavior deep insight analysis method and system based on big data artificial intelligence

    CN121937154A