Operation analysis method and system based on big data and medium
Through big data analysis methods and regression decision tree model, the problem of inaccurate customer churn prediction in the existing technology is solved, and the effect of timely intervention and recovery of users is achieved.
Patent Information
- Application Number
- CN202510820014.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-19
- Publication Date
- 2025-07-18
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Existing rules-based statistical reports or simple regression models are difficult to accurately predict the probability of customer churn and cannot formulate effective strategies to recover customers in a timely manner.
Using a big data-based operational analysis method, the enterprise's multi-dimensional feature vector is obtained, and the regression decision tree model is used to predict the churn, and intervention strategies are generated and visualized.
It improves the accuracy of customer churn prediction, can timely formulate intervention strategies to prevent user churn, and generates analysis reports that are easy for managers to view, which improves the convenience and practicality of analysis.
Smart Images

Figure CN120338870A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of data analysis, and particularly relates to an operation analysis method, system and medium based on big data. Background Art
[0002] In the field of enterprise operation, customer churn prediction is a typical high-value scenario. Research shows that the customer acquisition cost of a new customer for an enterprise is much higher than the retention cost of an old customer about to churn. Against this background, enterprises increasingly attach importance to customer churn prediction and customer retention work. How to predict in advance whether a customer has a high probability of churning, so as to take effective marketing measures and formulate reasonable marketing strategies to retain these customers has become an urgent problem for enterprises in customer management and business development.
[0003] Under the wave of the digital transformation of enterprises, the operation of enterprises gradually shifts from experience-driven to data-driven. Existing technologies analyze the churn probability through rule-based statistical reports or simple regression models. However, with the popularization of the Internet, Internet of Things and mobile terminals, the scale of data that enterprises can obtain has increased exponentially, covering multi-source heterogeneous information such as user behavior logs, transaction records, and social media interactions. Methods such as rule-based statistical reports or simple regression models are difficult to accurately predict the churn probability and cannot timely formulate corresponding strategies to retain customers. Summary of the Invention
[0004] The purpose of the present invention is to provide an operation analysis method, system and medium based on big data to solve the problem that existing methods such as rule-based statistical reports or simple regression models are difficult to accurately predict the churn probability and cannot timely formulate corresponding strategies to retain customers.
[0005] To achieve the above purpose, the present invention adopts the following technical solutions: In the first aspect, the present invention provides an operation analysis method based on big data, and the method includes: Obtain the operation data of the enterprise, where the operation data at least includes: basic attribute data, behavior data, transaction data of each user, and external data of the enterprise; Construct a multi-dimensional feature vector based on the basic attribute data, behavior data, transaction data of each user, and external data of the enterprise; Input the multi-dimensional feature vector into a pre-constructed user analysis model for churn prediction to obtain the churn probability of each user, where the user analysis model is a regression decision tree model; Generate an intervention strategy for each user based on the churn probability of each user; Generate an analysis report based on the churn probability of each user and the corresponding intervention strategy; Visually display the analysis report.
[0006] Preferably, the multi-dimensional feature vector includes: basic features and high-order features. The basic features at least include: RFM-derived features, behavior trend features, and external association features. The high-order features at least include: time series features, cross features, and Embedding features.
[0007] Preferably, an intervention strategy for each user is generated based on the churn probability of each user, including: Based on the churn probability of each user, a churn risk level for each user is constructed; Based on the churn risk level of each user, a corresponding strategy is matched from a preset strategy library as the intervention strategy for each user.
[0008] Preferably, the method further includes: constructing a user analysis model, including: Obtain sample data, preprocess the sample data to obtain a number of sample features and the corresponding feature values of each sample feature, construct a training set based on the number of sample features and the corresponding feature values of each sample feature, and use the training set as the root node of the decision tree; Use the training set as the object to be split, traverse all the sample features in the object to be split, and solve the split feature of the object to be split and the corresponding value of the split feature; Based on the split feature and the corresponding value of the split feature, divide the object to be split to obtain two new subsets, and use the two new subsets as the two child nodes of the decision tree; Use the two new subsets as the new objects to be split, and re-solve the split feature of the object to be split and the corresponding value of the split feature until a preset condition is reached to obtain a regression decision tree model.
[0009] Preferably, the preset condition is that the depth of the decision tree reaches a preset depth and / or the data volume in the divided child nodes is less than a preset quantity.
[0010] Preferably, traversing all the sample features in the object to be split and solving the split feature of the object to be split and the corresponding value of the split feature includes: Divide the object to be split with any sample feature and the corresponding feature value to obtain two divided subsets; Construct a split loss function for the object to be split; Based on the split loss function, calculate the loss values of the two divided subsets, traverse all the sample features, and obtain a loss value set; Use the sample feature corresponding to the minimum value in the loss value set and the feature value of the sample feature as the split feature of the object to be split and the corresponding value of the split feature, respectively.
[0011] Preferably, the expression of the split loss function is: ; Wherein, L is the segmentation loss function of the object to be segmented, x i is the i-th sample feature, is the label value corresponding to the i-th sample feature, R 1 is the first subset after partitioning, R 2 is the second subset after partitioning, c 1 is the first output value, c 2 is the second output value; wherein, the first output value and the second output value are determined by a machine learning algorithm.
[0012] Preferably, the machine learning algorithm includes: an input layer, a hidden layer, and an output layer; the first subset after partitioning and the second subset after partitioning are used as the input of the input layer, the output of the input layer is used as the input of the hidden layer, and the output of the hidden layer is used as the input of the output layer; the method further includes: Initializing the weights of the input layer, the biases of the hidden layer, and the input objects of the input layer, where the input objects are the first subset after partitioning and the second subset after partitioning; Based on the weights of the input layer, the biases of the hidden layer, and the input objects of the input layer, calculating the output matrix of the hidden layer of the machine learning algorithm; Performing regularization processing on the output matrix of the hidden layer to obtain a processed output matrix; Based on the processed output matrix and the expected output of the output layer of the machine learning algorithm, calculating the weights of the output layer of the machine learning algorithm; Based on the weights of the output layer and the output matrix of the hidden layer, determining the output values of the input objects, where the output values of the input objects are the first output value and the second output value.
[0013] In a second aspect, the present invention provides an operation analysis system based on big data for implementing the above-mentioned operation analysis method based on big data, and the system includes: A data acquisition module for acquiring the operation data of an enterprise, where the operation data at least includes: the basic attribute data, behavior data, and transaction data of each user, and the external data of the enterprise; A feature construction module for constructing a multi-dimensional feature vector based on the basic attribute data, behavior data, and transaction data of users and the external data of the enterprise; A churn prediction module for inputting the multi-dimensional feature vector into a pre-constructed user analysis model for churn prediction to obtain the churn probability of each user, where the user analysis model is a regression decision tree model; A strategy generation module for generating intervention strategies for each user based on the churn probability of each user; A report generation module, configured to generate an analysis report based on the churn probability of each user and the corresponding intervention strategy; A report display module, configured to visually display the analysis report.
[0014] Thirdly, the present invention provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the above-mentioned operation analysis method based on big data is implemented.
[0015] Fourthly, the present invention provides a computer-readable storage medium, on which a computer program is stored. When the program is executed by a processor, the above-mentioned operation analysis method based on big data is implemented.
[0016] Beneficial effects: 1. The user analysis module of the present invention is a regression decision tree model. The output of the regression decision tree model is a continuous value. Therefore, through the regression decision tree model, the churn probability of each user can be predicted; through the churn probability, corresponding intervention strategies can be formulated in a timely manner to prevent further user churn and be able to recover users in a timely manner; 2. The present invention generates an analysis report based on the churn probability of users and the intervention strategies generated for each user, which is convenient for management personnel to view the analysis results and improves the convenience and practicality of the analysis. Description of the Drawings
[0017] The drawings are used to provide a further understanding of the embodiments of the present invention and constitute a part of the specification. Together with the following specific embodiments, they are used to explain the embodiments of the present invention, but do not constitute a limitation to the embodiments of the present invention. In the drawings: Figure 1 is a flowchart of an operation analysis method based on big data provided by an embodiment of the present invention; Figure 2 is a block diagram of an operation analysis system based on big data provided by an embodiment of the present invention. Specific Embodiments
[0018] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the present invention in combination with the drawings and the descriptions of the embodiments or the prior art. Obviously, the following descriptions of the structures of the drawings are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts. It should be noted here that the descriptions of these embodiments are used to help understand the present invention, but do not constitute a limitation to the present invention.
[0019] Embodiment 1 Figure 1FIG. 1 is a flowchart of an operation analysis method based on big data provided by an embodiment of the present invention. Figure 1 As shown, this embodiment provides an operation analysis method based on big data, and the method includes: Step S10: Acquire the enterprise's operational data, which at least includes: basic attribute data, behavior data, and transaction data of each user and the enterprise's external data.
[0020] Among them, basic attribute data includes but is not limited to: user registration information (such as age, gender, region and registration channel), membership level (such as ordinary member, VIP and paid member).
[0021] Among them, behavioral data includes but is not limited to: APP (Application) / website activity (login frequency, page dwelling time, and number of function usage), interactive behavior (customer service consultation records, evaluation feedback, social media interaction).
[0022] Among them, transaction data includes but is not limited to: historical orders (consumption amount, frequency, and time of most recent consumption), and coupon usage (receipt rate, redemption rate).
[0023] Among them, the company's external data includes but is not limited to: competitor activity time (crawler obtains competitor promotion cycle) and macroeconomic indicators (such as industry prosperity index).
[0024] After obtaining the enterprise's operational data, it is necessary to preprocess the operational data, including but not limited to: Missing value processing: For numerical features, use the median or mean to fill, and for categorical features, use a separate mark as an "unknown" category; Outlier processing: Build business rule filtering (e.g., if the number of logins per day is greater than 100, it is considered a robot), and use IQR (interquartile range) or 3σ principle to eliminate outliers; Time window definition: Determine the "label": for example, not logging in and not making purchases for more than 30 days is considered churn; when dividing the training set and test set, split them by time window (to avoid data leakage).
[0025] Step S20: construct a multi-dimensional feature vector based on the user's basic attribute data, behavior data, transaction data and the enterprise's external data.
[0026] In this embodiment, the multi-dimensional feature vector includes: basic features and high-order features, the basic features include at least: RFM derivative features, behavioral trend features and external correlation features, and the high-order features include at least: time series features, cross features and Embedding features.
[0027] Among them, the RFM-derived features include: Recency (the number of days since the last consumption), Frequency (the number of consumptions in the past 90 days), and Monetary (the total consumption amount in the past 180 days). The behavioral trend features include: the ratio of the active days in the last 7 days to the active days in the past 30 days, and the page access depth (such as the completion rate of the steps from the home page to payment). The external correlation features include: the change in user activity during the promotion period of competing products (such as the decrease in the number of user logins during the major promotion of competing products).
[0028] Among them, the time series features include: the features statistically calculated using a sliding window (such as the average daily stay duration in the past 7 days) and the difference features (the difference between the current value and the average value of the previous week). The cross features include: the features of user level × consumption frequency (such as VIP users but no consumption in the past 30 days) and the features of coupon received but not used × decrease in unit price.
[0029] Among them, the Embedding feature is to generate embedding vectors for the user behavior sequence (such as click stream) using Word2Vec.
[0030] In this embodiment, after obtaining the above features, feature screening is required. For example: screening based on business experience to eliminate features unrelated to churn (such as user ID).
[0031] Step S30: Input the multi-dimensional feature vectors into a pre-constructed user analysis model for churn prediction to obtain the churn probability of each user. Among them, the user analysis model is a regression decision tree model; the output of the regression decision tree model in this embodiment is a continuous value. Therefore, the churn probability of each user can be predicted through the regression decision tree model.
[0032] Step S40: Generate an intervention strategy for each user based on the churn probability of each user.
[0033] Specifically, generating an intervention strategy for each user based on the churn probability of each user includes: Step S401: Based on the churn probability of each user, construct the churn risk level of each user; among them, the churn risk level is divided into: high risk, medium risk, and low risk.
[0034] The characteristic performance of high-risk users can be: last login > 7 days ago + consumption frequency decreased by 50% + coupon not used; The characteristic performance of medium-risk users can be: active but low consumption amount.
[0035] Step S402: Based on the churn risk level of each user, match the corresponding strategy from the preset strategy library as the intervention strategy for each user.
[0036] The intervention strategy of this embodiment is mainly targeted at high-risk users and medium-risk users. The strategy library is stored in the database in advance, and according to the corresponding risk level, the corresponding intervention strategy is matched from the database.
[0037] For example: For users with last login more than 7 days ago + consumption frequency decreased by 50% + no coupon used, the generated intervention strategy is: send targeted high-value coupons (such as 30 yuan off for every 100 yuan) + exclusive customer service follow-up.
[0038] Another example: For users who are active but have low consumption amount, the generated intervention strategy is: push low-price drainage products (such as 9.9 yuan trial pack).
[0039] Step S50: Generate an analysis report based on the churn probability of each user and the corresponding intervention strategy.
[0040] Step S60: Visualize the analysis report.
[0041] Therefore, through the churn probability, the present invention can timely formulate corresponding intervention strategies to prevent further user churn and be able to timely recover users; at the same time, an analysis report is generated through the churn probability of users and the intervention strategies generated for each user, which is convenient for managers to view the analysis results and improves the convenience and practicality of the analysis.
[0042] As a further optimization of this embodiment, the method further includes: constructing a user analysis model, including: Step a10: Obtain sample data, preprocess the sample data to obtain several sample features and the corresponding feature values of each sample feature, construct a training set based on the several sample features and the corresponding feature values of each sample feature, and use the training set as the root node of the decision tree; in this embodiment, corresponding labels are created for the sample data, such as historical data of users who have already churned within the enterprise, or public data sets on the Internet, which also include types such as basic user attribute data, behavior data, transaction data, and external data of the enterprise.
[0043] Step a20: Use the training set as the object to be split, traverse all sample features in the object to be split, and solve the split feature of the object to be split and the corresponding value of the split feature.
[0044] Step a30: Based on the split feature and the corresponding value of the split feature, divide the object to be split to obtain two new subsets, and use the two new subsets as the two child nodes of the decision tree.
[0045] Step a40: Take the two new subsets as new objects to be split, and re-solve the splitting features of the objects to be split and the corresponding values of the splitting features, that is, repeat steps a30 and a40 until a preset condition is reached to obtain a regression decision tree model, where the preset condition is that the depth of the decision tree reaches a preset depth and / or the amount of data in the divided child nodes is less than a preset number.
[0046] In this embodiment, the initial training set is used as the root node of the decision tree. Using the splitting method in step a30, the initial training set can be divided into two subsets, which are the child nodes generated by the decision tree. After the division operation reaches the preset condition, the last generated child node is used as the leaf node, and the leaf node is the final prediction result.
[0047] As a further optimization of this embodiment, traverse all sample features in the object to be split, and solve the splitting features of the object to be split and the corresponding values of the splitting features, including: Step a101: Divide the object to be split with any sample feature and the corresponding feature value to obtain two divided subsets; Step a102: Construct the splitting loss function of the object to be split; Step a103: Calculate the loss values of the two divided subsets based on the splitting loss function, traverse all sample features, and obtain a set of loss values; Step a104: Take the sample feature corresponding to the minimum value in the set of loss values and the feature value of the sample feature as the splitting feature of the object to be split and the corresponding value of the splitting feature, respectively.
[0048] In this embodiment, the expression of the splitting loss function is: ; In the formula, L is the splitting loss function of the object to be split, x i is the i-th sample feature, is the label value corresponding to the i-th sample feature, R 1 is the first divided subset, R 2 is the second divided subset, c 1 is the first output value, c 2 is the second output value; in this embodiment, usually the first output value and the second output value are the averages of the two subsets, that is to say, the decision tree fits a piecewise zero-order function, and these limited discrete fixed values greatly reduce the prediction accuracy of the decision tree for the churn probability.
[0049] Therefore, in order to improve the prediction accuracy of the regression decision tree model for the churn probability, the first output value and the second output value in this embodiment are determined by a machine learning algorithm.
[0050] Specifically, the machine learning algorithm includes an input layer, a hidden layer, and an output layer. The first partitioned subset and the second partitioned subset are used as the input to the input layer. The output of the input layer is used as the input to the hidden layer, and the output of the hidden layer is used as the input to the output layer.
[0051] Then, the method further includes: Step b10: Initialize the weights of the input layer, the biases of the hidden layer, and the input objects of the input layer. The input objects are the first partitioned subset and the second partitioned subset.
[0052] Step b20: Based on the weights of the input layer, the biases of the hidden layer, and the input objects of the input layer, calculate the output matrix of the hidden layer of the machine learning algorithm.
[0053] In this embodiment, the functional expression of the output matrix of the hidden layer of the machine learning algorithm is: ; ; In the formula, is the input object, is the weight of the j-th node of the input layer, is the bias of the j-th node of the hidden layer, is the output value of the j-th node of the hidden layer, is the output matrix of the hidden layer, () is the activation function. The activation function preferably uses the Sigmoid function; j = 1, 2,..., J, where J is the total number of nodes. Here, a node represents a neuron. For example, a node in the input layer represents an input neuron. In this embodiment, the total number of nodes in the input layer, the total number of nodes in the hidden layer, and the total number of nodes in the output layer are the same. The number of nodes in each layer can also be flexibly adjusted according to actual situations.
[0054] Step b40: Regularize the output matrix of the hidden layer to obtain the processed output matrix.
[0055] In this embodiment, the functional expression of the processed output matrix is: ; In the formula, is the processed output matrix, T is the matrix transpose symbol, and C is the regularization coefficient. In this embodiment, by regularizing the output matrix of the hidden layer, the generalization ability and robustness of the algorithm can be improved.
[0056] Step b50: Based on the processed output matrix and the expected output of the output layer of the machine learning algorithm, calculate the weights of the output layer of the machine learning algorithm.
[0057] In this embodiment, the functional expression of the weights of the output layer is: ; In the formula, is the weight matrix of the output layer, and Y is the expected output of the output layer.
[0058] Step b60: Based on the weights of the output layer and the output matrix of the hidden layer, determine the output values of the input object. The output values of the input object are the first output value and the second output value.
[0059] In this embodiment, the functional expression of the output values of the input object is: ; In the formula, is the output value of the input object. When the input object is R1, the output value of the input object is the first output value c1. When the input object is R2, the output value of the input object is the second output value c2.
[0060] Through the processing method of steps b10 to b60 in this embodiment, the prediction accuracy of the analysis model for the churn probability can be further improved.
[0061] Embodiment 2 Figure 2 is a block diagram of an operation analysis system based on big data provided by an embodiment of the present invention. As Figure 2 shown, this embodiment provides an operation analysis system based on big data for implementing the operation analysis method based on big data in Embodiment 1. The system includes: A data acquisition module for acquiring the operation data of the enterprise. The operation data at least includes: basic attribute data, behavior data, transaction data of each user, and external data of the enterprise; A feature construction module for constructing a multi-dimensional feature vector based on the basic attribute data, behavior data, transaction data of the user, and external data of the enterprise; A churn prediction module for inputting the multi-dimensional feature vector into a pre-constructed user analysis model for churn prediction to obtain the churn probability of each user. Among them, the user analysis model is a regression decision tree model; A strategy generation module for generating intervention strategies for each user based on the churn probability of each user; A report generation module for generating an analysis report based on the churn probability of each user and the corresponding intervention strategy; A report display module for visually displaying the analysis report.
[0062] This embodiment also provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it implements the big data-based operation analysis method in Embodiment 1.
[0063] This embodiment also provides a computer-readable storage medium, on which a computer program is stored. When the program is executed by a processor, it implements the big data-based operation analysis method in Embodiment 1.
[0064] The user analysis module of the present invention is a regression decision tree model. The output of the regression decision tree model is a continuous value. Therefore, through the regression decision tree model, the churn probability of each user can be predicted; corresponding intervention strategies can be formulated in a timely manner based on the churn probability to prevent users from further churning and be able to recover users in time; and an analysis report is generated based on the churn probability of users and the intervention strategies generated for each user, which is convenient for management personnel to view the analysis results and improves the convenience and practicality of the analysis.
[0065] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a system, or a computer program product. Therefore, the present application can adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can adopt the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk memories, CD-ROMs, optical memories, etc.) containing computer-usable program codes.
[0066] The present application is described with reference to the flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each flow and / or block in the flowcharts and / or block diagrams, as well as the combination of flows and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate a system for implementing the functions specified in Figure 1 one or more of the processes or multiple processes and / or blocks Figure 1 one or more of the blocks or multiple blocks.
[0067] The above are only the embodiments of the present application and are not used to limit the present application. For those skilled in the art, various changes and modifications can be made to the present application. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included within the scope of the claims of the present application.
Claims
1. An operation analysis method based on big data, characterized in that, The method includes: Obtaining the operation data of the enterprise, where the operation data at least includes: the basic attribute data, behavior data, transaction data of each user, and the external data of the enterprise; Constructing a multi-dimensional feature vector based on the basic attribute data, behavior data, transaction data of each user, and the external data of the enterprise; Inputting the multi-dimensional feature vector into a pre-constructed user analysis model for churn prediction to obtain the churn probability of each user, where the user analysis model is a regression decision tree model; Generating an intervention strategy for each user based on the churn probability of each user; Generating an analysis report based on the churn probability of each user and the corresponding intervention strategy; Visualizing the analysis report.
2. The operation analysis method based on big data according to claim 1, characterized in that The multi-dimensional feature vector includes: basic features and high-order features. The basic features at least include: RFM-derived features, behavior trend features, and external association features. The high-order features at least include: time series features, cross features, and Embedding features.
3. The operation analysis method based on big data according to claim 1, wherein Generating an intervention strategy for each user based on the churn probability of each user includes: Constructing a churn risk level for each user based on the churn probability of each user; Matching a corresponding strategy from a preset strategy library based on the churn risk level of each user as the intervention strategy for each user.
4. The operation analysis method based on big data according to claim 1, wherein The method further includes: constructing a user analysis model, including: Obtaining sample data, preprocessing the sample data to obtain a number of sample features and the corresponding feature values of each sample feature, constructing a training set based on the number of sample features and the corresponding feature values of each sample feature, and using the training set as the root node of the decision tree; Using the training set as the object to be split, traversing all the sample features in the object to be split, and solving the split feature of the object to be split and the corresponding value of the split feature; Dividing the object to be split based on the split feature and the corresponding value of the split feature to obtain two new subsets, and using the two new subsets as the two child nodes of the decision tree; Using the two new subsets as the new object to be split, re-solving the split feature of the object to be split and the corresponding value of the split feature until a preset condition is reached to obtain a regression decision tree model.
5. The operation analysis method based on big data according to claim 4, characterized in that, The preset condition is that the depth of the decision tree reaches a preset depth and / or the data volume in the child nodes after division is less than a preset quantity.
6. The operation analysis method based on big data according to claim 4, wherein Traversing all the sample features in the object to be split and solving the split feature of the object to be split and the corresponding value of the split feature includes: Dividing the object to be split with any sample feature and the corresponding feature value to obtain two divided subsets; Constructing a split loss function of the object to be split; Calculating the loss values of the two divided subsets based on the split loss function, traversing all the sample features to obtain a loss value set; Using the sample feature corresponding to the minimum value in the loss value set and the feature value of this sample feature as the split feature of the object to be split and the corresponding value of the split feature, respectively.
7. The operation analysis method based on big data according to claim 6, wherein The expression of the split loss function is: ; Where L is the segmentation loss function of the object to be segmented, x i is the feature of the i-th sample, is the label value corresponding to the feature of the i-th sample, R 1 is the first subset after partitioning, R 2 is the second subset after partitioning, c 1 is the first output value, c 2 is the second output value; among them, the first output value and the second output value are determined by machine learning algorithms.
8. The operation analysis method based on big data according to claim 7, wherein The machine learning algorithm includes: an input layer, a hidden layer, and an output layer; the first partitioned subset and the second partitioned subset are used as the input of the input layer, the output of the input layer is used as the input of the hidden layer, and the output of the hidden layer is used as the input of the output layer; the method further includes: Initializing the weights of the input layer, the biases of the hidden layer, and the input objects of the input layer, where the input objects are the first partitioned subset and the second partitioned subset; Calculating the output matrix of the hidden layer of the machine learning algorithm based on the weights of the input layer, the biases of the hidden layer, and the input objects of the input layer; Performing regularization processing on the output matrix of the hidden layer to obtain a processed output matrix; Calculating the weights of the output layer of the machine learning algorithm based on the processed output matrix and the expected output of the output layer of the machine learning algorithm; Determining the output values of the input objects based on the weights of the output layer and the output matrix of the hidden layer, where the output values of the input objects are the first output value and the second output value.
9. An operation analysis system based on big data, which is used to implement the operation analysis method based on big data according to any one of claims 1-8, characterized in that, The system includes: A data acquisition module for acquiring the operation data of the enterprise, where the operation data at least includes: the basic attribute data, behavior data, and transaction data of each user, and the external data of the enterprise; A feature construction module for constructing a multi-dimensional feature vector based on the basic attribute data, behavior data, and transaction data of the user and the external data of the enterprise; A churn prediction module for inputting the multi-dimensional feature vector into a pre-constructed user analysis model for churn prediction to obtain the churn probability of each user, where the user analysis model is a regression decision tree model; A strategy generation module for generating intervention strategies for each user based on the churn probability of each user; A report generation module for generating an analysis report based on the churn probability of each user and the corresponding intervention strategies; A report display module for visually displaying the analysis report.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the program is executed by a processor, it implements the big data-based operation analysis method described in any one of claims 1-8.
Citation Information
Patent Citations
User loss prediction method and system
CN106203679A
User loss prediction method and device
CN106250403A
Platform user loss prediction method
CN114022194A
Customer loss prediction method and device, computer equipment and storage medium
CN116049666A
Customer loss prediction method and device, terminal equipment and storage medium
CN116757325A