Explanatable employee demission prediction method and system
Through the combination of the random forest model and the SHAP interpretation model, the accuracy and interpretability of employee turnover predictions are solved, and a detailed explanation of the characteristic contribution is provided, which helps enterprises formulate effective employee retention strategies and improve the reliability and management efficiency of predictions.
Patent Information
- Application Number
- CN202510391241.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-31
- Publication Date
- 2025-07-11
AI Technical Summary
The prior art is insufficiently accurate and lacks interpretability in employee turnover forecasts, and traditional methods rely on qualitative analysis and machine learning models lack transparency, making it difficult to provide a specific explanation of feature impact.
The random forest model is used to combine the SHAP interpretation model, and a dynamic time window sampling strategy is constructed through multi-dimensional data collection, cleaning and standardization, and a feature importance map and SHAP value distribution map are combined to provide prediction and explanation of employee turnover tendencies.
Improve the accuracy and interpretability of employee turnover forecasts, and companies can identify the impact of key characteristics and formulate targeted retention measures to improve team stability and overall performance.
Smart Images

Figure CN120297490A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of human resource management, and more particularly, to an interpretable employee turnover prediction method and system. Background Art
[0002] With the rapid development of the global economy and increasing competition, attracting talents has become increasingly important for enterprises. The rising employee turnover rate not only increases the human resource cost of enterprises, but also has a negative impact on the stability of teams and the overall performance of enterprises. Therefore, accurately predicting employees' turnover intention and taking corresponding management measures in a timely manner have become key issues that enterprise managers urgently need to solve.
[0003] However, employee turnover behavior has the following special characteristics, which pose unique challenges to prediction:
[0004] ① Difficulty in observability: Turnover intention belongs to the inner activities of employees and is usually difficult to directly observe externally. Compared with other prediction tasks (such as sales prediction, equipment failure prediction, etc.), turnover intention lacks obvious external manifestations, increasing the difficulty of turnover prediction.
[0005] ② Time dependence: The relevance of turnover behavior in historical data to the current prediction time point will decay over time. For example, turnover cases two years ago may have less reference value for current prediction than recent cases.
[0006] ③ Multi-factor interactivity: Turnover decisions are usually formed by the joint action of multiple factors, rather than a single factor. There may be complex interaction relationships between these factors, increasing the complexity of prediction.
[0007] Moreover, the existing employee turnover prediction methods also have the following deficiencies:
[0008] ① Limitations of traditional methods: Traditional methods mainly rely on qualitative analysis methods such as questionnaires and employee interviews. These methods are limited by the subjectivity and finiteness of data, resulting in low accuracy and reliability of prediction.
[0009] ② Black box problem of machine learning models: Although machine learning models perform well in prediction accuracy, their "black box" characteristics lead to a lack of transparency and interpretability in the decision-making process, making it difficult to prescribe the right remedy.
[0010] ③Specific explanations for the lack of feature impact: Existing methods often can only simply give the probability of leaving the job or point out which features have a greater impact on leaving the job, lacking an explanation of how these features specifically affect employees' leaving the job. For example, although it is known that job satisfaction is an important influencing factor, it is impossible to specifically explain how the level of job satisfaction affects the probability of leaving the job, nor can it explain the specific contribution of different satisfaction levels to the leaving risk. The lack of such an explanation makes it difficult for enterprises to formulate targeted retention measures. Summary of the Invention
[0011] The present invention provides an interpretable employee leaving prediction method and system, which solves the problems of insufficient accuracy and poor interpretability in predicting employees' leaving tendency in existing enterprise human resource management.
[0012] In the first aspect, an interpretable employee leaving prediction method is provided in an embodiment of the present invention. The method includes the following processes:
[0013] Collect multi-dimensional historical data, and perform data cleaning, data standardization, and sample construction on the historical data to obtain a sample data set;
[0014] Construct a machine learning model, and perform model training on the machine learning model based on the sample data set to obtain a final machine learning model;
[0015] Construct a SHAP interpretation model, and use the SHAP interpretation model to interpret the prediction results obtained from model training to obtain the SHAP values of each feature corresponding to the prediction results, and draw a feature importance graph and a SHAP value distribution graph according to the SHAP values of each feature corresponding to the prediction results;
[0016] Use the final machine learning model to perform leaving prediction on the samples in the sample data set to obtain prediction results, and use the SHAP interpretation model to interpret the prediction results to obtain the SHAP values of each feature corresponding to the prediction results and the final prediction results.
[0017] In the above embodiment, the present invention performs employee leaving prediction and interpretation based on a machine learning model and a SHAP interpretation model, which can solve the problems of insufficient accuracy and poor interpretability in predicting employees' leaving tendency in enterprise human resource management.
[0018] As some optional implementation manners of the present application, the dimensions of the historical data include a basic attribute dimension, a work status dimension, and an external environment dimension.
[0019] In an embodiment of the present invention, the present invention comprehensively considers the basic attributes, external attributes, and current work status of employees, and conducts targeted design on the difficulties of the leaving problem, further improving the reliability and accuracy of the prediction.
[0020] In some alternative embodiments of the present application, the machine learning model uses a random forest model as the basic prediction model.
[0021] In some alternative embodiments of the present application, the core formula of the machine learning model is:
[0022]
[0023] where f(x) represents the prediction result of the random forest model for the sample x, and f t (x) represents the prediction result of the t-th decision tree for the sample x, and T represents the total number of decision trees.
[0024] In the embodiments of the present invention, the present invention uses the random forest algorithm to utilize the integrated learning advantage of multiple decision trees to achieve a robust prediction of the employee's turnover tendency.
[0025] In some alternative embodiments of the present application, the core formula of the SHAP interpretation model is:
[0026]
[0027] where f(x) represents the prediction result of the random forest model for the sample x, and φ i represents the contribution degree of the i-th feature to the prediction result, that is, the SHAP value, and x i represents the value of the i-th feature, and n represents the number of features.
[0028] In some alternative embodiments of the present application, the calculation formula of the SHAP value is:
[0029]
[0030] where φ i (x) represents the contribution degree of feature i to the prediction result of sample x, that is, the SHAP value, F represents the set of all features, S represents the feature subset that does not include feature i, and f x (S) represents the prediction result of using only the feature subset S to perform turnover prediction on the sample x, |S| represents the number of features in the feature subset S, and |F| represents the number of all features.
[0031] In some alternative embodiments of the present application, the calculation formula of the final prediction result is:
[0032]
[0033] where φ0 represents the benchmark value, that is, the average value of all sample prediction results, and φ i (x) represents the contribution of feature i to the predicted value of sample x, and n represents the number of features of sample x.
[0034] In a second aspect, the present invention provides an interpretable employee turnover prediction system, which includes:
[0035] A data preprocessing unit, which is used to collect multi-dimensional historical data, and perform data cleaning, data standardization, and sample construction on the historical data to obtain a sample data set;
[0036] A turnover prediction model unit, which is used to construct a machine learning model, and train the machine learning model based on the sample data set to obtain a final machine learning model;
[0037] A SHAP interpretation model unit, which is used to construct a SHAP interpretation model, and use the SHAP interpretation model to interpret the prediction results obtained from model training to obtain the SHAP values of each feature corresponding to the prediction results, and draw a feature importance graph and a SHAP value distribution graph according to the SHAP values of each feature corresponding to the prediction results;
[0038] A turnover prediction and interpretation unit, which uses the final machine learning model to predict the turnover of samples in the sample data set to obtain prediction results, and uses the SHAP interpretation model to interpret the prediction results to obtain the SHAP values of each feature corresponding to the prediction results and the final prediction results.
[0039] In a third aspect, the present invention provides a computer device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the interpretable employee turnover prediction method is implemented.
[0040] In a fourth aspect, the present invention provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the interpretable employee turnover prediction method is implemented.
[0041] The beneficial effects of the present invention are as follows:
[0042] 1. The present invention innovatively proposes a strategy of dynamic time-window sampling. Combining the time-window sliding technology, multiple samplings are carried out for a single employee at different time points. By treating these time-series slices as independent samples, not only is the scale of the data effectively expanded, but also the changing trend of the employee's leaving risk over time is accurately depicted. Through multiple samplings, the dynamic characteristics of the employee can be retained in detail. Compared with the traditional single-point cross-sectional prediction method, this method can more timely warn of risks. At the same time, since the number of leaving samples is usually small, by carrying out multiple samplings on in-service and leaving employees, the diversity of the samples is increased, and the model's ability to identify potential leaving patterns is enhanced. In addition, defining the problem as predicting the leaving probability of employees within a specific future period provides the enterprise with an opportunity for early intervention, which better meets the actual management needs.
[0043] 2. The present invention preferably uses the random forest algorithm, leveraging the advantage of ensemble learning of multiple decision trees to achieve a robust prediction of the employee's leaving tendency. In terms of interpretability, the SHAP value can quantitatively measure the positive and negative impacts of each feature on the leaving tendency, overcoming the deficiency of only staying at the ranking of feature importance, and providing more insightful decision-making references for the enterprise. By combining a large number of SHAP value distributions, common public risk factors can be identified, and at the same time, a personalized feature contribution map can be generated for a single employee, facilitating the formulation of differentiated retention plans.
[0044] 3. With the help of time-series sampling and SHAP analysis, the present invention can provide high-value support for enterprises in multiple aspects. For the group of employees with a rapidly increasing leaving probability, managers can promptly take intervention measures such as communication, job adjustment, or optimization of the incentive mechanism. Based on the monitoring of the leaving risk at multiple time points and SHAP analysis, the enterprise can track the effectiveness of management strategies and form a closed-loop optimization. The solution comprehensively considers the basic attributes, external attributes of employees, and the current working status, and is specifically designed for the difficulties of the leaving problem, further improving the reliability and accuracy of the prediction. BRIEF DESCRIPTION OF THE DRAWINGS
[0045] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following will briefly introduce the drawings required for use in the embodiments. It should be understood that the following drawings only show some embodiments of the present invention, and therefore should not be regarded as limiting the scope. For those of ordinary skill in the art, other related drawings can be obtained based on these drawings without creative efforts.
[0046] Figure 1 is the flowchart of the employee leaving prediction method described in the embodiments of the present invention;
[0047] Figure 2 is the structural block diagram of the employee leaving prediction system described in the embodiments of the present invention. Specific Embodiments
[0048] It should be understood that the specific embodiments described herein are merely for explaining the present application and are not intended to limit the present application.
[0049] To solve the problems of insufficient accuracy and poor interpretability in predicting employee turnover tendency in the existing enterprise human resource management. Embodiments of the present invention provide an interpretable employee turnover prediction method and system. Please refer to Figure 1 , Figure 1 which is a flowchart of the employee turnover prediction method, and the method process is as follows:
[0050] (1) Historical data preprocessing.
[0051] In the embodiments of the present invention, the historical data preprocessing process includes: collecting multi-dimensional historical data, and performing data cleaning, data standardization, and sample construction on the historical data to obtain a sample data set.
[0052] Specifically, historical data collection is the basis for employee turnover prediction, which involves obtaining employee-related data from channels such as the enterprise internal database.
[0053] Specifically, in order to comprehensively characterize the turnover risk of employees, historical data is collected from three dimensions:
[0054] ① Basic attribute dimension: including the personal basic information of employees, such as static characteristics such as age, gender, native place, education level, and marital status. These characteristics reflect the demographic characteristics and social background of employees and can explain the turnover tendency from the perspective of individual differences. For example, younger employees may be more inclined to try different career opportunities, while highly educated employees may have higher expectations for career development.
[0055] ② Work status dimension: including the current work situation of employees, such as dynamic characteristics such as department affiliation, job level, overtime hours, and performance level. These characteristics directly reflect the work engagement and development status of employees and are closely related to the turnover risk. For example, frequent overtime may lead to job burnout, and changes in performance ratings may affect employees' career confidence.
[0056] ③ External environment dimension: including the economic situation and life pressure of employees, such as environmental characteristics such as mortgage situation, car purchase situation, and work location. These characteristics reflect the external constraints and pressures faced by employees and are often important factors triggering turnover decisions. For example, a high mortgage may increase employees' need for job stability, while the convenience of the work location may affect job satisfaction.
[0057] Please refer to Table 1, which is an example of historical data in the above three dimensions:
[0058] Dimension type Continuous data Discrete data Basic attributes Age, length of service in the company Gender, education level, place of origin, marital status, type of institution Employment status Overtime hours, salary, job level Department, performance rating External environment Mortgage amount, credit card limit Car ownership, work location
[0059] Table 1 is an example of historical data in three dimensions
[0060] Specifically, through this multi-dimensional data collection, various factors affecting employee turnover can be comprehensively grasped, laying a solid data foundation for subsequent predictive analysis. The comprehensiveness and accuracy of the data directly affect the reliability of the prediction results.
[0061] In the embodiments of the present invention, due to the diverse and complex data sources, data quality problems may occur. For example, employee information may come from different departments or systems, and these data may vary in format, content, and update frequency. In addition, manually entered data is prone to errors or omissions, and the missing and inconsistent historical data will also affect the integrity and accuracy of the data. Therefore, data cleaning is a key link to ensure data quality.
[0062] In the embodiments of the present invention, data cleaning includes removing duplicate records, correcting incorrect data, filling in missing values, etc.; specifically, data cleaning includes the following steps:
[0063] ① Removing duplicate records: By comparing the key fields of the records, identify and delete duplicate data entries.
[0064] ② Correcting incorrect data: Check for logical errors and inconsistencies in the data, such as incorrect date formats, abnormal numerical values, etc., and make corrections.
[0065] ③ Filling in missing values: For missing data, fill it in reasonably according to the context information or using statistical methods.
[0066] In the embodiments of the present invention, data standardization is the key in the data preprocessing process. Its core purpose is to eliminate the dimensional differences of the data and ensure that data with different features can be effectively compared and analyzed on a unified scale. By implementing data standardization, the training effect and prediction accuracy of machine learning models can be significantly improved.
[0067] In particular, in actual business scenarios, the characteristic data of employees often have different dimensions and value ranges. For example, the value range of age is usually large, while the value of service years in the company is relatively small. Without standardization processing, these characteristics will have an adverse impact on the training of the model.
[0068] Specifically, the specific implementation steps of data standardization include:
[0069] ①Normalization processing: For continuous data, such as age, length of service in the company, salary, mortgage amount, credit card limit, etc., min-max or z-score normalization methods are adopted to scale data of different scales to the same scale, so as to eliminate the influence of data dimension or uneven data distribution and facilitate the training of machine learning models.
[0070] ②Encoding conversion: For discrete data, such as gender, department, education level, native place, marital status, car purchase situation, college category, etc., one-hot encoding is used to convert the data into numerical data to eliminate data type differences and facilitate the processing of machine learning models.
[0071] In the embodiments of the present invention, constructing samples is the key to building a machine learning model, and its purpose is to convert the preprocessed data into a standard format suitable for the training of machine learning models. Each sample consists of two parts: a feature vector and a label, and is used to train the model to identify and predict the turnover tendency of employees. Considering the sparsity characteristics of enterprise employee turnover data, the embodiments of the present invention innovatively apply a dynamic time window sampling strategy. This strategy constructs multiple sample points for the same employee at different time nodes (for example: January, March, April 2024, etc.), effectively capturing the dynamic change trajectory of the employee's state. This not only significantly improves the sensitivity to potential turnover risks, but also provides richer time series feature information for the model, thereby enhancing the prediction ability of the model.
[0072] Specifically, the feature vector, as the input part of the sample, contains all-round information of the employee at a specific reference time point, covering multiple dimensions such as basic information, work performance, and performance evaluation. These features not only reflect the personal attributes of the employee, but also reflect their development status and environmental factors in the enterprise. For example, if January 1, 2024 is taken as the reference time point, the feature vector will include complete data such as age, length of service in the company, department affiliation, educational background, geographical information, family status, etc. at this time point. Through this multi-dimensional feature construction, the model can comprehensively grasp various factors affecting the turnover tendency of employees and improve the accuracy and reliability of prediction.
[0073] Specifically, the label is the output part of the sample and is used to represent whether the employee will leave the company during the prediction window period. The setting of the prediction window period needs to comprehensively consider multiple key factors. If the window period is too short, the number of turnover samples will be insufficient, exacerbating the problem of sample imbalance. At the same time, it cannot effectively capture the gradual process of turnover, reducing the warning effect, and it is difficult for enterprises to implement effective retention measures in a timely manner. If the window period is too long, the uncertainty of prediction will increase, major changes may occur in influencing factors, the correlation between features and labels will weaken, affecting the accuracy of the model, and the timeliness of the prediction result will be insufficient, delaying the opportunity for management intervention.
[0074] Specifically, based on the above analysis, the embodiment of the present invention optimally sets the prediction window period to 6 months. This time span can not only ensure an adequate number of positive samples but also reserve a reasonable intervention space for enterprises. Specifically, if the reference time point is January 1, 2024, the label will reflect the departure situation of the employee during the period from January 1, 2024 to June 30, 2024. This setting enables the model to accurately learn the departure behavior patterns of employees within a specific time period, providing an effective prediction basis for enterprises.
[0075] In the embodiment of the present invention, through the above refined sample construction, combined with comprehensive feature information collection and scientific prediction window period setting, the embodiment of the present invention lays a solid data foundation for the subsequent machine learning model training, effectively improving the accuracy and practical value of employee departure tendency prediction.
[0076] (2) Construct and train a machine learning model.
[0077] In the embodiment of the present invention, the process of constructing and training a machine learning model includes: constructing a machine learning model and performing model training on the machine learning model based on the training set to obtain the final machine learning model.
[0078] In the embodiment of the present invention, in order to achieve accurate prediction of employee departure tendency, various machine learning models can be used, including but not limited to: random forest model, gradient boosting decision tree model, support vector machine model, neural network model.
[0079] Specifically, the embodiment of the present invention uses a random forest model as the basic prediction model. As an advanced ensemble learning algorithm, the random forest model can significantly improve the prediction accuracy and model robustness by constructing multiple decision trees and synthesizing their prediction results.
[0080] Specifically, the core formula of the random forest model is:
[0081]
[0082] where f(x) represents the prediction result of the random forest model for sample x, f t (x) represents the prediction result of the t-th decision tree for sample x, and T represents the total number of decision trees. Each decision tree is trained by randomly selecting features and samples during the training process, thereby enhancing the generalization ability of the model and adapting to different business scenarios.
[0083] In the embodiments of the present invention, in the specific implementation of model construction, the preprocessed sample dataset is first divided into a training set, a validation set, and a test set. Among them, the training set is used for model training, the validation set is used for tuning model parameters, and the test set is used to evaluate the actual performance of the model. A random forest model is constructed using the Sklearn library in Python. The hyperparameters of the model, such as the number of decision trees, the maximum depth, etc., are optimized through the Grid Search method to ensure that the model reaches the best performance. Finally, the model with the best performance on the validation set is selected as the final model to ensure its generalization ability and stability in actual applications. Specifically, the embodiments of the present invention use a variety of evaluation metrics to comprehensively evaluate the model performance, including precision, recall, F1-score, accuracy, etc., to ensure the applicability and reliability of the model in different business scenarios.
[0084] (3) Construct and apply the SHAP interpretation model.
[0085] In the embodiments of the present invention, the process of constructing and applying the SHAP interpretation model includes: constructing the SHAP interpretation model, and using the SHAP interpretation model to interpret the prediction results obtained from model training to obtain the SHAP values corresponding to the prediction results for each feature, and drawing a feature importance graph and a SHAP value distribution graph based on the SHAP values corresponding to the prediction results for each feature.
[0086] Specifically, in order to improve the interpretability and transparency of the model, the embodiments of the present invention introduce the SHAP interpretation model, which can accurately quantify the contribution degree of each feature to the prediction results of the machine learning model, thereby providing a profound analysis for the model decision-making process.
[0087] Specifically, the core formula of the SHAP interpretation model is:
[0088]
[0089] where f(x) represents the prediction result of the random forest model for the sample x, φ i represents the contribution degree of the i-th feature to the prediction result, that is, the SHAP value, x i represents the value of the i-th feature, and n represents the number of features.
[0090] Specifically, the calculation of SHAP values is based on the Shapley value in cooperative game theory, ensuring that the contribution of each feature to the prediction result is fairly allocated. There are various ways to calculate SHAP values. For example, tree-based models (such as random forests, gradient boosting machines, etc.) usually use tree-based SHAP value calculation methods, while neural network models use neural network-based SHAP value calculation methods. Given that the machine learning model constructed in the embodiments of the present invention is a random forest model, a tree-based SHAP value calculation method is adopted.
[0091] Specifically, first, use the SHAP library in Python to construct a SHAP interpretation model. Based on the optimal machine learning model constructed, interpret the samples in the training set and test set, and calculate the contribution degree of each feature to the prediction result. Then draw the feature importance graph and the SHAP value distribution graph to visually display the influence degree of each feature on the model prediction. The SHAP interpretation model provides three levels of analysis perspectives for enterprises, helping managers deeply understand the turnover risk and formulate corresponding management strategies:
[0092] (1) Feature importance analysis: The SHAP feature importance graph identifies the key features that have the most significant impact on the turnover prediction result through the analysis of all samples. This analysis helps enterprises prioritize the limited management resources on the most critical influencing factors. For example:
[0093] If the performance level is identified as a key influencing factor, the enterprise should focus on improving the performance management system: construct multi-dimensional evaluation indicators, including work results, ability performance, and development potential, etc.; establish a regular performance communication mechanism to timely discover problems and formulate improvement plans; provide training guidance or job adjustments for employees with poor performance.
[0094] If the salary level is an important influencing factor, then it is necessary to: conduct regular market salary research to ensure salary competitiveness; optimize the salary structure to balance fixed salary and performance bonus; establish a transparent salary promotion channel.
[0095] (2) Feature influence direction analysis: The SHAP value distribution graph reveals the specific influence ways of each feature on the turnover tendency:
[0096] Positive influence: A positive SHAP value indicates that this feature will increase the turnover tendency. For example, if the overtime hours show a positive influence, it means that excessive overtime will increase the turnover risk. The enterprise should: ① Optimize the work process to improve work efficiency; ② Reasonably allocate human resources to avoid long-term overtime; ③ Establish a flexible work system to balance work and life.
[0097] Negative impacts: Negative SHAP values indicate that the feature reduces the turnover tendency. For example, if the service length shows a negative impact, it means that the longer the length of service, the lower the turnover tendency. The enterprise can: ① Provide a perfect training system for newly recruited employees; ② Design a long-term incentive plan to enhance employees' sense of belonging; ③ Establish a career development path and provide a clear promotion route.
[0098] (3) Multi-level early warning analysis: The SHAP interpretation model supports multi-level analysis from individuals to groups, providing support for precise management:
[0099] Individual level: Intuitively display the specific turnover risk factors of each employee through a personalized feature contribution interpretation graph. For example: The high turnover risk of employee A may stem from low job satisfaction and limited promotion opportunities; The turnover tendency of employee B may be related to high-intensity overtime and team atmosphere. This personalized analysis enables managers to formulate targeted retention measures.
[0100] Group level: Identify the commonly existing turnover risk factors by summarizing and analyzing the SHAP value distributions of a large number of samples: Discover common problems, such as common career development bottlenecks; Identify the turnover characteristics of specific groups, such as the career aspirations of young employee groups; Predict potential turnover risk trends and deploy preventive measures in advance.
[0101] These interpretation messages not only help to understand the decision-making process of the model, but also provide strategic insights for enterprise managers, support them in making wise decisions in a complex business environment, improve the overall performance of the enterprise and employee satisfaction, so as to achieve sustainable development and the enhancement of competitive advantages.
[0102] (4) Turnover tendency prediction and interpretation.
[0103] In the embodiments of the present invention, the process of turnover tendency prediction and interpretation includes: Using the final machine learning model to perform turnover prediction on the samples in the test set to obtain a prediction result, and using the SHAP interpretation model to interpret the prediction result to obtain the SHAP values corresponding to each feature for the prediction result and the final prediction result.
[0104] In the embodiments of the present invention, turnover prediction is first performed:
[0105] Using the constructed random forest model, perform turnover prediction on the samples in the test set. Based on all the feature data of the employee before the reference time point, predict the probability of the employee having a turnover behavior within the next 6 months. This probability is calculated by averaging the prediction results of all decision trees:
[0106]
[0107] where, f t(x) represents the probability of turnover predicted by the t-th decision tree for the sample x (ranging from 0 to 1), and T is the total number of decision trees. The finally output probability value ranges from 0 to 1, and the larger the value, the higher the probability of turnover.
[0108] In the embodiment of the present invention, then the SHAP value calculation is carried out:
[0109] For each prediction result, it is necessary to calculate the SHAP value of each feature to explain its contribution to the prediction. For the feature i of the sample x, its SHAP value φ i is calculated based on the Shapley value concept in game theory:
[0110]
[0111] where φ i (x) represents the contribution degree of the feature i to the prediction result of the sample x, F represents the set of all features, S represents the feature subset that does not contain the feature i, and f x (S) represents the prediction result of predicting the turnover of the sample x only using the feature subset S, |S| represents the number of features in the feature subset S, and |F| represents the number of all features.
[0112] In the embodiment of the present invention, the prediction result of each sample can finally be expressed as:
[0113]
[0114] where φ0 represents the benchmark value, that is, the average value of all sample prediction results, and φ i (x) represents the contribution of the feature i to the predicted value of the sample x, and n represents the number of features of the sample x.
[0115] To more clearly show the purpose, technical solution and advantages of the present invention, the following takes the data of a certain enterprise as an example to show the steps of the embodiment of the present invention.
[0116] S1: Historical data preprocessing.
[0117] In the embodiment of the present invention, first, multi-dimensional data such as the basic information, work situation, and performance evaluation of employees are obtained from the enterprise. To ensure the accuracy and effectiveness of the analysis results, it is necessary to comprehensively preprocess this data. The data processing process mainly includes data cleaning, missing value filling, outlier detection and processing, etc., to ensure the integrity and consistency of the data.
[0118] Specifically, during the data processing, sample records are created in a dynamic time window manner. Specifically, sample records are created for each employee at different time points. Taking an employee as an example, in the embodiments of the present invention, different time points such as January 2024, March 2024, and April 2024 can be selected as reference time points to create corresponding sample records respectively. Feature values such as "current age" and "current length of service in the company" in each record, as well as the label of "whether leaving the company within 6 months", will be updated according to the specific reference time point. This dynamic sampling method can not only completely record the time changes of the employee status, but also expand the amount of training data, effectively improving the prediction accuracy.
[0119] Please refer to Table 2, which shows part of the data after preprocessing historical data. Among them, columns such as "Department A" and "Department B" are the results after one-hot encoding processing of department information. One-hot encoding converts categorical variables (such as departments) into multiple binary features, and each feature represents a specific department. The value of True indicates that the employee belongs to this department, and False indicates not belonging. This encoding method helps the machine learning model better process categorical data. After the above data preprocessing, the enterprise can more comprehensively master the dynamic changes of employee characteristics, laying a solid foundation for the subsequent construction of the leaving prediction model.
[0120]
[0121] Table 2 is part of the data after preprocessing historical data
[0122] S2: Construct and train a machine learning model.
[0123] In the embodiments of the present invention, the Python Sklearn library is used to construct and train a machine learning model. During the model training process, hyperparameters of the model, such as the number of decision trees and the maximum depth, are optimized to ensure the best performance of the model. Finally, the model that performs best on the validation set is selected as the final model and evaluated on the test set to ensure the generalization ability of the model. Please refer to Table 3, Table 3 shows the relevant parameters of the machine learning model:
[0124] Parameter name Parameter value Number of decision trees 100 Maximum depth 10 Minimum samples for splitting 2 Minimum samples for leaf nodes 1 Number of iterations 100 Proportion of training set 60% Proportion of validation set 20% Proportion of test set 20%
[0125] Table 3 shows the relevant parameters of the machine learning model. Specifically, please refer to Table 4, Table 4 shows the relevant indicators of the machine learning model on the validation set:
[0126] Precision Recall F1-score Sample size Not left the company 0.93 0.84 0.88 347 Left the company 0.68 0.85 0.75 143
[0127] Table 4 is the relevant indicator of the machine learning model on the validation set
[0128] S3: Construct and apply the SHAP interpretation model.
[0129] In the embodiments of the present invention, first, the SHAP library of Python is used to interpret the samples in the training set to obtain the contribution of each feature to the prediction result. Then, the feature importance graph and the SHAP value distribution graph are drawn to intuitively show the influence degree of each feature on the model prediction.
[0130] Specifically, by drawing the feature importance graph and the SHAP value distribution graph, the following key findings can be obtained:
[0131] ① Influence of age: Older employees (the older the age) are usually less likely to leave, while younger employees (the younger the age) are more likely to choose to leave. This may be because older employees often have more stable family responsibilities and career plans.
[0132] ② Influence of working years (company tenure): The longer the working time in the company (the greater the company tenure), the lower the likelihood of leaving. This indicates that the sense of belonging of employees to the company will increase with the increase of working time.
[0133] ③ Influence of overtime situation: Employees who often work overtime (the longer the average overtime duration) are actually less likely to leave. This may be because these employees are often the core backbones of the company and have better development prospects and treatment.
[0134] ④ Department differences: Employees in Department A are more likely to leave than employees in other departments, which prompts the enterprise to pay special attention to the management and working environment of this department. Although employees in Department B also have a certain tendency to leave, the degree is lighter than that of Department A.
[0135] Specifically, these findings provide a clear management direction for the enterprise. For example, the enterprise can:
[0136] ① Provide more career development opportunities and training for young employees.
[0137] ② Focus on employees with shorter working years and help them better integrate into the team.
[0138] ③ Investigate the specific problems of Department A and improve its working environment and management methods.
[0139] ④ Reasonably allocate work tasks and balance the work intensity of employees.
[0140] ⑤ Through these targeted management measures, the enterprise can more effectively reduce the employee turnover rate and improve the team stability.
[0141] S4: Prediction and interpretation of turnover intention.
[0142] Specifically, using the constructed optimal machine learning model, predict the turnover intention of the samples in the test set. Please refer to Table 5, and Table 5 shows some prediction results:
[0143]
[0144] Table 5 shows some of the prediction results
[0145] Specifically, based on the prediction results in Table 5, enterprise managers can identify which employees have a higher probability of leaving within half a year, and thus take corresponding management measures to reduce the turnover rate, improve employees' job satisfaction and the overall performance of the enterprise.
[0146] In summary, the embodiments of the present invention can effectively predict the probability of employees leaving within a specific time period and provide an in-depth analysis of the model decision-making process by constructing and applying an interpretable turnover prediction method based on the SHAP explanation model and the machine learning model. Through the ensemble learning method of the random forest model, the accuracy of the prediction and the robustness of the model are significantly improved. The introduction of the SHAP explanation model makes the decision-making process of the model transparent, and managers can clearly understand the impact of each feature on the prediction result, so as to make more informed management decisions. Information such as feature importance ranking, feature impact direction, individual prediction explanation, and group trend analysis provides strategic insights for enterprise managers, supports them in formulating more targeted employee retention strategies in a complex business environment, reduces the turnover rate, and improves employees' job satisfaction and the overall performance of the enterprise.
[0147] In addition, in one embodiment, based on the same inventive concept as the foregoing embodiment, the embodiments of the present invention provide an interpretable employee turnover prediction system. Please refer to Figure 2 , Figure 2 which is a schematic structural diagram of the employee turnover prediction system. The system corresponds one-to-one with the method of Embodiment 1. The system includes:[[]]
[0148] A data preprocessing unit, which is used to collect multi-dimensional historical data and perform data cleaning, data standardization, and sample construction on the historical data to obtain a sample data set;
[0149] A turnover prediction model unit, which is used to construct a machine learning model and perform model training on the machine learning model based on the sample data set to obtain a final machine learning model;
[0150] A SHAP explanation model unit, which is used to construct a SHAP explanation model and use the SHAP explanation model to explain the prediction results obtained from model training to obtain the SHAP values of each feature corresponding to the prediction results, and draw a feature importance graph and a SHAP value distribution graph according to the SHAP values of each feature corresponding to the prediction results;
[0151] An employee turnover prediction and explanation unit, which uses a final machine learning model to predict the turnover of samples in a sample dataset to obtain a prediction result, and uses a SHAP explanation model to explain the prediction result to obtain the SHAP value corresponding to the prediction result of each feature and the final prediction result.
[0152] It should be noted that each unit in the interpretable employee turnover prediction system in this embodiment corresponds one by one to each step in the interpretable employee turnover prediction method in the foregoing embodiment. Therefore, the specific implementation manners and achieved technical effects of this embodiment can refer to the implementation manners of the foregoing interpretable employee turnover prediction method, which will not be elaborated here.
[0153] In addition, in one embodiment, the present application further provides a computer device, which includes a processor, a memory, and a computer program stored in the memory. When the computer program is run by the processor, it implements the method in the foregoing embodiment.
[0154] In addition, in one embodiment, the present application further provides a computer storage medium, on which a computer program is stored. When the computer program is run by the processor, it implements the method in the foregoing embodiment.
[0155] In some embodiments, the computer-readable storage medium may be a memory such as FRAM, ROM, PROM, EPROM, EEPROM, flash memory, magnetic surface memory, optical disc, or CD-ROM; or it may be various devices including one or any combination of the above memories. The computer may be various computing devices including smart terminals and servers.
[0156] In some embodiments, the executable instructions may be in the form of a program, software, software module, script, or code, written in any form of programming language (including compiled or interpreted languages, or declarative or procedural languages), and may be deployed in any form, including being deployed as an independent program or being deployed as a module, component, subroutine, or other unit suitable for use in a computing environment.
[0157] As an example, the executable instructions may or may not correspond to files in the file system, may be stored as part of a file that stores other programs or data, for example, stored in one or more scripts in a HyperText Markup Language (HTML) document, stored in a single file dedicated to the program being discussed, or stored in multiple cooperating files (for example, files that store one or more modules, subroutines, or code portions).
[0158] As an example, the executable instructions may be deployed to execute on one computing device, or on multiple computing devices located at one location, or on multiple computing devices distributed across multiple locations and interconnected by a communication network.
[0159] It should be noted that, in this document, the term "including", "comprising" or any other variant thereof is intended to cover non-exclusive inclusion, such that a process, method, article or system comprising a series of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article or system. Without further limitation, an element defined by the statement "comprising an..." does not exclude the presence of additional identical elements in the process, method, article or system comprising that element.
[0160] The serial numbers of the embodiments of the present application above are only for description and do not represent the superiority or inferiority of the embodiments.
[0161] Through the description of the above embodiments, those skilled in the art can clearly understand that the above-described embodiment methods can be implemented by means of software plus a necessary general hardware platform. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation. Based on such an understanding, the technical solution of the present application, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. The computer software product is stored in a storage medium (such as a read-only memory / random access memory, magnetic disk, optical disk), and includes several instructions for causing a multimedia terminal device (which can be a mobile phone, a computer, a television receiver, or a network device, etc.) to execute the methods described in the various embodiments of the present application.
[0162] The above are only the preferred embodiments of the present application and do not limit the patent scope of the present application accordingly. Any equivalent structural or equivalent process transformation made by using the specification and drawings of the present application, or directly or indirectly applied in other related technical fields, are similarly included in the patent protection scope of the present application.
Claims
1. An interpretable employee turnover prediction method, characterized in that, The method includes the following processes: Collect historical data in multiple dimensions, and perform data cleaning, data standardization, and sample construction on the historical data to obtain a sample data set; Build a machine learning model, and perform model training on the machine learning model based on the sample data set to obtain a final machine learning model; Build a SHAP interpretation model, and use the SHAP interpretation model to interpret the prediction results obtained from model training to obtain the SHAP values corresponding to the prediction results for each feature, and draw a feature importance graph and a SHAP value distribution graph according to the SHAP values corresponding to the prediction results for each feature; Use the final machine learning model to predict employee turnover for the samples in the sample data set to obtain prediction results, and use the SHAP interpretation model to interpret the prediction results to obtain the SHAP values corresponding to the prediction results for each feature and the final corrected prediction results.
2. The interpretable employee turnover prediction method according to claim 1, wherein The dimensions of the historical data include basic attribute dimension, work status dimension, and external environment dimension.
3. An interpretable employee turnover prediction method according to claim 1, wherein The machine learning model uses a random forest model as the basic prediction model.
4. An interpretable employee turnover prediction method according to claim 3, wherein The core formula of the machine learning model is: Among them, f(x) represents the prediction result of the random forest model for the sample x, and f t (x) represents the prediction result of the t-th decision tree for the sample x, and T represents the total number of decision trees.
5. The interpretable employee turnover prediction method according to claim 1, wherein The core formula of the SHAP interpretation model is: Among them, f(x) represents the prediction result of the random forest model for the sample x, and φ i represents the contribution degree of the i-th feature to the prediction result, that is, the SHAP value, and x i represents the value of the i-th feature, and n represents the number of features.
6. The interpretable employee turnover prediction method according to claim 1, characterized in that, The calculation formula of the SHAP value is: Among them, φ i (x) represents the contribution degree of feature i to the prediction result of sample x, that is, the SHAP value, F represents the set of all features, S represents the feature subset that does not contain feature i, and f x (S) represents the prediction result of using only the feature subset S to predict the turnover of sample x, |S| represents the number of features in the feature subset S, and |F| represents the number of all features.
7. An interpretable employee turnover prediction method according to claim 6, characterized in that The calculation formula of the final corrected prediction result is: Among them, φ0 represents the reference value, that is, the average value of the prediction results of all samples, and φ i (x) represents the contribution of feature i to the predicted value of sample x, and n represents the number of features of sample x.
8. An interpretable employee turnover prediction system, characterized in that, The device includes: A data preprocessing unit, which is used to collect historical data in multiple dimensions, and perform data cleaning, data standardization, and sample construction on the historical data to obtain a sample data set; An employee turnover prediction model unit, which is used to build a machine learning model, and perform model training on the machine learning model based on the sample data set to obtain a final machine learning model; A SHAP interpretation model unit, which is used to build a SHAP interpretation model, and use the SHAP interpretation model to interpret the prediction results obtained from model training to obtain the SHAP values corresponding to the prediction results for each feature, and draw a feature importance graph and a SHAP value distribution graph according to the SHAP values corresponding to the prediction results for each feature; An employee turnover prediction and interpretation unit, which uses the final machine learning model to predict employee turnover for the samples in the sample data set to obtain prediction results, and uses the SHAP interpretation model to interpret the prediction results to obtain the SHAP values corresponding to the prediction results for each feature and the final prediction results.
9. A computer device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the computer program, it implements the interpretable employee turnover prediction method according to any one of claims 1-7.
10. A computer-readable storage medium, characterized in that, A computer program is stored on the computer-readable storage medium, and when the computer program is executed by the processor, it implements the interpretable employee turnover prediction method according to any one of claims 1-7.
Citation Information
Cited By
Programs, information processing devices, methods, and systems
JP7842428B1