Employee performance prediction and optimization system based on big data analysis

By building a big data-based employee performance prediction system, using public and enterprise data sets to train employee performance models, and combining KL divergence and gradient descent algorithm to optimize model parameters, the problem of low accuracy of employee performance prediction in the existing technology is solved, and higher prediction accuracy and generalization ability are achieved.

CN120509778APending Publication Date: 2025-08-19HENAN PROVINCE SECOND GEOLOGICAL BRIGADE CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510589761.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-08
Publication Date
2025-08-19

AI Technical Summary

Technical Problem

The existing employee performance forecasting model has low accuracy in prediction results due to the single data type.

Method used

Using an employee performance prediction and optimization system based on big data analysis, the target employee's public data set and enterprise data set are obtained, the first employee performance model and the second employee performance model are constructed and trained, and the second employee performance model is optimized using the KL divergence and loss function, and the parameters are adjusted in combination with the gradient descent algorithm to obtain the target employee performance model.

Benefits of technology

It improves the accuracy and generalization ability of employee performance forecasts, and provides more reliable employee screening reference.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120509778A_ABST
    Figure CN120509778A_ABST
Patent Text Reader

Abstract

The invention provides an employee performance prediction and optimization system based on big data analysis, and the system comprises a data set obtaining module which is used for obtaining a target data set of a target post engaged by a target employee; the target data set comprises a public data set and an enterprise data set; the model construction module is used for obtaining a target employee performance model based on the target data set; the resume information acquisition module is used for acquiring resume information of to-be-screened employees and extracting educational background information, working age limit information, skill categories and proficiency information in the resume information; and the performance prediction module is used for inputting the educational background information, the working age limit information, the skill category and the proficiency information in the resume information into the target employee performance model to obtain the predicted performance of the to-be-screened employees. Through the scheme provided by the invention, the accuracy of the employee performance prediction result can be effectively improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of employee performance prediction, and in particular to an employee performance prediction and optimization system based on big data analysis. Background Art

[0002] With the development of computer technology, big data has become a significant force driving social progress. For businesses, the advancement of big data makes it possible to make informed decisions. For example, companies can leverage industry data for industry analysis and predict industry trends, providing valuable insights for strategic decisions. Another example is building datasets based on historical employee performance data to train performance prediction models and predict future performance.

[0003] However, the data types used in the training of existing employee performance prediction models are too single. Generally, only the historical performance data of employees are used to train employee performance models, which leads to low accuracy of employee performance prediction results. Summary of the Invention

[0004] The purpose of the present invention is to overcome the deficiencies in the prior art and provide an employee performance prediction and optimization system based on big data analysis to improve the accuracy of employee performance prediction results.

[0005] The present invention is achieved through the following technical solutions:

[0006] An employee performance prediction and optimization system based on big data analysis, including:

[0007] A data set acquisition module is configured to acquire a target data set for a target position held by a target employee; wherein each data sample in the target data set includes at least educational background information, years of experience information, target skill category and proficiency information, and performance information; the target data set includes a public data set and an enterprise data set, wherein the enterprise data set refers to a data set acquired from internal employee data of an enterprise;

[0008] A model building module, configured to obtain a target employee performance model based on the target data set;

[0009] The resume information acquisition module is used to obtain the resume information of the employees to be screened and extract the educational background information, working years information, skill categories and proficiency information in the resume information;

[0010] The performance prediction module is used to input the educational information, working years information, skill category and proficiency information in the resume information into the target employee performance model to obtain the predicted performance of the employee to be screened.

[0011] In one embodiment, obtaining a target employee performance model based on the target data set includes:

[0012] Using the public dataset in the target dataset, constructing and training a first employee performance model;

[0013] Using the enterprise dataset in the target dataset, constructing and training a second employee performance model;

[0014] Determining a first KL divergence from the first employee performance model to the second employee performance model, and a second KL divergence from the second employee performance model to the first employee performance model;

[0015] Determining a loss function of the second employee performance model based on the first KL divergence and the second KL divergence;

[0016] Based on the loss function, the parameters of the second employee performance model are adjusted to obtain a target employee performance model.

[0017] In one embodiment, determining a first KL divergence from the first employee performance model to the second employee performance model includes:

[0018] Inputting each data sample in the enterprise data set in the target data set into the first employee performance model and the second employee performance model respectively, to obtain a first output distribution corresponding to the first employee performance model and a second output distribution corresponding to the second employee performance model;

[0019] Calculating a first KL divergence from the first employee performance model to the second employee performance model using a KL divergence formula based on the first output distribution and the second output distribution;

[0020] The first KL divergence calculation formula is:

[0021]

[0022] in, represents the first KL divergence, P(x) represents the first output distribution, Q(x) represents the second output distribution, and x represents the xth data sample in the enterprise data set.

[0023] In one embodiment, determining the loss function of the second employee performance model based on the first KL divergence and the second KL divergence includes:

[0024] Obtaining the salary change of the target position from the public dataset;

[0025] Based on the magnitude relationship between the salary change and a preset change threshold, determining a magnitude relationship between a first weight corresponding to the first KL divergence and a second weight corresponding to the second KL divergence; the sum of the first weight and the second weight is 1;

[0026] Determine the first weight and the second weight using a random algorithm based on a magnitude relationship and a sum of the first weight and the second weight;

[0027] Based on the first weight and the second weight, a weighted sum is performed on the first KL divergence and the second KL divergence to obtain a loss function of the second employee performance model.

[0028] In one embodiment, determining the relationship between a first weight corresponding to the first KL divergence and a second weight corresponding to the second KL divergence based on the relationship between the salary change and the preset change threshold includes:

[0029] If the salary change is greater than a preset change threshold, it is determined that the first weight is greater than the second weight; or, if the salary change is not greater than the preset change threshold, it is determined that the first weight is not greater than the second weight.

[0030] In one embodiment, adjusting the parameters of the second employee performance model based on the loss function to obtain the target employee performance model includes:

[0031] Based on the loss function, adjusting the parameters of the second employee performance model using a gradient descent algorithm;

[0032] The second employee performance model that minimizes the loss function is determined as the target employee performance model.

[0033] In one embodiment, when obtaining the public dataset in the target dataset, obtaining the target dataset of the target position that the target employee is engaged in includes:

[0034] Crawling an initial public data set associated with the target position, where the data sample format of the initial public data set is a text format;

[0035] Using natural language processing (NPL) technology to process each data sample in the initial public dataset, obtaining educational information, years of work experience information, at least one skill category and proficiency information, and performance information of each data sample in the initial public dataset;

[0036] Based on the at least one skill category and proficiency information of each data sample in the initial public data set, a correlation analysis method or a feature importance analysis method is used to filter out the target skill category and proficiency information associated with the target position from the at least one skill category and proficiency information.

[0037] In one embodiment, obtaining a target dataset of a target position held by a target employee includes:

[0038] Obtain the educational background information, years of work experience, target skill category and proficiency information, and performance information of each sample in the public dataset, and the educational background information, years of work experience, target skill category and proficiency information, and performance information of each sample in the enterprise dataset;

[0039] If the first sample in the public data set does not contain the performance information, the mapping relationship between the educational information, years of work experience information, target skill category and proficiency information and the performance information in the enterprise data set is used to obtain the performance information of the first sample based on the educational information, years of work experience information, target skill category and proficiency information in the first sample.

[0040] In one embodiment, the first employee performance model is a classification feature enhanced CatBoost model, and the second employee performance model is a natural gradient boosting NGBoost model.

[0041] In one embodiment, the system further includes a judgment module;

[0042] The judgment module is used to judge whether the employee to be screened meets the recruitment conditions based on the predicted performance of the employee to be screened and a preset performance threshold.

[0043] The present invention provides an employee performance prediction and optimization system based on big data analysis, comprising: a data set acquisition module, used to acquire a target data set of a target position engaged in by a target employee; wherein each data sample in the target data set includes at least educational background information, years of work experience information, target skill category and proficiency information and performance information, and the target data set includes a public data set and an enterprise data set, and the enterprise data set refers to a data set acquired from employee data within the enterprise; a model construction module, used to obtain a target employee performance model based on the target data set; a resume information acquisition module, used to acquire resume information of an employee to be screened, and extract educational background information, years of work experience information, skill category and proficiency information from the resume information; and a performance prediction module, used to input the educational background information, years of work experience information, skill category and proficiency information from the resume information into the target employee performance model to obtain the predicted performance of the employee to be screened.

[0044] The technical solution provided by the present invention trains a target employee performance model by using a public data set and an enterprise data set containing educational information, years of work experience, target skill categories and proficiency information, and performance information that are strongly correlated with employee performance. As a result, the trained target performance model can not only achieve high accuracy in employee performance prediction within the enterprise; at the same time, combined with the public data set, it effectively improves the generalization ability of the target employee performance model, further effectively improves the accuracy of target employee performance prediction, and also provides a reliable reference for the screening of employees to be screened. BRIEF DESCRIPTION OF THE DRAWINGS

[0045] Figure 1 This is a schematic diagram of the structure of the employee performance prediction and optimization system based on big data analysis provided by the present invention;

[0046] Figure 2 A flow chart of the method for obtaining a target employee performance model provided by the present invention;

[0047] Figure 3 A schematic diagram comparing the performance prediction results of the first employee performance model and the target employee performance model provided by the present invention;

[0048] Figure 4 A schematic diagram comparing the performance prediction results of the second employee performance model and the target employee performance model provided by the present invention;

[0049] Figure 5 A schematic diagram comparing different evaluation indicators of different models provided by the present invention;

[0050] Figure 6 A flowchart of a method for determining the loss function of the second employee performance model provided by the present invention;

[0051] Figure 7 A flowchart of a method for obtaining a target data set for a target position performed by a target employee provided by the present invention. DETAILED DESCRIPTION

[0052] The technical solutions in the embodiments of the present application will be described clearly and completely below in conjunction with the drawings in the embodiments of the present application.

[0053] It should be understood that some embodiments described herein are merely used to explain the technical solutions of the present application and are not intended to limit the technical scope of the present application.

[0054] like Figure 1 As shown, the present invention provides an employee performance prediction and optimization system 100 based on big data analysis, comprising:

[0055] The data set acquisition module 101 is configured to acquire a target data set for a target position held by a target employee; wherein each data sample in the target data set includes at least educational background information, years of experience information, target skill category and proficiency information, and performance information. The target data set includes a public data set and an enterprise data set, wherein the enterprise data set refers to a data set acquired from internal employee data of an enterprise;

[0056] In one embodiment, the target employee refers to an employee who is currently engaged in or has been engaged in the target position; wherein, the target position can be any position provided by various enterprises, institutions, universities, etc., such as data analyst, product manager, counselor, etc., and the present invention is not limited to this.

[0057] In one embodiment, text recognition technology, such as natural language processing (NLP) technology and optical character recognition (OCR) technology, can be used to obtain educational information, years of work experience, target skill category and proficiency information, and performance information of each data sample in the target data set.

[0058] In one embodiment, a public dataset refers to a dataset crawled from online big data, such as a dataset consisting of data samples containing educational information, years of experience, target skill categories, proficiency information, and performance information publicly available on public web pages (such as Zhiyouji). It can also be a dataset consisting of data samples containing educational information, years of experience, target skill categories, proficiency information, and performance information obtained and organized by enterprise personnel from third-party recruitment websites, or a combination of the above two or more datasets. In one embodiment, an enterprise dataset refers to a dataset consisting of data samples containing educational information, years of experience, target skill categories, proficiency information, and performance information obtained from internal employee data by an enterprise with employee performance prediction needs, or by a third party cooperating with the enterprise.

[0059] A model building module 102 is used to obtain a target employee performance model based on the target data set;

[0060] In one embodiment, the target employee performance model may adopt a neural network model, a machine learning model, etc. Using a data set to train the model is an existing mature technology and will not be described in detail here.

[0061] The resume information acquisition module 103 is used to obtain the resume information of the employee to be screened and extract the educational background information, working years information, skill category and proficiency information from the resume information;

[0062] In one embodiment, similar to the data set acquisition module 101, the educational background information, working experience information, skill category and proficiency information in the resume information may also be extracted using NLP technology, OCR technology, etc.

[0063] The performance prediction module 104 is used to input the educational background information, working experience information, skill category and proficiency information in the resume information into the target employee performance model to obtain the predicted performance of the employee to be screened.

[0064] The technical solution provided by the present invention trains a target employee performance model by using a public data set and an enterprise data set containing educational information, years of work experience, target skill categories and proficiency information, and performance information that are strongly correlated with employee performance. As a result, the trained target performance model can not only achieve high accuracy in employee performance prediction within the enterprise; at the same time, combined with the public data set, it effectively improves the generalization ability of the target employee performance model, further effectively improves the accuracy of target employee performance prediction, and also provides a reliable reference for the screening of employees to be screened.

[0065] In one embodiment, since public data sets and enterprise data sets usually have certain differences, such as the educational information, years of work information, skill categories and proficiency information, and performance information in the data samples of the public data sets are usually obtained through data statistics, such as taking the educational information, years of work information, skill categories and proficiency information within a certain interval, and then using the median or average to obtain the performance information within this interval. Therefore, compared with the enterprise data set, the generalization ability of the model trained using the public data set is usually better than the model trained using the enterprise data set. In order to obtain better generalization ability, the employee performance prediction and optimization system 100 based on big data analysis provided by the present invention will use the public data set and the enterprise data set to train the first employee performance prediction model and the second employee performance prediction model respectively, and then combine the first employee performance model and the second employee performance model to make the final target employee performance prediction model have stronger generalization ability and higher performance prediction accuracy.

[0066] Specifically, such as Figure 2 As shown, the target employee performance model is obtained based on the target data set, including:

[0067] Step 201: construct and train a first employee performance model using the public dataset in the target dataset;

[0068] In one embodiment, the first employee performance model is a CatBoost model, a machine learning model that enhances classification features. CatBoost is based on the Gradient Boosting Decision Tree (GBDT) algorithm, which iteratively trains weak learners (decision trees) to build a strong model. The core concept is to gradually correct the prediction errors of the previous model, thereby achieving higher prediction accuracy.

[0069] Step 202: construct and train a second employee performance model using the enterprise dataset in the target dataset;

[0070] In one embodiment, the second employee performance model is a natural gradient boosting (NGBoost) model in a machine learning model.

[0071] Step 203: determining a first KL divergence from the first employee performance model to the second employee performance model, and a second KL divergence from the second employee performance model to the first employee performance model;

[0072] In one embodiment, for the trained first employee performance model and the second employee performance model, the performance prediction results obtained for the same data sample are different; therefore, for the same input, the outputs of the first employee performance model and the second employee performance model can be regarded as two discrete probability distributions, and therefore the first KL divergence from the first employee performance model to the second employee performance model and the second KL divergence from the second employee performance model to the first employee performance model can be calculated.

[0073] Specifically, the same data samples mentioned above can be data samples in a public data set or data samples in an enterprise data set.

[0074] Taking the same data sample as the data sample in the target data set as an example: determining the first KL divergence from the first employee performance model to the second employee performance model includes:

[0075] Inputting each data sample in the enterprise data set in the target data set into the first employee performance model and the second employee performance model respectively, to obtain a first output distribution corresponding to the first employee performance model and a second output distribution corresponding to the second employee performance model;

[0076] Calculating a first KL divergence from the first employee performance model to the second employee performance model using a KL divergence formula based on the first output distribution and the second output distribution;

[0077] The first KL divergence calculation formula is:

[0078]

[0079] in, represents the first KL divergence, P(x) represents the first output distribution, Q(x) represents the second output distribution, and x represents the xth data sample in the enterprise data set.

[0080] Taking the same data sample as the data sample in the target data set as an example: determining the second KL divergence from the second employee performance model to the first employee performance model includes:

[0081] Inputting each data sample in the enterprise data set in the target data set into the first employee performance model and the second employee performance model respectively, to obtain a first output distribution corresponding to the first employee performance model and a second output distribution corresponding to the second employee performance model;

[0082] Calculating a second KL divergence from the second employee performance model to the first employee performance model using a KL divergence formula based on the first output distribution and the second output distribution;

[0083] The second KL divergence calculation formula is:

[0084]

[0085] Step 204: determining a loss function of the second employee performance model based on the first KL divergence and the second KL divergence;

[0086] In one embodiment, weights of the first KL divergence and the second KL divergence may be preset, and a weighted sum of the first KL divergence and the second KL divergence may be determined as the loss function of the second employee performance model.

[0087] In one embodiment, the sum of the first KL divergence and the second KL divergence, or the difference between the first KL divergence and the second KL divergence, may be directly determined as the loss function of the second employee performance model.

[0088] Step 205: Based on the loss function, adjust the parameters of the second employee performance model to obtain a target employee performance model.

[0089] Specifically, adjusting the parameters of the second employee performance model based on the loss function to obtain the target employee performance model includes:

[0090] Based on the loss function, adjusting the parameters of the second employee performance model using a gradient descent algorithm;

[0091] The second employee performance model that minimizes the loss function is determined as the target employee performance model.

[0092] Take the data analyst position at a certain company as an example: When training the performance model of the target employee, the company selects a public dataset and a corporate dataset. The public dataset contains 10,000 samples, each of which includes the target employee's education information, years of work experience, Python and SQL proficiency information, and performance information. Due to the smaller data volume, the corporate dataset contains 2,000 samples, each of which includes the target employee's education information, years of work experience, Python and SQL proficiency information, and performance information.

[0093] A company used a public dataset containing 10,000 samples to train the first employee performance model, and used a company dataset containing 10,000 samples to train the second employee performance model. Before model training, the public and company datasets were standardized, duplicate values were deleted, and missing values were filled. For example, Z-score was used for standardization and the mean was used for missing value filling.

[0094] After preprocessing the data samples in the public dataset and the enterprise dataset, the samples contained in the public dataset and the enterprise dataset are divided according to the preset ratios of the test set, validation set, and training set, as shown in Table 1. Table 1 is a schematic table of the ratios and sample sizes of the training set, validation set, and test set.

[0095] Dataset type Training set (70%) Validation set (15%) Test set (15%) Public datasets 7000 1500 1500 Enterprise Datasets 1400 300 3000

[0096] When training the first employee performance model and the second employee performance model, the corresponding training set can be used to determine the model parameters, and then the corresponding validation set can be used to adjust the model parameters. Finally, the corresponding test set can be used to test the performance of the model.

[0097] After the first employee performance model and the second employee performance model are trained, 300 to 5000 target samples can be randomly selected from the enterprise sample set as input, and the trained first employee performance model and the second employee performance model can be input respectively to obtain the probability distribution of the target samples in the first employee performance model and the second employee performance model; then, the first KL divergence and the second KL divergence are calculated based on the probability distribution of the target samples in the first employee performance model and the second employee performance model, and the loss function constructed by the weighted sum of the first KL divergence and the second KL divergence is used to constrain the second employee target performance model; that is, the loss function constructed by the weighted sum of the first KL divergence and the second KL divergence is used as the loss function, and the second employee performance model is retrained using the training set, validation set and test set divided in the enterprise data set to obtain the target employee performance model.

[0098] In order to evaluate the accuracy of the prediction results of the first employee performance model, the second employee performance model and the target employee performance model for employee performance prediction, 200 samples can be selected from at least one set of public data sets or enterprise data sets as test samples to test the accuracy of the performance prediction results of each model. In this example, 100 data samples are selected from the public data set and the enterprise data set as test samples, where the performance prediction results of the first employee performance model and the second employee performance model are compared with the performance prediction results of the target employee performance model, respectively. Figure 3 and Figure 4 As shown. Figure 3 and Figure 4 It can be seen that the predicted value of the target employee performance model is closer to the actual performance value than the first employee performance model and the second employee performance model. In particular, the maximum deviation between the predicted value and the actual performance value of the first employee performance model and the second employee performance model is significantly greater than that of the target employee performance model. It should be noted that in order to clearly reflect the comparison between the first employee performance model or the second employee performance model and the target employee performance model, Figure 3 and Figure 4 In the figure, only 100 examples of the 200 test samples are given to avoid unclear drawings due to too many test samples.

[0099] In order to further evaluate the model performance of the first employee performance model, the second employee performance model and the target employee performance model, the above 200 test samples were also selected, and the educational information, working experience information, Python and SQL proficiency information in the test samples were used as input to the first employee performance model, the second employee performance model and the target employee performance model respectively, and the performance prediction values of the first employee performance model, the second employee performance model and the target employee performance model were obtained; then the performance prediction values corresponding to each model were compared with the actual performance values to obtain the mean square error (MES), mean absolute error (MAE) and R that reflect the model prediction effect. 2 Evaluation indicators such as scores and error standard deviation indicators that reflect the stability of the model prediction effect.

[0100] As shown in Table 2 below, Table 2 is a performance comparison table of the trained first employee performance model, second employee performance model, and target employee performance model on 200 test samples.

[0101] Model MSE MAE <![CDATA[R 2 ]]> Error standard deviation First Employee Performance Model 12.5 3.1 0.85 1.2 Second employee performance model 8.2 2.4 0.72 0.9 Target employee performance model 6.8 2.0 0.91 0.6

[0102] In order to more intuitively represent the performance differences of the trained first employee performance model, the second employee performance model, and the target employee performance model on 200 test samples, you can refer to Figure 5 Schematic diagram of the comparison of different evaluation indicators of different models shown.

[0103] As shown in Table 2 and Figure 5 As shown, compared with the first employee performance model trained using the public dataset, the MSE of the target employee performance model is reduced by 45.6%, and compared with the second employee performance model trained using the enterprise dataset, the MSE of the target employee performance model is reduced by 17%. That is, compared with the first employee performance model and the second employee performance model, the difference between the predicted value and the true value of the employee performance of the target employee performance model is smaller. Furthermore, compared with the first employee performance model trained using the public dataset, the MAE of the target employee performance model is reduced by 35.4%, and compared with the second employee performance model trained using the enterprise dataset, the MSE of the target employee performance model is reduced by 16.7%. That is, compared with the first employee performance model and the second employee performance model, the difference between the predicted value and the true value of the employee performance of the target employee performance model is smaller. Furthermore, it can be seen from the above table that the R 2 The R of the second employee performance model is 0.85. 2 The R of the target employee performance model is 0.72. 2 is 0.91. Obviously, the R 2 Both are less than 0.9, so the fit of the first employee performance model and the second employee performance model is insufficient, and because the second employee performance model is trained with a smaller enterprise data set, the fit of the second employee performance model is even worse. However, using the target employee performance model optimized by KL divergence, its R 2The accuracy of the first and second KL divergences is 0.91, indicating a better fit and more accurate prediction of employee performance. Furthermore, whether it is the first employee performance model, the second employee performance model, or the target employee performance model constrained by the loss function constructed from the weighted sum of the first and second KL divergences, they all represent the relationship between the independent variables (the target employee's educational background, years of experience, target skill categories, and proficiency) and the dependent variable (performance information). Therefore, although the target employee performance model is based on the second employee performance model, it can also effectively learn the relationship between independent and dependent variables in the public dataset, effectively enhancing the generalization ability of the target employee performance model. Furthermore, since the standard deviation of the error of the first employee performance model is 1.2, the standard deviation of the error of the second employee performance model is 0.9, and the standard deviation of the error of the target employee performance model is 0.6, that is, the first employee performance model trained using the public data set is greatly affected by industry fluctuations and data authenticity because its data source is the public data set, and therefore the model has poor robustness; the second employee performance model trained using the enterprise data set is more stable and true and accurate because its data source is the enterprise data set, and therefore the model has higher robustness than the first employee performance model; however, the target employee performance model has better robustness and stronger model stability, which is more conducive to accurately predicting employee performance.

[0104] It can be seen that, through the method provided in steps 201 to 205 based on the target data set, a target employee performance model is obtained, the first target employee performance model is trained by the public data set, and the second target employee performance model is trained by the enterprise data set; wherein, the first target employee performance model has good performance on the public data set, and the second target employee performance model has good performance on the enterprise data set; further, by using the loss function constructed by the first KL divergence and the second KL divergence to optimize the second target employee performance model that has good performance on the enterprise data set, the generalization ability of the target employee performance model obtained is significantly enhanced, the target employee performance model is used to adapt to the characteristics of the enterprise itself, and it can also be combined with the industry-wide data in the public data set to more objectively and accurately evaluate the performance of employees in a specific enterprise. Wherein, a specific enterprise refers to an enterprise that uses the employee performance prediction and optimization system 100 based on big data analysis provided by the present invention to predict employee performance.

[0105] In one embodiment, for a specific enterprise, changes in the industry in which it is located are generally strongly positively correlated with performance, for example, a good industry development trend indicates high performance, and a poor industry development trend indicates low performance. The industry development trend is most obviously reflected in the employee's salary data. Therefore, the salary change of the target industry in which the target employee is engaged can be used to adjust the weights of the first KL divergence and the second KL divergence constituting the loss function. A large salary change indicates that the industry is unstable. In this case, the proportion of the enterprise data set representing the specific enterprise should be appropriately reduced, that is, the weight corresponding to the second KL divergence should be appropriately reduced. Conversely, a small salary change indicates that the industry is stable. In this case, the proportion of the enterprise data set representing the specific enterprise should be appropriately increased, that is, the weight corresponding to the second KL divergence should be appropriately increased.

[0106] Based on this, in one embodiment, if Figure 6 As shown, determining the loss function of the second employee performance model based on the first KL divergence and the second KL divergence includes:

[0107] Step 301: Obtain the salary change of the target position from the public dataset;

[0108] In one embodiment, the public dataset can more objectively reflect industry changes compared to the enterprise dataset.

[0109] Step 302: Based on the relationship between the salary change and a preset change threshold, determine the relationship between a first weight corresponding to the first KL divergence and a second weight corresponding to the second KL divergence; the sum of the first weight and the second weight is 1.

[0110] Specifically, if the salary change is greater than a preset change threshold, it is determined that the first weight is greater than the second weight; or if the salary change is not greater than the preset change threshold, it is determined that the first weight is not greater than the second weight.

[0111] For example, if the salary change is greater than a preset change threshold, the first weight can be set to any value greater than 0.5, and the second weight can be set to any value less than 0.5. Conversely, if the salary change is not greater than the preset change threshold, the first weight can be set to any value less than 0.5, and the second weight can be set to any value greater than 0.5.

[0112] In one embodiment, the salary change can be an annual salary change, a quarterly salary change, or a salary change at fixed annual intervals, but the present invention is not limited thereto. Specifically, public datasets typically contain data from multiple consecutive years or quarters, and the corresponding salary change can be obtained by using the difference between the data.

[0113] Step 303: Determine the first weight and the second weight using a random algorithm based on the magnitude relationship and sum of the first weight and the second weight;

[0114] Step 304 : Based on the first weight and the second weight, perform weighted summation on the first KL divergence and the second KL divergence to obtain a loss function of the second employee performance model.

[0115] In one embodiment, any company in a given industry typically has multiple positions, each with different skill requirements. However, public datasets may contain statistical data for a particular industry that includes multiple skill categories and proficiency levels. For example, the skill requirements for a data analyst position are typically Python and SQL, but public datasets may also include Java, C, and other skill categories and proficiency levels. Therefore, to obtain a more accurate dataset, it is necessary to filter out skill categories and proficiency information that are more relevant to the target position from data samples in the public dataset.

[0116] Based on this, in one embodiment, if Figure 7 As shown, when obtaining the public data set in the target data set, obtaining the target data set of the target position that the target employee is engaged in includes:

[0117] Step 401: crawling an initial public data set associated with the target position, where the data sample format of the initial public data set is text format;

[0118] In one embodiment, the initial public data set refers to a data set crawled from network big data, such as a data set composed of data samples containing educational information, years of work experience information, target skill categories and proficiency information and performance information disclosed on public web pages (such as Zhiyouji, etc.); it can also be a data set composed of data samples containing educational information, years of work experience information, target skill categories and proficiency information and performance information obtained and organized by corporate personnel from third-party recruitment websites; it can also be a collection of the above two or more data sets.

[0119] Step 402: Process each data sample in the initial public dataset using natural language processing (NPL) technology to obtain educational background information, years of work experience information, at least one skill category and proficiency information, and performance information of each data sample in the initial public dataset;

[0120] Step 403: Based on the at least one skill category and proficiency information of each data sample in the initial public data set, a correlation analysis method or a feature importance analysis method is used to filter out the target skill category and proficiency information associated with the target position from the at least one skill category and proficiency information.

[0121] In one embodiment, obtaining a target dataset of a target position held by a target employee includes:

[0122] Obtain the educational background information, years of work experience, target skill category and proficiency information, and performance information of each sample in the public dataset, and the educational background information, years of work experience, target skill category and proficiency information, and performance information of each sample in the enterprise dataset;

[0123] If the first sample in the public data set does not contain the performance information, the mapping relationship between the educational information, years of work experience information, target skill category and proficiency information and the performance information in the enterprise data set is used to obtain the performance information of the first sample based on the educational information, years of work experience information, target skill category and proficiency information in the first sample.

[0124] In one embodiment, the mapping relationship can be pre-stored in the employee performance prediction and optimization system 100 based on big data analysis in the form of a preset table. For example, the first column of the table contains educational information indicating the highest educational level, the second column contains years of work experience information indicating years of work experience, the third column contains target skill categories and proficiency information indicating target skill categories and proficiency, and the fourth column contains performance information indicating performance. If the first sample in the public dataset does not contain the performance information, it is only necessary to traverse the table and filter out rows from the table that have the same educational information, years of work experience information, target skill categories, and proficiency information as the first sample. The fourth column of the same row will be the performance information of the first sample.

[0125] In one embodiment, the mapping relationship can also be a random forest model constructed based on the random forest algorithm. The performance information of the first sample can be obtained by simply inputting the educational information, working years information, target skill category and proficiency information of the first sample into the random forest model.

[0126] In one embodiment, since the employee performance prediction and optimization system 100 based on big data analysis provided by the present invention can be used to predict the predicted performance of employees to be screened, the employee performance prediction and optimization system 100 based on big data analysis provided by the present invention can also determine whether the employees to be screened meet the recruitment conditions based on the predicted performance of the employees to be screened. Based on this, the employee performance prediction and optimization system 100 based on big data analysis provided by the present invention further includes a judgment module;

[0127] The judgment module is used to judge whether the employee to be screened meets the recruitment conditions based on the predicted performance of the employee to be screened and a preset performance threshold.

[0128] It should be noted that the above description is merely a specific embodiment of the present application, but the scope of protection of the present application is not limited thereto. Any changes or substitutions that can be easily conceived by a person skilled in the art within the technical scope disclosed in the present application should be included in the scope of protection of the present application. The scope of protection of the present application should be based on the scope of protection of the claims.

Claims

1. An employee performance prediction and optimization system based on big data analysis, characterized in that: include: A data set acquisition module is configured to acquire a target data set for a target position held by a target employee; wherein each data sample in the target data set includes at least educational background information, years of experience information, target skill category and proficiency information, and performance information; the target data set includes a public data set and an enterprise data set, wherein the enterprise data set refers to a data set acquired from internal employee data of an enterprise; A model building module, configured to obtain a target employee performance model based on the target data set; The resume information acquisition module is used to obtain the resume information of the employees to be screened and extract the educational background information, working years information, skill categories and proficiency information in the resume information; The performance prediction module is used to input the educational information, working years information, skill category and proficiency information in the resume information into the target employee performance model to obtain the predicted performance of the employee to be screened.

2. The system according to claim 1, wherein: The step of obtaining a target employee performance model based on the target data set includes: Using the public dataset in the target dataset, constructing and training a first employee performance model; Using the enterprise dataset in the target dataset, constructing and training a second employee performance model; Determining a first KL divergence from the first employee performance model to the second employee performance model, and a second KL divergence from the second employee performance model to the first employee performance model; Determining a loss function of the second employee performance model based on the first KL divergence and the second KL divergence; Based on the loss function, the parameters of the second employee performance model are adjusted to obtain a target employee performance model.

3. The system according to claim 2, characterized in that Determining a first KL divergence from the first employee performance model to the second employee performance model includes: Inputting each data sample in the enterprise data set in the target data set into the first employee performance model and the second employee performance model respectively, to obtain a first output distribution corresponding to the first employee performance model and a second output distribution corresponding to the second employee performance model; Calculating a first KL divergence from the first employee performance model to the second employee performance model using a KL divergence formula based on the first output distribution and the second output distribution; The first KL divergence calculation formula is: in, represents the first KL divergence, P(x) represents the first output distribution, Q(x) represents the second output distribution, and x represents the xth data sample in the enterprise data set.

4. The system according to claim 2, wherein: The determining of the loss function of the second employee performance model based on the first KL divergence and the second KL divergence includes: Obtaining the salary change of the target position from the public dataset; Based on the magnitude relationship between the salary change and a preset change threshold, determining a magnitude relationship between a first weight corresponding to the first KL divergence and a second weight corresponding to the second KL divergence; the sum of the first weight and the second weight is 1; Determine the first weight and the second weight using a random algorithm based on a magnitude relationship and a sum of the first weight and the second weight; Based on the first weight and the second weight, a weighted sum is performed on the first KL divergence and the second KL divergence to obtain a loss function of the second employee performance model.

5. The system according to claim 4, characterized in that The determining, based on the magnitude relationship between the salary change and the preset change threshold, a magnitude relationship between a first weight corresponding to the first KL divergence and a second weight corresponding to the second KL divergence includes: If the salary change is greater than a preset change threshold, it is determined that the first weight is greater than the second weight; or, if the salary change is not greater than the preset change threshold, it is determined that the first weight is not greater than the second weight.

6. The system according to claim 2, wherein: The step of adjusting the parameters of the second employee performance model based on the loss function to obtain a target employee performance model includes: Based on the loss function, adjusting the parameters of the second employee performance model using a gradient descent algorithm; The second employee performance model that minimizes the loss function is determined as the target employee performance model.

7. The system according to claim 1, wherein: When obtaining the public data set in the target data set, obtaining the target data set of the target position that the target employee is engaged in includes: Crawling an initial public data set associated with the target position, where the data sample format of the initial public data set is a text format; Using natural language processing (NPL) technology to process each data sample in the initial public dataset, obtaining educational information, years of work experience information, at least one skill category and proficiency information, and performance information of each data sample in the initial public dataset; Based on the at least one skill category and proficiency information of each data sample in the initial public data set, a correlation analysis method or a feature importance analysis method is used to filter out the target skill category and proficiency information associated with the target position from the at least one skill category and proficiency information.

8. The system according to claim 1, wherein: The step of obtaining a target dataset of a target position in which a target employee is employed includes: Obtain the educational background information, years of work experience, target skill category and proficiency information, and performance information of each sample in the public dataset, and the educational background information, years of work experience, target skill category and proficiency information, and performance information of each sample in the enterprise dataset; If the first sample in the public data set does not contain the performance information, the mapping relationship between the educational information, years of work experience information, target skill category and proficiency information and the performance information in the enterprise data set is used to obtain the performance information of the first sample based on the educational information, years of work experience information, target skill category and proficiency information in the first sample.

9. The system according to claim 2, wherein: The first employee performance model is a classification feature enhanced CatBoost model, and the second employee performance model is a natural gradient boosting NGBoost model.

10. The system according to claim 1, wherein: The system further includes a judgment module; The judgment module is used to judge whether the employee to be screened meets the recruitment conditions based on the predicted performance of the employee to be screened and a preset performance threshold.