Resource variable prediction method and system, computer equipment and storage medium
By acquiring and adjusting the initial resource variables of the target object in the resource variable prediction in the logistics field, the problems of low prediction accuracy and difficulty in abnormal detection in the prior art are solved, and higher prediction accuracy and abnormal detection capabilities are achieved.
Patent Information
- Application Number
- CN202311834355.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-12-27
- Publication Date
- 2025-06-27
AI Technical Summary
The prior art has low accuracy in the logistics field of resource variable prediction, especially the prediction accuracy at the beginning of the month, and it is difficult to detect abnormal situations in predicting resource variables.
By obtaining the initial resource variable of the target object, determining the target object set from the preset object set based on its working data, and abnormal detection of sample resource variables is performed. If the abnormal detection result indicates that the initial resource variable is abnormal, it is adjusted to obtain the target resource variable.
It improves the accuracy of resource variable prediction, can adaptively detect exceptions of initial resource variables, filter out the target objects of exceptions, and adjust them when the initial resource variable is abnormal, thereby improving the accuracy of prediction.
Smart Images

Figure CN120219093A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of resource variable prediction, and particularly to a resource variable prediction method, system, computer device and storage medium. Background Art
[0002] Currently, resource variables can be used to reflect the work situation of users. Resource variables can include fixed resources and reward resources. In the related art, both fixed resources and reward resources need to be actually calculated according to the work situation of users in the current month. This calculation method uses the way of manual verification and calculation of various data, resulting in low calculation efficiency and high error probability.
[0003] However, in the logistics field, resource variable prediction is affected by many factors, such as special situations like scheduling time, holiday peaks, heavy rain, etc., making the calculation of resource variables not linearly increasing. Only relying on traditional methods to achieve the prediction of resource variables of users in the logistics field and obtaining the actual resource variables obtained by employees as of the current day will result in higher prediction accuracy closer to the end of the month and lower accuracy at the beginning of the month, thus leading to low prediction accuracy and large errors. Based on this, how to provide a resource variable prediction method to improve the accuracy of resource prediction has become an urgent technical problem to be solved. Summary of the Invention
[0004] The present invention provides a resource variable prediction method, system, computer device and storage medium, aiming to improve the accuracy of resource variable prediction.
[0005] To solve the above technical problems, an embodiment of the present invention provides a resource variable prediction method, including:
[0006] Obtain the initial resource variable of the target object;
[0007] Determine the target object set corresponding to the target object from a preset object set according to the work data of the target object; wherein, the target object set includes sample objects;
[0008] Perform anomaly detection on the initial resource variable according to the sample resource variables of the sample objects to obtain an anomaly detection result;
[0009] If the anomaly detection result indicates that the initial resource variable is abnormal, adjust the initial resource variable to obtain the target resource variable of the target object.
[0010] As a preferred solution, performing anomaly detection on the initial resource variable according to the sample resource variables of the sample objects to obtain an anomaly detection result includes:
[0011] Sort the sample resource variables and the initial resource variable to obtain a sorting result;
[0012] Determine the first comparison variable and the second comparison variable according to the sorting result;
[0013] Perform anomaly detection on the initial resource variable based on the first comparison variable and the second comparison variable to obtain the anomaly detection result.
[0014] As a preferred solution, perform anomaly detection on the initial resource variable according to the sample resource variable of the sample object to obtain the anomaly detection result, including:
[0015] Determine the sample cumulative variable of the sample object based on the sample resource variable, and determine the sample cumulative delivery volume of the sample object;
[0016] Determine the target cumulative variable of the target object according to the initial resource variable, and determine the target cumulative delivery volume of the target object;
[0017] Calculate the average resource variable based on the sample cumulative variable and the target cumulative variable, and calculate the average delivery volume based on the sample cumulative delivery volume and the target cumulative delivery volume;
[0018] Perform anomaly detection on the initial resource variable according to the target cumulative variable, the target cumulative delivery volume, the average delivery volume and the average resource variable to obtain the anomaly detection result.
[0019] As a preferred solution, perform anomaly detection on the initial resource variable according to the target cumulative variable, the target cumulative delivery volume, the average delivery volume and the average resource variable to obtain the anomaly detection result, including:
[0020] Construct sample comparison data according to the sample cumulative variable and the sample cumulative delivery volume, and construct target comparison data according to the target cumulative variable and the target cumulative delivery volume;
[0021] Calculate the average comparison data based on the sample comparison data and the target comparison data, and calculate the variance comparison data based on the sample comparison data and the target comparison data;
[0022] Calculate the sample distance based on the sample comparison data, the average comparison data and the variance comparison data;
[0023] Calculate the target distance based on the target comparison data, the average comparison data and the variance comparison data;
[0024] Calculate the average distance based on the sample distance and the target distance;
[0025] Determine the first distance difference between the target distance and the average distance, and determine the second distance difference between the sample distance and the average distance;
[0026] If the first distance difference is greater than the second distance difference, calculate the prediction deviation data based on the first distance difference;
[0027] Anomaly detection is performed on the initial resource variable based on the prediction deviation data and a preset threshold to obtain the anomaly detection result.
[0028] As a preferred solution, obtaining the initial resource variable of the target object includes:
[0029] Determine the number of days the target object should be present within a preset period, and determine the current attendance days of the target object;
[0030] Determine the initial cumulative variable of the target object;
[0031] Calculate the initial resource variable based on the initial cumulative variable, the current attendance days, and the number of days the target object should be present.
[0032] As a preferred solution, calculating the initial resource variable based on the initial cumulative variable, the current attendance days, and the number of days the target object should be present includes:
[0033] Calculate the preliminary resource variable based on the initial cumulative variable, the current attendance days, and the number of days the target object should be present;
[0034] Determine the target cumulative delivery volume of the target object based on the current attendance days;
[0035] Determine the historical average attendance days of the target object based on the target cumulative delivery volume;
[0036] Calculate the adjustment threshold based on the historical average attendance days and the current attendance days;
[0037] Adjust the preliminary resource variable according to the adjustment threshold to obtain the initial resource variable.
[0038] As a preferred solution, before determining the target object set corresponding to the target object from the preset object set according to the work data of the target object, the resource variable prediction method further includes constructing the preset object set, specifically:
[0039] Obtain the delivery area and job type of the sample object;
[0040] Perform clustering processing on the sample objects based on the delivery area and job type to obtain the preset object set.
[0041] To solve the same technical problem, an embodiment of the present invention further provides a resource variable prediction system, including: a resource variable acquisition module, an object classification module, an anomaly detection module, and an anomaly adjustment module;
[0042] Among them, the resource variable acquisition module is used to obtain the initial resource variable of the target object;
[0043] The object classification module is used to determine the target object set corresponding to the target object from the preset object set according to the work data of the target object; among them, the target object set includes sample objects;
[0044] The anomaly detection module is used to perform anomaly detection on the initial resource variables according to the sample resource variables of the sample object to obtain the anomaly detection result;
[0045] The anomaly adjustment module is used to adjust the initial resource variables if the anomaly detection result indicates that the initial resource variables are abnormal, so as to obtain the target resource variables of the target object.
[0046] To solve the same technical problem, an embodiment of the present invention further provides a computer device, including a processor and a memory. The memory is used to store a computer program, and when the computer program is executed by the processor, a resource variable prediction method is implemented.
[0047] To solve the same technical problem, an embodiment of the present invention further provides a computer-readable storage medium, which stores a computer program, and when the computer program is executed by the processor, a resource variable prediction method is implemented.
[0048] Implementing the embodiments of the present invention, determining the target object set corresponding to the target object from the preset object set according to the working data of the target object realizes the classification and differential processing of the target object. Anomaly detection is performed on the initial resource variables of the target object according to the sample resource variables corresponding to the target object set. If the anomaly detection result indicates that the initial resource variables are abnormal, the initial resource variables of the target object can be adjusted to obtain the target resource variables of the target object. It can be seen from this that the embodiments of the present invention can adaptively detect whether the initial resource variables are abnormal, adaptively screen out the target objects with abnormal initial resource variables, and make adjustments when the initial resource variables are abnormal, thereby improving the accuracy of predicting the resource variables of the target object. Description of the Drawings
[0049] Figure 1 : It is a schematic flowchart of an embodiment of a resource variable prediction method provided by the present invention;
[0050] Figure 2 : It is a flowchart of a prediction method for initial resource variables in an embodiment of a resource variable prediction method provided by the present invention;
[0051] Figure 3 : It is a flowchart of anomaly detection for initial resource variables in an embodiment of a resource variable prediction method provided by the present invention;
[0052] Figure 4 : It is a schematic structural diagram of an embodiment of a resource variable prediction system provided by the present invention. Detailed Embodiments
[0053] The technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0054] In the logistics field, according to the attendance and performance of couriers in the current month, the cumulative resource variable of the courier up to the current time can be obtained. At the same time, affected by many factors such as scheduling time and holiday peaks, the calculation of the resource variable is not linearly increasing, and the prediction accuracy is relatively low. The existing traditional method obtains the predicted resource variable actually obtained by the employee up to the current day. Therefore, the closer to the end of the month, the higher the prediction accuracy, and the lower the accuracy at the beginning of the month. Moreover, it is impossible to detect abnormal situations of the predicted resource variable. In view of this scenario, how to predict the employee's resource variable and detect abnormalities in the estimated resource variable has become a major problem in the prediction of employee resource variables in the logistics industry. Therefore, in order to improve the prediction accuracy of the resource variable and adaptively detect abnormal employee resource variables, the present invention combines the historical salary distribution of employees with scheduling and pick-up / delivery situations to perform weighted prediction of the resource variable, and can dynamically give the predicted resource variable of the employee for the current month every day; classify employees based on the clustering algorithm, and use the Mahalanobis distance and Grubbs test to realize the adaptive detection of abnormalities in the initially estimated resource variable of employees, so as to focus on and strongly remind employees with abnormal predicted resource variables, which is convenient for enterprises to conduct financial planning and employee management.
[0055] It should be noted that in each specific embodiment of the present invention, when it comes to relevant processing that needs to be carried out according to data related to the user's identity or characteristics, such as user information, user delivery data, user resource variable data, and user location information, the user's permission or consent will be obtained first. Moreover, the collection, use, and processing of these data will comply with relevant laws, regulations, and standards. In addition, when the embodiments of the present invention need to obtain sensitive personal information of users, the user's separate permission or separate consent will be obtained through pop-up windows or by jumping to a confirmation page. After clearly obtaining the user's separate permission or separate consent, the necessary user-related data for the normal operation of the embodiments of the present invention will be obtained.
[0056] Please refer to Figure 1 , which is a schematic flowchart of a resource variable prediction method provided by an embodiment of the present invention. The salary prediction and anomaly detection method includes steps 101 to 104, and the specific steps are as follows:
[0057] Step 101: Obtain the initial resource variable of the target object.
[0058] It should be noted that the resource variable is used to reflect the work situation of employees, which can be the accrued amount or the monthly salary. The target object represents an employee, and the initial resource variable represents the predicted monthly salary or the monthly accrued amount of the employee.
[0059] Exemplarily, for the logistics field, when the target object is an employee and the initial resource variable is the predicted monthly salary, to obtain the predicted monthly salary of the employee, an improved weighted average prediction method based on the traditional average method can be selected for salary weighted prediction. Specifically, according to the employee's monthly attendance data and historical pick-up and delivery volume data, the predicted monthly salary of the employee can be dynamically obtained every day.
[0060] In an alternative embodiment, step 101 includes steps S11 to S13, and the specific steps are as follows:
[0061] S11: Determine the number of days the target object should be present within a preset time period, and determine the current attendance days of the target object.
[0062] In one embodiment, exemplarily, when the target object is an employee and the initial resource variable is the predicted monthly salary, the preset time period is the current month, the number of days the employee should be present is obtained through the monthly attendance data, and the monthly attendance data includes the monthly attendance schedule, actual attendance situation data, and the monthly cumulative pick-up and delivery volume. According to the employee's monthly attendance schedule and actual attendance situation data, the number of days the employee should be present and the current attendance days can be respectively counted. The current attendance days refer to the attendance days corresponding to the current date, and the month refers to the current month. Among them, the actual attendance situation data is obtained by searching the waybill wide table database according to the employee number field and the current date.
[0063] It should be noted that employees in the logistics industry will record whether they are present on the corresponding device when receiving and delivering packages, and record each pick-up and delivery behavior according to the waybill number. These pick-up and delivery information will be imported into the company's waybill wide table database for employees. By searching the waybill wide table database, the pick-up and delivery volume and the attendance information corresponding to the employee can be obtained according to the date and the employee number field.
[0064] S12: Determine the initial cumulative variable of the target object.
[0065] In one embodiment, according to the accrual table of the target object, the resource variables accrued to the current day are statistically calculated to obtain the initial cumulative variable, and the initial cumulative variable refers to the resource variable accrued to the current date.
[0066] S13: Calculate the initial resource variable based on the initial cumulative variable, the current attendance days, and the number of days the employee should be present. In an alternative embodiment, step S13 includes steps S1311 to S1315, and the specific steps are as follows:
[0067] S1311: Calculate a preliminary resource variable based on an initial cumulative variable, the current number of attendance days, and the number of days to be attended.
[0068] In one embodiment, the preliminary resource variable represents an initial estimated monthly salary. Based on the initial cumulative variable, the current number of attendance days, and the number of days to be attended, the salary corresponding to the scheduled days of the current month is calculated using the average method to obtain the preliminary resource variable.
[0069] In one embodiment, by way of example, when the target object is an employee and the preliminary resource variable is the initial estimated monthly salary, based on obtaining the initial cumulative variable corresponding to the employee and the current number of attendance days, the average method is used to obtain the initial estimated monthly salary corresponding to the employee as of the current date. By way of example, if employee A has attended work for 3 days from the 1st to the 5th of a certain month (i.e., the current number of attendance days is 3 days), the initial cumulative variable is 600 yuan, and the scheduled number of days in the month is 20 days (i.e., the number of days to be attended is 20 days), then the preliminary resource variable predicted for employee A in the month on the 5th is 600 / 3*20 = 4000 yuan.
[0070] S1312: Determine the target cumulative delivery volume of the target object according to the current number of attendance days.
[0071] In one embodiment, the target cumulative delivery volume may refer to the cumulative delivery volume of the target object in a certain month as of the current number of attendance days. The target cumulative delivery volume can be obtained by accumulating all the delivery volumes corresponding to the current number of attendance days. For example, if employee A has attended work for 3 days from the 1st to the 5th of a certain month, where the delivery volume on the first day of attendance is 400 tickets, the delivery volume on the second day of attendance is 1000 tickets, and the delivery volume on the third day of attendance is 100 tickets, then the target cumulative delivery volume of employee A as of the 5th is 2000 tickets. It can be understood that the target cumulative delivery volume can be counted only according to the delivery volume, or according to the delivery volume and the receipt volume, and the embodiments of the present invention do not make specific limitations in this regard.
[0072] It should be noted that by searching the waybill wide-table database, the number of waybills under different receiving and delivering behaviors of each target object is calculated according to the employee ID dimension, and it is used as the receiving and delivering volume of the target object on the current day. S1313: Determine the historical average number of attendance days of the target object according to the target cumulative delivery volume.
[0073] In one embodiment, the historical average attendance days may refer to the attendance days used by the target object to reach the target cumulative delivery volume per month within the historical time. Search the waybill wide table database to obtain the historical cumulative receipt and delivery volume of the target object. Based on the target cumulative delivery volume, determine the cumulative working days corresponding to the historical cumulative receipt and delivery volume to obtain the historical attendance days. Perform an averaging process on multiple historical attendance days to obtain the historical average attendance days. For example, according to the confirmation by searching the waybill wide table database, it takes an average of 6 days for the historical cumulative delivery volume to reach 2000 tickets. Therefore, the historical attendance days can be confirmed as 6.
[0074] S1314: Calculate the adjustment threshold according to the historical average attendance days and the current attendance days.
[0075] In one embodiment, the adjustment threshold can be obtained by dividing the historical average attendance days by the current attendance days.
[0076] It should be noted that due to the specific particularity of employees' work in the logistics industry, in addition to the influence of attendance and scheduling, employees' resource variables are also affected by other external factors such as weather, scheduling time, and holiday peaks. This makes it possible for the preliminary resource variables calculated only based on the supposed attendance days and the current attendance days to have errors. The preliminary resource variables may be too high or too low. Especially at the beginning of the month, if an employee delivers a large number of waybills on the first day and has a high salary for that day, the predicted preliminary resource variable for the current month may be too high. Therefore, to avoid this situation, the adjustment threshold is analyzed through the historical average attendance days and the current attendance days of the employee, and the adjustment threshold is used as a weight to further correct the preliminary resource variable to ensure the accuracy of the modified resource variable (i.e., the initial resource variable).
[0077] S1315: Adjust the preliminary resource variable according to the adjustment threshold to obtain the initial resource variable.
[0078] In one embodiment, by way of example, when the target object is an employee, the preliminary resource variable is the initial estimated monthly salary, and the initial resource variable is the predicted salary for the current month, the adjustment threshold is used as the weight to adjust the preliminary resource variable. Divide the preliminary resource variable by the adjustment threshold to obtain a more accurate predicted resource variable with the weight factor taken into account, that is, the initial resource variable.
[0079] It should be noted that the weighted calculation is based on the adjustment threshold calculated from the historical average attendance days instead of the historical cumulative salary. This is because the nature of the work, job type, and work location of the courier often change, and there are significant differences in the salary calculation rules for different job types and work areas. Therefore, the historical salary calculation rules may vary greatly from the current month's salary calculation rules and are not of reference value. By analyzing the target cumulative delivery volume and historical average attendance days, the salary influencing factors of employees can be more accurately determined, improving the accuracy of salary prediction.
[0080] In one embodiment, the flowchart of the prediction method for the initial resource variable is as Figure 2 shown. Stat the current attendance days and initial cumulative variable of the employee, and calculate the preliminary resource variable; process the historical delivery volume data, calculate the target cumulative delivery volume and historical average attendance days, and calculate the adjustment threshold; determine the initial resource variable (final predicted salary) based on the preliminary resource variable and the adjustment threshold. Exemplarily, Employee B has worked for 2 days as of the 5th of the month (i.e., the current attendance days is 2), the cumulative delivery volume is 2000 tickets (i.e., the target cumulative delivery volume is 2000), the cumulative salary is 4000 yuan (i.e., the initial cumulative variable is 4000), and the scheduled working days in the current month is 20 days (i.e., the expected attendance days is 20). Then the initial predicted salary (i.e., the preliminary resource variable) is 4000 / 2*20 = 40000. However, there is a holiday event at the beginning of the month, which is a peak period for delivery. This salary prediction is obviously unreasonable. Then calculate the adjustment threshold based on the historical delivery order data. The monthly working days for this employee to reach the average cumulative delivery volume of 2000 tickets in history is 6 days (i.e., the historical average attendance days is 6). Then the adjusted predicted salary (i.e., the initial resource variable) is 40000 / (6 / 2) = 13333. Similarly, if Employee C has worked for 5 days on the 10th and only has a delivery volume of 1000 tickets, and the number of days to reach a delivery volume of 1000 tickets in history is 2 days, then the weighted estimated salary of this employee is (1000 / 5*20) / (2 / 5) = 10000. It can be seen that the present invention combines the delivery scenarios of logistics and express delivery and the historical delivery behaviors of employees. Through salary weighted prediction, a reasonable adjustment of the original estimated salary is achieved based on the adjustment threshold, effectively avoiding prediction errors caused by holiday peaks or heavy rain and epidemics, and improving the accuracy of prediction.
[0081] It should be noted that the resource variable budget plays an important role in the enterprise's human resource management. It helps the enterprise reasonably formulate the salary standards for employees, promotes employees' work enthusiasm and satisfaction, and also plays an important role in the enterprise's business decisions. For the enterprise, accurately estimating the resource variables of employees helps it better conduct financial planning and management. In the logistics field, based on the attendance and performance of couriers in the current month, the cumulative resource variables of couriers up to the current time can be obtained, and it is hoped to predict the total accrual amount that the employee can receive in the current month every day. For employees, dynamically displaying the resource variables can inform employees of the impact of accruals on their salaries during the work process and strengthen their work awareness.
[0082] Step 102: Determine the target object set corresponding to the target object from the preset object set according to the work data of the target object; wherein, the target object set includes sample objects.
[0083] In one embodiment, before determining the target object set corresponding to the target object from the preset object set according to the work data of the target object, the resource variable prediction method further includes constructing the preset object set, specifically:
[0084] Obtain the delivery areas and job types of the sample objects; perform clustering processing on the sample objects based on the delivery areas and job types to obtain the preset object set.
[0085] It should be noted that the work data represents the work type-related data of the target object such as the job type and the pick-up and delivery area type, and the preset object set is the clustering result of the sample objects.
[0086] In one embodiment, in the logistics field, employees are divided into different job types according to their work natures: small-piece couriers, large-piece couriers, heavy-cargo couriers, international-piece couriers, and mobile couriers, etc. The pick-up and delivery areas are also divided into different categories: CBD, residential areas, remote townships, mountainous areas, schools and hospitals, and industrial areas, etc.
[0087] Implementing the embodiments of the present invention, due to the particularity of the logistics industry, each region and job will have its specific salary payment rules, and the pick-up and delivery difficulties are also different under different regional types and job types. If the abnormality of the initial resource variables is uniformly detected and estimated, the detected employees will be inaccurate. Therefore, it is necessary to use the clustering algorithm to classify and differentiate employees under different regions and job types, and corresponding accrual rules can also be given to facilitate the analysis of the abnormality detection of the initial resource variables of employees under different classifications and improve the accuracy of the abnormality detection of the initial resource variables.
[0088] Optionally, by way of example, when the sample object is an employee, clustering the sample object based on the delivery area and job type to obtain a preset object set, which specifically includes steps 1021 to 1023. The specific steps are as follows:
[0089] Step 1021: Obtain the employee wide-table data in the waybill wide-table database to get the work situation data of all employees.
[0090] In one embodiment, the work situation data includes pick-up and delivery data, attendance data, customer complaint data, and pick-up and delivery duration data. Through the employee wide-table in the waybill wide-table database, fields such as the pick-up and delivery volume, attendance days, whether a customer complaint is generated for this waybill, and pick-up and delivery duration data of each employee can be obtained, so as to obtain the work situation data of the employees.
[0091] Step 1022: Perform data standardization processing on the work situation data of all employees and calculate the feature data of all employees.
[0092] In one embodiment, the feature data includes the attendance days in the current month, pick-up and delivery volume, customer complaint volume, and average pick-up and delivery duration. Perform data preprocessing, that is, data standardization processing, on the work situation data of all employees (i.e., sample objects), and calculate the feature data such as the attendance days in the current month, pick-up and delivery volume, customer complaint volume, and average pick-up and delivery duration of each employee after standardization.
[0093] Step 1023: Use the feature data of all employees as the input features of the K-means clustering algorithm, and perform clustering division on all employees according to the preset number of clustering categories by using the K-means clustering algorithm to obtain the clustering result (i.e., the preset object set); among them, the preset number of clustering categories is the best number of categories selected by calculating the silhouette coefficient under different numbers of categories through experiments.
[0094] In one embodiment, use the feature data of employees as the feature input for clustering, and perform clustering division on employees by using the K-means clustering algorithm. The K-means clustering algorithm will distinguish employees with similar features. Among them, the best number of categories corresponding to the preset number of clustering categories is 4. The reason is that: calculate the silhouette coefficient under different clustering numbers according to the experimental results. The silhouette coefficient is used to describe the similarity between samples in the clustering data. The closer it is to 1, the relatively better the intra-class similarity and inter-class separation degree are. By calculating the silhouette coefficient under different numbers of clusters, it is determined that the clustering effect is the best when the number of categories is 4.
[0095] It should be noted that the K-means clustering algorithm is an iterative clustering analysis algorithm. Its steps are as follows: initially divide the data into K groups, then randomly select K objects as the initial clustering centers, and then calculate the distance between each object and each seed clustering center, and assign each object to the clustering center closest to it. The clustering centers and the objects assigned to them represent a cluster. Each time a sample is assigned, the clustering center of the cluster will be recalculated based on the existing objects in the cluster. This process will be repeated continuously until a certain termination condition is met. The termination condition can be that no (or the minimum number of) objects are reassigned to different clusters, no (or the minimum number of) clustering centers change anymore, or the sum of squared errors is locally minimized. The K-means clustering algorithm can prune the tree based on the categories of fewer known clustering samples to determine the classification of some samples. To overcome the inaccuracy of clustering with a small number of samples, the algorithm itself has an optimization iteration function. It iteratively corrects and prunes on the already obtained clusters to determine the clustering of some samples, optimizing the unreasonable parts of the initial supervised learning sample classification. Since it only targets some small samples, it can reduce the overall clustering time complexity. Therefore, the K-means clustering algorithm is selected to classify the characteristic data of employees.
[0096] In one embodiment, count the number of employees in different regional types and job types under these four categories. Designate the delivery areas such as large institutions like schools and hospitals, residential areas, remote townships and mountainous areas, and commercial-residential mixtures as difficult areas, and other regional types as non-difficult areas. Designate the jobs such as large-item couriers, mobile couriers, and mobile large-item couriers as difficult delivery jobs, and other job types as non-difficult delivery jobs. According to the delivery area and job type, the employees are divided into 4 types, that is, 4 preset object sets can be obtained. It can be understood that the delivery area can refer to the area corresponding to when an object is delivered, or it can refer to the area corresponding to both delivery and receipt. The embodiments of the present invention do not make specific limitations on this.
[0097] In some embodiments, the target object set corresponding to the target object can be determined from the 4 preset object sets based on the regional type and job type of the target object.
[0098] Step 103: Perform anomaly detection on the initial resource variable according to the sample resource variable of the sample object to obtain the anomaly detection result.
[0099] In one embodiment, in the logistics field, since the delivery difficulty is also different under different delivery areas and different job types, it is necessary to classify the employees under different delivery areas and job types and then detect the salaries of abnormal employees. The process of anomaly detection of the initial resource variable is as follows Figure 3As shown in the figure, for the anomaly detection of the initial resource variable, it can be determined by the quartile statistical analysis method, or by the similarity detection method, or by combining the above two methods. The embodiments of the present invention do not make specific limitations on this.
[0100] It should be noted that employees are classified according to their job types and delivery areas, and each area and job will have its specific salary payment rules. The resource variable distributions of the sample objects concentrated in various target objects are different. Through quartile statistical analysis, it can be preliminarily determined whether the target object is an abnormal object. For example, the target objects with too high or too low initial resource variables are regarded as abnormal objects. The prediction method for the initial resource variable in step 101 can make relatively accurate predictions for most employees, but there will still be cases of abnormal predictions every day. For example, some newly recruited employees do not have historical delivery and pickup volumes as a reference, or there are other reward and punishment policies that cause the predicted salary of the employee to be too high or too low. For these employees, we need to perform anomaly detection and conduct manual review or strong reminders.
[0101] In an optional embodiment, the anomaly detection result is determined by the quartile statistical analysis method, including steps S11 to S13. The specific steps are as follows:
[0102] S11: Sort the sample resource variable and the initial resource variable to obtain a sorting result.
[0103] In an embodiment, all sample objects in the target object set are integrated with the target object to obtain an integrated object set. The resource variable distribution of the integrated object set is statistically analyzed, and the upper and lower quartiles of the resource variable of the integrated object set are calculated. It should be noted that the sample resource variable is the resource variable of the sample object, and the sample resource variable can be determined according to the method described in the embodiments of the present invention or other methods. The embodiments of the present invention do not make specific limitations on this. The resource variables (including sample resource variables and target resource variables) in the integrated object set are sorted from small to large to obtain a sorting result.
[0104] S12: Determine the first comparison variable and the second comparison variable according to the sorting result.
[0105] In an embodiment, the calculation process of the upper and lower quartiles in statistics is as follows: First, the data is arranged in ascending order to obtain a sorting result. Then, the positions of the first 25% and the last 25% are calculated, that is, the first quartile and the last quartile of the data. Finally, the first quartile and the last quartile are respectively used as the values of the upper and lower quartiles to obtain the upper and lower quartiles. It should be noted that the first comparison variable is the first 25% (upper quartile), and the second comparison variable is the last 25% (lower quartile).
[0106] S13: Perform anomaly detection on the initial resource variable based on the first comparison variable and the second comparison variable to obtain an anomaly detection result.
[0107] In one embodiment, calculate the upper and lower limits of the anomaly resource variable by calculating the upper and lower quartiles of the integrated object set. Specifically, the upper and lower limits of the anomaly resource variable = upper quartile ± 2 * (upper quartile - lower quartile). Compare the initial resource variable of the target object with the upper and lower limits of the anomaly resource variable to obtain an anomaly detection result.
[0108] It should be noted that if the initial resource variable of the target object is greater than the upper limit of the anomaly resource variable or the initial resource variable is less than the lower limit of the anomaly resource variable, it is determined that the initial resource variable of the target object is abnormal.
[0109] It can be understood that the initial resource variable being greater than the upper limit of the anomaly resource variable or less than the lower limit of the anomaly resource variable does not necessarily mean that the initial resource variable is truly abnormal. In some embodiments, it can also be combined with its pick-up and delivery situation for judgment to improve the accuracy of the anomaly detection result. For example, since the resource variable is positively correlated with the pick-up and delivery volume, if the initial resource variable of the target object is greater than the upper limit of the anomaly resource variable but the pick-up and delivery volume is very low, an anomaly detection result that the initial resource variable is abnormal can be obtained. Or, if the initial resource variable is less than the lower limit of the anomaly resource variable but the pick-up and delivery volume is very high, an anomaly detection result that the initial resource variable is abnormal can be obtained.
[0110] In an alternative embodiment, determine the anomaly detection result according to the similarity detection method, including steps S21 to S24. The specific steps are as follows:
[0111] S21: Determine the sample cumulative variable of the sample object based on the sample resource variable and determine the sample cumulative pick-up and delivery volume of the sample object.
[0112] In one embodiment, the sample cumulative variable can refer to the cumulative resource variable of the sample object as of the current month. The sample cumulative pick-up and delivery volume can refer to the cumulative pick-up and delivery volume of the sample object as of the current month.
[0113] S22: Determine the target cumulative variable of the target object based on the initial resource variable and determine the target cumulative pick-up and delivery volume of the target object.
[0114] In some embodiments, the target cumulative variable can refer to the variable obtained by accumulating the adjusted resource variable (i.e., the initial resource variable) of the target object as of the current month. The target cumulative pick-up and delivery volume can refer to the cumulative pick-up and delivery volume of the target object as of the current month.
[0115] S23: Calculate the average resource variable based on the sample cumulative variable and the target cumulative variable, and calculate the average delivery volume based on the sample cumulative delivery volume and the target cumulative delivery volume.
[0116] In one embodiment, calculate the mean of the sample cumulative variables and the target cumulative variables of all sample objects in the target object set to obtain the average resource variable. Calculate the mean of the sample cumulative delivery volumes and the target cumulative delivery volumes of all sample objects in the target object set, the average delivery volume. It can be understood that all sample objects in the target object set and the target object are integrated to obtain an integrated object set. Step S23 is equivalent to calculating the average delivery volume and the average resource variable of this integrated object set.
[0117] S24: Perform anomaly detection on the initial resource variable based on the target cumulative variable, the target cumulative delivery volume, the average delivery volume, and the average resource variable to obtain an anomaly detection result.
[0118] In an alternative embodiment, perform anomaly detection on the initial resource variable based on the target cumulative variable, the target cumulative delivery volume, the average delivery volume, and the average resource variable to obtain an anomaly detection result, including steps S241 to S248. The specific steps are as follows:
[0119] S241: Construct sample comparison data based on the sample cumulative variable and the sample cumulative delivery volume, and construct target comparison data based on the target cumulative variable and the target cumulative delivery volume.
[0120] In one embodiment, the sample cumulative variable and the sample cumulative delivery volume can be represented in two-dimensional vector form to obtain sample comparison data (sample cumulative delivery volume, sample cumulative variable). Similarly, the target cumulative variable and the target cumulative delivery volume can be represented in two-dimensional vector form to obtain target comparison data (target cumulative delivery volume, target cumulative variable).
[0121] S242: Calculate the average comparison data based on the sample comparison data and the target comparison data, and calculate the variance comparison data based on the sample comparison data and the target comparison data.
[0122] In one embodiment, calculate the mean of the comparison data (including sample comparison data and target comparison data) of all objects in the integrated object set to obtain the mean vector μ (i.e., the average comparison data). Calculate the covariance of the comparison data (including sample comparison data and target comparison data) of all objects in the integrated object set to obtain the covariance matrix Σ (i.e., the variance comparison data).
[0123] S243: Calculate the sample distance based on the sample comparison data, the average comparison data, and the variance comparison data.
[0124] In some embodiments, the sample distance can be calculated based on the Mahalanobis distance for the sample comparison data, the average comparison data, and the variance comparison data.
[0125] S244: Calculate the target distance based on the target comparison data, the average comparison data, and the variance comparison data.
[0126] In some embodiments, the target distance can be calculated based on the Mahalanobis distance for the target comparison data, the target comparison data, and the variance comparison data.
[0127] It can be understood that when the sample objects and the target objects in the target object set are integrated into an integrated object set, steps S243 and S244 are essentially to calculate the distances of each object (including sample objects or target objects) in this integrated object set to determine the similarity between each object in this integrated object set and the "average vector" of this integrated object set. Taking the calculation of the sample distance d using the Mahalanobis distance method as an example, the calculation process is as follows: According to the two-dimensional vector a = (x, y) of the sample object (i.e., the sample comparison data (sample cumulative delivery volume, sample cumulative variable)), the mean vector μ and the covariance matrix Σ are calculated, and based on the mean vector and the covariance matrix, the sample distance d is calculated. The formula is as follows:
[0128] d = sqrt((a - μ)'Σ^(-1)(a - μ))
[0129] Where, sqrt represents taking the square root; ' represents vector transpose; Σ^(-1) represents the inverse of the matrix.
[0130] S245: Calculate the average distance based on the sample distance and the target distance.
[0131] In one embodiment, the average distance is mean(d), which represents the average value of the sample distance and the target distance.
[0132] S246: Determine the first distance difference between the target distance and the average distance, and determine the second distance difference between the sample distance and the average distance.
[0133] S247: If the first distance difference is greater than the second distance difference, calculate the prediction deviation data based on the first distance difference.
[0134] In one embodiment, when the first distance difference is greater than all the second distance differences, it indicates that the target object is the point farthest from the mean. At this time, based on the first distance difference and the following formula, the G value of the prediction deviation data is calculated:
[0135] G = (max(abs(d - mean(d))) / std.dev(d))
[0136] Among them, max(abs(d - mean(d))) represents the difference between the point farthest from the mean and the average value (i.e., the first distance deviation), and std.dev(d) represents the standard deviation.
[0137] S248: Based on the prediction deviation data and a preset threshold, perform anomaly detection on the initial resource variables to obtain an anomaly detection result.
[0138] In one embodiment, according to a preset significance level, determine whether the G statistic of the current farthest point is greater than the critical value obtained by looking up the table. The preset threshold is the critical value obtained by looking up the table. By setting the significance level, judge the G statistic of the current farthest point to determine whether the G statistic of the current farthest point is greater than the critical value obtained by looking up the table. If the G statistic of the current farthest point is less than the critical value obtained by looking up the table, the G statistics corresponding to the resource variables of all objects in the current integration object set all meet the significance level. Among them, the significance level is a pre-determined value, generally represented by alpha, which is exactly opposite in direction to the confidence probability (the sum is 1).
[0139] It should be noted that a significance test is performed based on the test statistic G to determine whether there is a significant difference between the detection samples. The G value needs to be compared with the critical value G0 given in the Grubbs table. If it is greater than G0, it is determined that the predicted salary may be an outlier and should be excluded. The critical value G0 is related to the confidence probability p and the sample size n. The confidence probability can be set to 0.95. Since the number of employees of different types in each city is different, the n value is also inconsistent. According to these two parameters, the corresponding critical value can be obtained from the Grubbs table. When the number of employees in a certain target object set is very large, the critical value cannot be obtained by looking up the table, and the outlier can be directly excluded by the 3sigma principle.
[0140] It should be noted that when the prediction deviation value is greater than the preset threshold, it indicates that the target object is an abnormal object, and an anomaly detection result of abnormal initial resource variables of the target object can be obtained. On the contrary, an anomaly detection result of normal initial resource variables can be obtained.
[0141] In some embodiments, the object corresponding to the current farthest point can be marked as a similarity abnormal object in the current integration object set, the current farthest point is removed, and the G statistic of the removed current farthest point is calculated according to the Mahalanobis distances of the objects after removal in the current integration object set until the G statistic of the removed current farthest point is not greater than the critical value obtained by looking up the table, and all similarity abnormal objects in the current integration object set that are marked are obtained.
[0142] Implementing the embodiments of the present invention, through the anomaly detection method combining Mahalanobis distance and Grubbs test, it is possible to effectively discover whether the target object is an abnormal object, find out that the relationship between its pick-up and delivery volume and the cumulative salary does not conform to the normal situation, and can manually review the resource variables of the abnormal object, thereby improving the accuracy of resource variable prediction.
[0143] Step 104: If the anomaly detection result indicates that the initial resource variable is abnormal, adjust the initial resource variable to obtain the target resource variable of the target object.
[0144] In some embodiments, when the initial resource variable of the target object is detected as abnormal, the initial resource variable can be adjusted to obtain the target resource variable. It can be understood that the target resource variable can be used as the resource variable for the final prediction of the target object. The method for adjusting the initial resource variable can include manual review and adjustment, adjustment based on a formula, or adjustment based on artificial intelligence, etc., and the embodiments of the present invention do not make specific limitations on this.
[0145] Implementing the embodiments of the present invention, determining the target object set corresponding to the target object from the preset object set according to the work data of the target object realizes the classification and differential processing of the target object. Performing anomaly detection on the initial resource variable of the target object according to the sample resource variable corresponding to the target object set. If the anomaly detection result indicates that the initial resource variable is abnormal, the initial resource variable of the target object can be adjusted to obtain the target resource variable of the target object. Thus, it can be seen that the embodiments of the present invention can adaptively detect the anomaly of the initial resource variable, adaptively screen out the target objects with truly abnormal initial resource variables, and improve the accuracy of resource variable prediction for the target object.
[0146] See Figure 4 , Figure 4 is the structural schematic diagram of Embodiment 2 of a resource variable prediction system provided by the present invention. As Figure 4 shown, the resource variable prediction system includes a resource variable acquisition module 401, an object classification module 402, an anomaly detection module 403, and an anomaly adjustment module 404;
[0147] Among them, the resource variable acquisition module 401 is used to acquire the initial resource variable of the target object;
[0148] The object classification module 402 is used to determine the target object set corresponding to the target object from the preset object set according to the work data of the target object; among them, the target object set includes sample objects;
[0149] The anomaly detection module 403 is used to perform anomaly detection on the initial resource variable according to the sample resource variable of the sample object to obtain the anomaly detection result;
[0150] The anomaly adjustment module 404 is used to adjust the initial resource variable if the anomaly detection result indicates that the initial resource variable is anomalous, so as to obtain the target resource variable of the target object.
[0151] Implementing the embodiments of the present invention, determining the target object set corresponding to the target object from the preset object set according to the working data of the target object realizes the classification and differential processing of the target object. Anomaly detection is performed on the initial resource variable of the target object according to the sample resource variable corresponding to the target object set. If the anomaly detection result indicates that the initial resource variable is anomalous, the initial resource variable of the target object can be adjusted to obtain the target resource variable of the target object. It can be seen that the embodiments of the present invention can adaptively detect the anomaly of the initial resource variable, and adaptively screen out the target objects with truly anomalous initial resource variables, improving the accuracy of predicting the resource variables of the target objects.
[0152] The above-mentioned resource variable prediction system can implement a resource variable prediction method in the above method embodiments. The optional items in the above method embodiments are also applicable to this embodiment and will not be elaborated here. The remaining content of the embodiments of this application can refer to the content of the above method embodiments and will not be repeated in some embodiments.
[0153] In addition, the embodiments of the present application also provide a computer device. The computer device includes a processor and a memory. The memory is used to store a computer program. When the computer program is executed by the processor, the steps in any of the above method embodiments are implemented.
[0154] The embodiments of the present application also provide a computer-readable storage medium. The computer-readable storage medium stores a computer program. When the computer program is executed by the processor, the steps in any of the above method embodiments are implemented.
[0155] The above specific embodiments further elaborate the purpose, technical solutions and beneficial effects of the present invention. It should be understood that the above are only specific embodiments of the present invention and are not used to limit the protection scope of the present invention. It is particularly pointed out that for those skilled in the art, any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.
Claims
1. A method for predicting resource variables, characterized in that, The resource variable prediction method includes: Obtaining the initial resource variable of the target object; Determining the target object set corresponding to the target object from a preset object set according to the working data of the target object; wherein, the target object set includes sample objects; Performing anomaly detection on the initial resource variable according to the sample resource variables of the sample objects to obtain an anomaly detection result; If the anomaly detection result indicates that the initial resource variable is abnormal, adjusting the initial resource variable to obtain the target resource variable of the target object.
2. The resource variable prediction method according to claim 1, wherein The performing anomaly detection on the initial resource variable according to the sample resource variables of the sample objects to obtain an anomaly detection result includes: Sorting the sample resource variables and the initial resource variable to obtain a sorting result; Determining a first comparison variable and a second comparison variable according to the sorting result; Performing anomaly detection on the initial resource variable based on the first comparison variable and the second comparison variable to obtain the anomaly detection result.
3. The resource variable prediction method according to claim 1, wherein The performing anomaly detection on the initial resource variable according to the sample resource variables of the sample objects to obtain an anomaly detection result includes: Determining the sample cumulative variable of the sample object based on the sample resource variable, and determining the sample cumulative delivery volume of the sample object; Determining the target cumulative variable of the target object according to the initial resource variable, and determining the target cumulative delivery volume of the target object; Calculating an average resource variable according to the sample cumulative variable and the target cumulative variable, and calculating an average delivery volume according to the sample cumulative delivery volume and the target cumulative delivery volume; Performing anomaly detection on the initial resource variable according to the target cumulative variable, the target cumulative delivery volume, the average delivery volume and the average resource variable to obtain the anomaly detection result.
4. The resource variable prediction method according to claim 3, wherein The performing anomaly detection on the initial resource variable according to the target cumulative variable, the target cumulative delivery volume, the average delivery volume and the average resource variable to obtain the anomaly detection result includes: Constructing sample comparison data according to the sample cumulative variable and the sample cumulative delivery volume, and constructing target comparison data according to the target cumulative variable and the target cumulative delivery volume; Calculating average comparison data according to the sample comparison data and the target comparison data, and calculating variance comparison data according to the sample comparison data and the target comparison data; Calculating a sample distance according to the sample comparison data, the average comparison data and the variance comparison data; Calculating a target distance according to the target comparison data, the average comparison data and the variance comparison data; Calculating an average distance according to the sample distance and the target distance; Determining a first distance difference between the target distance and the average distance, and determining a second distance difference between the sample distance and the average distance; If the first distance difference is greater than the second distance difference, calculating prediction deviation data based on the first distance difference; Performing anomaly detection on the initial resource variable based on the prediction deviation data and a preset threshold to obtain the anomaly detection result.
5. The resource variable prediction method according to claim 1, wherein The obtaining of the initial resource variable of the target object includes: Determining the number of days the target object should be present during a preset period, and determining the current attendance days of the target object; Determining the initial cumulative variable of the target object; Calculating the initial resource variable based on the initial cumulative variable, the current attendance days, and the number of days the target object should be present.
6. The resource variable prediction method according to claim 5, wherein The calculating the initial resource variable based on the initial cumulative variable, the current attendance days, and the number of days the target object should be present includes: Calculating a preliminary resource variable based on the initial cumulative variable, the current attendance days, and the number of days the target object should be present; Determining the target cumulative delivery volume of the target object based on the current attendance days; Determining the historical average attendance days of the target object based on the target cumulative delivery volume; Calculating an adjustment threshold based on the historical average attendance days and the current attendance days; Adjusting the preliminary resource variable according to the adjustment threshold to obtain the initial resource variable.
7. The resource variable prediction method according to claim 1, characterized in that Before determining the target object set corresponding to the target object from the preset object set according to the work data of the target object, the resource variable prediction method further includes constructing the preset object set, specifically: Obtaining the delivery area and job type of the sample object; Performing clustering processing on the sample object based on the delivery area and the job type to obtain the preset object set.
8. A resource variable prediction system, characterized in that, Including: A resource variable obtaining module, an object classification module, an anomaly detection module, and an anomaly adjustment module; Among them, the resource variable obtaining module is used to obtain the initial resource variable of the target object; The object classification module is used to determine the target object set corresponding to the target object from the preset object set according to the work data of the target object; among them, the target object set includes sample objects; The anomaly detection module is used to perform anomaly detection on the initial resource variable according to the sample resource variable of the sample object to obtain an anomaly detection result; The anomaly adjustment module is used to adjust the initial resource variable to obtain the target resource variable of the target object if the anomaly detection result indicates that the initial resource variable is abnormal.
9. A computer device, characterized in that, Including a processor and a memory, the memory is used to store a computer program, and when the computer program is executed by the processor, it implements the resource variable prediction method according to any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that, It stores a computer program, and when the computer program is executed by the processor, it implements the resource variable prediction method according to any one of claims 1 to 7.