A big data-based human resource data management method
By constructing a behavior density score V1 and a credibility enhancement score V2, the problems of high-value data identification and redundancy in human resource data management are solved, dynamic cleaning decisions are made, and the accuracy and credibility of the system's data management are improved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- SHEBAO INFORMATION TECH SHANGHAI CO LTD
- Filing Date
- 2026-01-27
- Publication Date
- 2026-04-10
AI Technical Summary
Existing human resources data management systems are unable to effectively identify and retain high-value historical data, resulting in serious data redundancy problems that affect system performance and data credibility.
By constructing a behavior density score V1 and a personnel data credibility enhancement score V2, and combining Euclidean structure and temporal sparsity offsetting factors, the activity level of personnel behavior and data credibility are dynamically evaluated, thereby achieving synchronous consistency calibration and cleaning decisions in multi-terminal environments.
Effectively identify high-behavioral-value data, reduce the risk of accidental deletion, improve the accuracy and reliability of data management, reduce redundant data, and enhance system performance and data security.
Smart Images

Figure CN121581826B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of data management, in particular to a human resource data management method based on big data. BACKGROUND
[0002] With the continuous improvement of the degree of enterprise digitization, the application of big data technology in the field of human resource management has gradually become an important basis for fine management of enterprises. In the field of human resource management, personnel behavior analysis, business participation evaluation, and post performance judgment core modules are gradually relying on data support from multi-source systems. In this process, various personnel operation logs, business participation records and behavior event data generated by the widely deployed industrial equipment monitoring terminals, office system terminals and business execution system terminals of enterprises become important data sources for human resource systems to conduct data modeling, behavior analysis and performance auxiliary judgment. Therefore, personnel interaction behavior raw data from different terminals and synchronous deviation data in the multi-terminal running environment have become specific data categories that need to be focused on in the current human resource data management system, and are also the core data basis that the big data driven human resource system must rely on.
[0003] At present, in the existing human resource data management process, the above-mentioned personnel interaction behavior data and multi-source synchronous deviation data usually have diverse sources and non-uniform structures, and a large amount of historical data has not been systematically organized and cleaned. In this case, the system often classifies, filters or retains personnel data in a fixed rule, static threshold or manual screening manner, which cannot accurately reflect the actual value of personnel data in the current business cycle. In addition, the existing cleaning method mainly takes "whether an exception occurs" as the core judgment basis, and lacks comprehensive analysis ability of key elements such as the degree of association between personnel data and business, the degree of behavior activity, and the consistency of data synchronization. Therefore, when facing the continuously accumulated historical data, the system cannot effectively judge which data should be retained and which data can be cleaned, resulting in a large amount of low-value or no-value historical records continuously accumulated in the system, thereby making the data redundancy problem more serious.
[0004] The causes of the above problems mainly lie in that the personnel behavior data presents the characteristics of high frequency generation, cross-terminal recording, time sequence dispersion and inconsistent format, and the existing data cleaning mechanism lacks the ability to dynamically judge these characteristics. On the one hand, static rules cannot adapt to the changes of different business cycles, resulting in that some data which have not been used for a long time are still regarded as "to be retained" by the system, forming data sedimentation phenomenon; on the other hand, due to the lack of quantitative judgment mechanism for the value of personnel behavior and the correlation strength of data, the system cannot accurately identify those key data which have high business value although they are long in history. The above situations may eventually lead to abnormal effects such as decline of system processing performance, increase of data retrieval delay, distortion of personnel portrait model and increase of data deviation relied by business decision, which seriously affect the data credibility and management efficiency of enterprises in human resource management. SUMMARY
[0005] In view of the deficiencies of the prior art, the present application provides a human resource data management method based on big data, which solves the problems mentioned in the background art.
[0006] To achieve the above object, the present application is implemented by the following technical scheme: a human resource data management method based on big data, comprising the following steps:
[0007] S1, through data source access configuration of a plurality of device terminals and data acquisition interfaces of a human resource system, personnel interaction behavior data and personnel multi-source synchronization deviation data are collected to form a personnel interaction behavior original data set and a personnel multi-source synchronization deviation data set, and the personnel interaction behavior original data set is preprocessed in a personnel data correlation analysis module to obtain a behavior density feature vector;
[0008] S2, the behavior density feature vector is processed by structured calculation in the personnel data correlation analysis module, a behavior density score V1 is constructed, and the behavior density score V1 is compared with a preset behavior density threshold Vth1 to obtain a preliminary comparison result;
[0009] S3, when the preliminary comparison result automatically triggers a personnel data credibility calibration module, the personnel multi-source synchronization deviation data set is analyzed and processed to obtain a personnel data credibility enhancement score V2, and the personnel data credibility enhancement score V2 is compared with a preset credibility judgment threshold Vth2 to obtain a final personnel data cleaning judgment result;
[0010] S4, based on the final personnel data cleaning judgment result, a system-level cleaning control mechanism is started to execute cleaning rules and data recovery actions, and cleaning control records and recovery records are output to a data management log module.
[0011] Preferably, the S1 comprises S11;
[0012] S11, deploying a device-end data sending plug-in at a system communication layer of a monitoring terminal of a plurality of industrial devices, the device-end data sending plug-in being connected with a log output interface, a clock recording interface and a patrol task recording interface of the industrial device monitoring terminal through a device communication protocol adaptation module, and converting terminal operation data, terminal timestamp data and patrol trajectory data generated by the industrial device monitoring terminal into a transmittable format;
[0013] deploying a data access management plug-in at a data access layer on a human resource system side, the data access management plug-in being connected with the device-end data sending plug-in through a message transmission channel;
[0014] and setting two collection points in each monitoring terminal, the collection points including a personnel behavior parameter collection point and a multi-source synchronization parameter collection point; and setting a data structure analysis plug-in in each collection point; the data structure analysis plug-in including a behavior time analysis plug-in and a synchronization deviation detection plug-in;
[0015] The personnel behavior parameter collection point is set at a software application layer of the monitoring terminal, and a behavior event analysis plug-in is installed, the behavior event analysis plug-in being connected with an operation interface event log module of the industrial device monitoring terminal, and collecting operation frequency information, operation confirmation behavior information and patrol range identification information of personnel on an operation interface of the industrial device monitoring terminal;
[0016] The multi-source synchronization parameter collection point is set at a system bottom layer running environment of the monitoring terminal, and a synchronization deviation detection plug-in is installed, the synchronization deviation detection plug-in being connected with a system clock management module and a running state recording module of the monitoring terminal, and collecting operation timestamp offset information and behavior interruption recording information of personnel between different industrial device monitoring terminals.
[0017] Preferably, the S1 further comprises S12;
[0018] S12, extracting a personnel interactive behavior original data set and a personnel multi-source synchronization deviation data set in real time based on the data structure analysis plug-in set in the collection point; and packing the personnel interactive behavior original data set and the personnel multi-source synchronization deviation data set into a standardized data frame through the device-end data sending plug-in, and sending the standardized data frame to the data access layer on the human resource system side through the device-end data sending plug-in;
[0019] The personnel interactive behavior original data set includes operation frequency A of personnel at a device terminal, operation confirmation event quantity B submitted by the personnel and independent area quantity C of area patrol participated by the personnel;
[0020] The multi-source synchronous deviation data set of the personnel includes a time stamp offset D between personnel records at different monitoring terminals and an abnormal interruption number E of personnel in device-side behavior.
[0021] Preferably, the S1 further comprises S13.
[0022] S13, after the human resource data correlation analysis module receives the personnel interaction behavior original data set, pre-processing the personnel interaction behavior original data set to obtain a behavior density feature vector; the pre-processing includes data integrity preprocessing, time alignment processing and data standardization processing;
[0023] The data integrity preprocessing detects missing values based on the original sampling frequency for each parameter sequence in the personnel interaction behavior original data set, and repairs the missing values by mean interpolation for adjacent time periods.
[0024] The time alignment processing aligns the personnel interaction behavior original data set with the same acquisition frequency to a unified time reference point.
[0025] The data standardization processing performs data standardization processing on the personnel interaction behavior original data set after time alignment by using the minimum and maximum standardization processing method, eliminates the unit dimension influence between all parameters in the personnel interaction behavior original data set, and then performs secondary aggregation on the processed personnel interaction behavior original data set to obtain a behavior density feature vector.
[0026] Preferably, the S2 comprises S21.
[0027] S21, based on the behavior density feature vector obtained by pre-processing, the personnel data correlation analysis module uses an improved Euclidean structure combined with a time sequence sparse offset factor calculation method to calculate and output a behavior density score V1 as a measure of personnel interaction behavior activity.
[0028] The behavior density score V1 is calculated and output by the following algorithm formula:
[0029] In the formula, ln represents a natural logarithm function, and Tr represents the number of days from the current personnel's last device-side behavior.
[0030] Preferably, the S2 further comprises S22.
[0031] S22, the human resource system presets a behavior density threshold Vth1 based on the mean value of the behavior density score V1 in the full behavior value distribution of historical personnel interaction behavior data.
[0032] And based on the real-time acquisition of the behavior density score V1 and the behavior density threshold Vth1, a preliminary comparison is made; the specific comparison content is as follows:
[0033] When the behavior density score V1 is greater than or equal to the behavior density threshold Vth1, the preliminary comparison result is determined to be sufficient behavior value, and the personnel related data is marked as normal behavior value data;
[0034] When the behavior density score V1 is less than the behavior density threshold Vth1, the preliminary comparison result is determined to be insufficient behavior value, triggering the personnel data credibility calibration module to enter the next stage of the deep credibility calibration process.
[0035] Preferably, the S3 comprises S31;
[0036] S31, after triggering the personnel data credibility calibration module to enter the next stage of the deep credibility calibration process, the personnel data credibility calibration module extracts the time stamp offset D between different monitoring terminals and the number of abnormal interruptions E of the behavior of the personnel at the equipment end from the personnel multi-source synchronous deviation data set;
[0037] And the personnel data credibility enhancement score V2 is generated by calculating the behavior density score V1, the time stamp offset D between different monitoring terminals and the number of abnormal interruptions E of the behavior of the personnel at the equipment end through the personnel data credibility calibration module, and the synchronization consistency level and data credibility of the personnel related data in the multi-terminal environment are quantitatively analyzed;
[0038] The personnel data credibility enhancement score V2 is calculated and output by the following calculation formula:
[0039] ; In the formula, Indicates the synchronization deviation intensity.
[0040] Preferably, the S3 further comprises S32;
[0041] S32, the credibility determination threshold Vth2 is determined according to the statistical result of the historical multi-terminal synchronization deviation sample, and the credibility determination threshold Vth2 is the minimum acceptable credibility standard value of the personnel data credibility enhancement score V2;
[0042] After the personnel data credibility enhancement score V2 is calculated, the personnel data credibility enhancement score V2 is compared with the credibility determination threshold Vth2 for the second time by the personnel data credibility calibration module, and the determination logic of the second comparison processing is:
[0043] When the personnel data credibility enhancement score V2 is greater than or equal to the credibility determination threshold Vth2, it is determined that the personnel-related data is credible historical data, and the current personnel-related data is written into the data refinement retention area;
[0044] When the personnel data credibility enhancement score V2 is less than the credibility determination threshold Vth2, it is determined that the personnel-related data is untrustworthy historical data, and the current personnel is added to the data cleaning list and a corresponding cleaning action record is generated.
[0045] Preferably, the S4 comprises S41;
[0046] S41, after receiving the final cleaning determination result of the personnel data as personnel-related data being untrustworthy historical data, automatically calling a data cleaning scheduling module, and executing a cleaning rule of a system-level cleaning control mechanism on the target personnel data according to the scheduling instruction of the data cleaning scheduling module; the cleaning rule comprises a deletion rule, a merging rule and a freezing rule;
[0047] The deletion rule is determined to be a deletable record when the personnel data is not referenced by any system module in the latest business cycle, and the behavior density score V1 and the personnel data credibility enhancement score V2 of the data are both lower than the behavior density threshold Vth1 and the credibility determination threshold Vth2, and the current state is maintained for two consecutive cycles, and the current record is deleted;
[0048] The merging rule is to merge repeated records into a single storage unit when multiple data records have the same personnel identifier, the same time window and the content repetition ratio reaches 80%;
[0049] The freezing rule is to mark the current data record as a frozen state when the data record does not meet any one of the deletion rule and the merging rule, but the synchronous deviation intensity is greater than 0.8, the data in the frozen state reenters the credibility calibration process in the next cycle and is not deleted temporarily;
[0050] After the execution of the cleaning rule of the system-level cleaning control mechanism ends, a cleaning control record including the processing type, the processing time, the processing result and the corresponding personnel identifier is automatically generated for each processed data record, and is written into the data management log module.
[0051] Preferably, the S4 further comprises S42;
[0052] S42, when the final cleaning determination result of the personnel data is "highly credible important historical data", the system-level cleaning control mechanism sends a recovery instruction to the data recovery management module, and the data recovery management module executes a data recovery action according to the recovery instruction, removes the data record of the corresponding personnel from the cleaning list and restores it to the data refinement retention area.
[0053] The system-level cleaning control mechanism performs a consistency check before the data recovery action is executed, and the consistency check comprises:
[0054] The behavior density score V1 is greater than or equal to a behavior density threshold Vth1.
[0055] The personnel data credibility enhancement score V2 is greater than or equal to a credibility determination threshold Vth2.
[0056] The synchronization deviation intensity ≤0.8.
[0057] When all three conditions are met, the recovery operation can be executed.
[0058] After the system-level cleaning control mechanism completes the data recovery action, a recovery rule is executed, and the recovery rule is as follows:
[0059] If the target record is in a frozen state, the system will recover from the frozen area to the main data area.
[0060] If the record has been in the data cleaning list, the cleaning flag is cancelled and the record is recovered to the refined retention area.
[0061] The present application provides a human resource data management method based on big data. The method has the following advantages:
[0062] (1) The method introduces a behavior density feature vector and a behavior density score V1, realizes an accurate quantitative determination mechanism based on multi-dimensional behavior characteristics and behavior timeliness, enables the human resource system to dynamically evaluate the activity and behavior value of personnel interaction behavior more in line with actual business rules, preprocesses, standardizes and multi-dimensionally combines operation frequency parameters, operation confirmation event parameters and inspection coverage range parameters, and uses a natural logarithmic decay function to model the time value of behavior, so that the present application can effectively reflect the objective phenomenon that the intensity of personnel behavior decays over time. Compared with the existing behavior determination method which relies on fixed conditions and static rules, the present application can automatically identify high behavior value data when facing a large amount of historical behavior data, avoid deleting records that still have business value, and thus improve the accuracy and reliability of the human resource system in behavior value evaluation.
[0063] (2) The method realizes the dynamic quantitative evaluation of the synchronization consistency and the credibility of personnel data in the multi-terminal environment by constructing the personnel data credibility enhancement score V2, fusing and calculating the behavior density score V1 and the multi-terminal synchronization deviation intensity, and solving the problem that the traditional data cleaning process cannot identify the data records with high behavior value but synchronization difference. The personnel data credibility enhancement score V2 measures the overall deviation intensity of the cross-terminal time offset and the behavior interruption number in the Euclidean deviation structure, and then uses the inverse decay model to suppress the influence of deviation on the data credibility, so that the greater the synchronization deviation, the lower the credibility, and the smaller the synchronization deviation, the closer the credibility to the behavior value itself. By adding the credibility calibration link before the data cleaning trigger, the application can significantly reduce the risk of false deletion caused by the asynchronization of cross-device records, make the system retain more high-value data with actual business contribution, and improve the data credibility and data security of the system.
[0064] (3) The method realizes the intelligent, correlation strength driven data cleaning control of historical human resource data by combining the behavior value judgment mechanism of the behavior density score V1 and the synchronization consistency calibration mechanism of the personnel data credibility enhancement score V2, forming a dynamic layered cleaning system of "behavior value judgment, credibility calibration, and final cleaning decision". The dynamic cleaning system can automatically adjust the cleaning priority according to the real-time correlation strength of personnel data and business behavior: when the behavior density score V1 is higher than the threshold value, it is directly marked as high-value data; when the behavior density score V1 is lower than the threshold value, the personnel data credibility enhancement score V2 is obtained based on the calculation of the synchronization deviation intensity; and finally, the cleaning or recovery decision is output through the secondary determination of the threshold Vth2. Compared with the traditional static cleaning strategy, the application realizes the dynamic, differentiated and intelligent cleaning rules, so that the system can effectively eliminate redundant low-value data under the premise of ensuring data integrity, significantly reduce system storage redundancy, improve retrieval performance and model training data quality, and retain high-value historical behavior data, thereby improving the intelligent level and business effectiveness of human resource data management as a whole. BRIEF DESCRIPTION OF DRAWINGS
[0065] Figure 1 A step schematic diagram of the human resource data management method based on big data of the application;
[0066] Figure 2 A general structure framework schematic diagram of the application. DETAILED DESCRIPTION
[0067] With reference to the accompanying drawings, the technical solutions in the embodiments of the present application will be described clearly and completely. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments of the present application. Based on the embodiments in the present application, all the other embodiments obtained by those ordinarily skilled in the art without creative work fall within the scope of the present application.
[0068] Embodiment 1; the present application provides a big data-based human resource data management method, please refer to Figure 1 and Figure 2 , comprising the following steps:
[0069] S1, by configuring the data source access of the data acquisition interface of a plurality of equipment terminals and human resource systems, collecting personnel interaction behavior data and personnel multi-source synchronization deviation data, forming personnel interaction behavior original data set and personnel multi-source synchronization deviation data set, and preprocessing the personnel interaction behavior original data set in the personnel data correlation analysis module to obtain the behavior density feature vector;
[0070] S2, in the personnel data correlation analysis module, the behavior density feature vector is processed by structural calculation, the behavior density score V1 is constructed, and the behavior density score V1 and the preset behavior density threshold Vth1 are compared preliminarily, and the preliminary comparison result is obtained;
[0071] S3, when the preliminary comparison result automatically triggers the personnel data credibility calibration module, the personnel multi-source synchronization deviation data set is analyzed and processed, the personnel data credibility enhancement score V2 is calculated, and the personnel data credibility enhancement score V2 and the preset credibility judgment threshold Vth2 are compared twice, and the final cleaning judgment result of personnel data is obtained;
[0072] S4, based on the final cleaning judgment result of personnel data, the system-level cleaning control mechanism is started to execute the cleaning rules and data recovery actions, and the cleaning control record and recovery record are output to the data management log module.
[0073] In this embodiment, the method realizes a complete closed-loop processing flow from multi-source data collection, behavior value calculation, credibility calibration to final cleaning decision through the continuous execution of S1-S4. First, S1 performs data source access configuration on multiple industrial equipment terminals and human resource systems, pre-processes personnel behavior data and synchronization deviation data generated by different terminals, and unifies them into a structured form, so that the system can avoid data distortion problems caused by differences in equipment types and inconsistent sampling periods, thereby ensuring that the behavior density feature vector required for subsequent calculation has comparability and calculability. Then in step S2, the behavior density score V1 is constructed by improving the Euclidean structure and time sparsity penalty term, so that the system can effectively distinguish between "current personnel behavior with business value" and "high-frequency behavior with expired time effectiveness", avoiding the misjudgment problem caused by traditional static threshold. Then in S3, the behavior density score V1 is combined with the cross-terminal timestamp offset D and the number of behavior interruptions E to construct the personnel data credibility enhancement score V2, which is used to calibrate "false abnormal" data caused by device clock differences, log loss or record delay in a multi-terminal environment, so that the system can identify data that has actual business value although it is behavior sparse, and avoid false cleaning caused by synchronization errors. Finally, through S4, the system performs differentiated system-level cleaning control using deletion, merging, freezing and recovery strategies, so that data cleaning is no longer "one size fits all", but dynamic cleaning according to the dual judgment of behavior value and credibility, and allows the recovery of important historical data with high credibility when necessary, ensuring that the system reduces data redundancy without losing key business records. Through the above process, the present embodiment effectively reduces redundant data, improves the retention rate of high-value data, avoids misjudgment caused by cross-terminal synchronization deviation, and significantly enhances the precision, stability and business interpretability of human resource data management, achieving a dynamic and intelligent cleaning effect that traditional static cleaning mechanisms cannot achieve.
[0074] Embodiment 2; in particular: S1 includes S11;
[0075] S11, a device-side data sending plug-in is deployed at the system communication layer of the monitoring terminal of the multiple industrial equipment, the device-side data sending plug-in is connected with the log output interface, the clock record interface and the inspection task record interface of the industrial equipment monitoring terminal through a device communication protocol adaptation module, and converts terminal operation data, terminal timestamp data and inspection trajectory data generated by the industrial equipment monitoring terminal into a transmittable format;
[0076] A data access management plug-in is deployed at the data access layer of the human resource system side, and the data access management plug-in is connected with the device-side data sending plug-in through a message transmission channel;
[0077] And set two collection points in each monitoring terminal, the collection points include personnel behavior parameter collection points and multi-source synchronization type parameter collection points; meanwhile, set data structure analysis plug-ins in each collection point; the data structure analysis plug-ins include behavior time analysis plug-ins and synchronization deviation detection plug-ins;
[0078] The personnel behavior parameter collection points are set in the software application layer of the monitoring terminal, and the behavior event analysis plug-ins are installed, the behavior event analysis plug-ins are connected with the operation interface event log module of the industrial equipment monitoring terminal, and the operation frequency information, operation confirmation behavior information and inspection range identification information of personnel on the operation interface of the industrial equipment monitoring terminal are collected;
[0079] The multi-source synchronization type parameter collection points are set in the system bottom layer running environment of the monitoring terminal, and the synchronization deviation detection plug-ins are installed, the synchronization deviation detection plug-ins are connected with the system clock management module and the running state recording module of the monitoring terminal, and the operation timestamp offset information and behavior interruption recording information of personnel between different industrial equipment monitoring terminals are collected.
[0080] S1 further includes S12;
[0081] S12, based on the data structure analysis plug-ins set in the collection points, real-time extraction of personnel interaction behavior original data set and personnel multi-source synchronization deviation data set; and the personnel interaction behavior original data set and the personnel multi-source synchronization deviation data set are packaged into standardized data frames through the device end data sending plug-in, and the standardized data frames are sent to the data access layer of the human resource system side through the device end data sending plug-in, and the data access management plug-in is deployed;
[0082] The personnel interaction behavior original data set includes the operation frequency A of personnel on the device terminal, the operation confirmation event number B submitted by personnel and the independent area number C of personnel participating in regional inspection;
[0083] The personnel multi-source synchronization deviation data set includes the timestamp offset D of personnel recorded between different monitoring terminals and the abnormal interruption number E of personnel performing behavior on the device end.
[0084] Among them, the timestamp offset D of personnel recorded between different monitoring terminals is obtained by difference calculation of the operation time of the same personnel on different monitoring terminals through the synchronization deviation detection plug-in.
[0085] S1 further includes S13;
[0086] S13, after the personnel interaction behavior original data set is received by the human resource data correlation analysis module, the personnel interaction behavior original data set is preprocessed to obtain a behavior density feature vector; the preprocessing includes data integrity preprocessing, time alignment processing and data standardization processing;
[0087] The data integrity preprocessing detects missing values based on the original sampling frequency by the sequence of each parameter in the personnel interaction behavior original data set, and repairs the missing values by the mean value of adjacent time periods.
[0088] The time alignment processing aligns the personnel interaction behavior original data set of the same acquisition frequency to a unified time reference point by the personnel interaction behavior original data set after the data integrity preprocessing.
[0089] The data standardization processing eliminates the unit dimension influence between all parameters in the personnel interaction behavior original data set by the minimum maximum standardization processing method, and then performs secondary aggregation on the personnel interaction behavior original data set after the time alignment processing to obtain the behavior density feature vector.
[0090] In the embodiment, the method realizes stable, structured and computable input of data from different industrial equipment terminals through the step-by-step processing of S11-S13. First, the actual purpose of deploying the device end data sending plug-in in the monitoring terminal system communication layer is to convert logs, clocks and inspection trajectories generated by different equipment manufacturers and different bottom layer protocols into a unified data format. If this conversion link is missing, the original multi-terminal data will not be able to enter the unified analysis process due to inconsistent protocols and different field meanings, resulting in that the behavior characteristics cannot be quantified and the feature vector is distorted. Further, the personnel behavior parameter collection point and the multi-source synchronous parameter collection point are respectively set at the equipment end, and the core role is to avoid the mixing of behavior data and synchronous deviation data. If mixed, it will cause the behavior intensity index and the equipment abnormal index to interfere with each other, so that the subsequent algorithm cannot distinguish between "true weak behavior" and "device abnormality leading to behavior loss", and finally cause misjudgment.
[0091] The data structure analysis plug-in set by S12 enables the system to directly extract complete behavior raw data sets (A, B, C) and synchronization deviation data sets (D, E). The significance of this design lies in that the three parameters of behavior reflect "what the personnel did", and the two parameters of synchronization deviation reflect "whether the equipment records accurately". If the two dimensions are not distinguished, the system cannot determine whether the behavior is weak because the personnel behavior is indeed sparse or because the equipment timestamp drift causes the record to be missing. By packaging the two types of data into a standard format through the data sending plug-in, the data packet loss or out-of-order phenomenon caused by network jitter, device cache differences and other factors can be avoided, thereby maintaining the integrity of the data link. In the preprocessing stage of S13, through missing data repair, time alignment and minimum maximum standardization processing, the multiple behavior parameters originally inconsistent in amplitude, across devices and across sampling periods are converted into a unified and computable feature space. Its real purpose is to eliminate the "pseudo behavior features" introduced by data source differences. For example, the operation frequency of the same personnel on different devices may be different due to the influence of the device UI structure. If not standardized, the device interface difference will be mistaken for personnel behavior difference, leading to distortion of the behavior model. The role of time alignment processing is to avoid the "behavior out-of-order" caused by the writing delay of different devices, thereby ensuring that the behavior density feature vector reflects the true behavior intensity, rather than the device writing delay. In summary, through strict hierarchical design of data acquisition link, data classification extraction, structure analysis and feature preprocessing, the embodiment realizes the real extraction of personnel behavior features and the isolated control of device deviation, so that the behavior density feature vector can accurately express the behavior density of personnel in actual business. The direct effects include: the behavior feature dimension is more pure, the influence of synchronization deviation on behavior calculation is completely stripped off, and the consistency of cross-device behavior data is significantly improved, thereby establishing a solid data foundation for subsequent behavior density score calculation and credibility calibration. Finally, this high-quality data input mechanism not only reduces the misjudgment rate of the subsequent cleaning link, but also significantly improves the judgment accuracy of the system on historical data, making the entire data management process more stable, reliable and interpretable.
[0092] Embodiment 3; in particular: S2 comprises S21;
[0093] S21, based on the behavior density feature vector obtained through preprocessing, the personnel data correlation analysis module adopts a calculation method combining improved Euclidean structure and time sequence sparsity offset factor to calculate and output behavior density score V1 as a measure of personnel interaction behavior activity;
[0094] The behavior density score V1 is calculated and output by the following algorithm formula:
[0095] ; wherein, ln represents a natural logarithm function, and Tr represents the number of days from the current time to the last time the current personnel performed an action on the device;
[0096] In the formula, in order to quantify the activity level of personnel behavior, the personnel data correlation analysis module needs to calculate the behavior density score V1; the calculation of the behavior density score V1 is based on two mathematically explicit structures: one is the Euclidean norm structure derived from the linear algebra field of mathematics; and the other is the time sparsity penalty term composed of the logarithmic function derived from the field of mathematical analysis.
[0097] Firstly, the Euclidean norm is a classic formula for measuring the length of a multidimensional vector and is widely used to measure the size of a multidimensional quantity; in the present scheme, the behavior characteristics of personnel are encapsulated as three-dimensional dimensionless characteristic values, corresponding to the standardized operation frequency characteristic value, the standardized operation confirmation event characteristic value, and the standardized patrol coverage range characteristic value; the behavior density characteristic vector thus formed is an important input for the present scheme to characterize the intensity of personnel behavior; substituting the behavior density characteristic vector into the Euclidean norm structure yields the numerator, and the larger the value, the higher the overall intensity of personnel behavior, representing a higher degree of behavior density; secondly, in order to make the score reflect the time value of behavior,
[0098] The present scheme introduces the time sparsity penalty term 1+ln(1+Tr) composed of the logarithmic function, wherein Tr represents the number of days from the current time to the last time the current personnel performed an action on the device; the logarithmic function ln(1+Tr) has the mathematical property of monotonically increasing and slowly growing, which can simulate the natural decay process of the timeliness of behavior, so that the contribution value is gradually weakened as the behavior is farther away from the current time, but without the sudden drop problem that may be caused by linear decay; therefore, by placing the above Euclidean norm in the numerator and the time sparsity penalty term in the denominator, the behavior density score formula of the present formula can be constructed.
[0099] In the formula, both the numerator and the denominator are dimensionless quantities, so the behavior density score V1 is also a dimensionless result, maintaining the dimensional consistency of the formula. The physical meaning of this combination form is that the numerator of the formula reflects the overall behavior energy of personnel under the multi-dimensional behavior characteristics, while the denominator controls the decay of the timeliness of behavior, thereby constructing a quantitative model of "behavior intensity decay over time", i.e., the more frequent and closer to the current, the higher the behavior density score V1 value; the more sparse and distant, the lower the behavior density score V1 value; the score not only explicitly reflects the mathematical size of the behavior density, but also conforms to the business rule that the "behavior value" weakens over time in the human resource scenario.
[0100] S2 also includes S22;
[0101] S22, the human resource system according to the distribution of historical personnel interaction behavior data, through the behavior density score V1 average value of the behavior value sufficient time, preset as the behavior density threshold Vth1;
[0102] And based on the real-time acquisition of behavior density score V1 and behavior density threshold Vth1, preliminary comparison is carried out to determine whether the behavior value of personnel interaction behavior in the current business cycle reaches the minimum use value level required by the system, so as to decide whether the personnel data needs to enter the next stage of credibility calibration process; The specific comparison content is as follows:
[0103] When the behavior density score V1 is greater than or equal to the behavior density threshold Vth1, the preliminary comparison result is determined as sufficient behavior value, and the personnel related data is marked as normal behavior value data;
[0104] When the behavior density score V1 is less than the behavior density threshold Vth1, the preliminary comparison result is determined as insufficient behavior value, triggering the personnel data credibility calibration module to enter the next stage of deep credibility calibration process.
[0105] In this embodiment, the method S2 realizes the first dynamic filtering mechanism of the personnel behavior value by calculating the behavior density score V1 and preliminarily comparing it with the behavior density threshold Vth1. The reason for using the calculation method of combining the improved Euclidean structure with the time sparsity penalty term is that the simple superposition of behavior parameters such as operation frequency, confirmation event or inspection range cannot truly reflect the "current behavior value" of the personnel. For example, a certain personnel has very intensive operation behavior in a month ago, but has not performed any task recently. If the traditional cumulative behavior counting method is used, it will lead to the misjudgment of the system that the personnel is still a "high value personnel", which is difficult to reflect the time decay effect of the behavior. In order to avoid this error, the multi-dimensional behavior characteristics are fused through the Euclidean norm in the present embodiment, so that the behavior intensity presents a real "behavior energy" in mathematics; and then the behavior contribution exceeding the time window is naturally attenuated through the logarithmic sparsity penalty term, so that the score can reflect the "time decay law" of the behavior, thereby constructing a quantitative model of the behavior naturally decreasing with time. This processing essentially solves the problem of distinguishing between "strong behavior but not timely" and "weak behavior but recently active", so that the system can judge whether the personnel data in the current period has the minimum use value according to the real business requirements. Further, the purpose of setting the behavior density threshold Vth1 is to form an objective basis for the behavior value baseline. The threshold is calculated based on the historical "high value behavior sample", which can avoid the deviation caused by subjective setting of the threshold. For example, if the enterprise business is in the peak period, the behavior frequency will rise as a whole, and the fixed threshold will cause a large number of normal data to be misjudged as low value. The present embodiment dynamically adapts to the business changes through the sample mean, so that the threshold has self-adaptability. When the behavior density score V1 is lower than the threshold, it can be quickly judged that the behavior data has insufficient contribution in the current period, and needs to be further entered into the credibility calibration process; otherwise, it can be directly marked as normal behavior value data. In summary, by calculating the behavior density score V1 and performing preliminary comparison, the present embodiment effectively realizes the pre-judgment of the behavior value, so that the system can filter out most of the low-density behavior data lacking business significance in advance, thereby reducing the subsequent calculation load, reducing the proportion of invalid data participating in calibration, improving the data processing efficiency, and building a reliable behavior value basis for the entire cleaning decision link.
[0106] Embodiment 4; in particular: S3 includes S31;
[0107] S31, after triggering the personnel data credibility calibration module to enter the next stage of the deep credibility calibration process, extracting the time stamp offset D between personnel records in different monitoring terminals and the number of abnormal interruptions E of the personnel in the device end behavior from the personnel multi-source synchronous deviation data set by the personnel data credibility calibration module;
[0108] The behavior density score V1 is calculated by the personnel data credibility calibration module together with the timestamp offset D between personnel records in different monitoring terminals and the number of abnormal interruptions E of the personnel's behavior at the equipment end, to generate a personnel data credibility enhancement score V2, to quantitatively analyze the synchronization consistency level and data credibility of personnel-related data in a multi-terminal environment;
[0109] The personnel data credibility enhancement score V2 is calculated by the following formula:
[0110] In the formula, represents the synchronization deviation intensity;
[0111] In the formula, when the behavior density score V1 does not reach the preset behavior density threshold Vth1, it is necessary to further determine whether the personnel data in a multi-source equipment environment still has sufficient credibility; for this purpose, the personnel data credibility enhancement score V2 is constructed, and its formula structure is derived from a classical model that can be verified in the mathematical discipline, and is improved on the basis of this application;
[0112] First, the behavior density score V1 comes from the formula in the foregoing steps of the application, which is a comprehensive behavior value index constructed based on the Euclidean norm in the field of linear algebra and the logarithmic penalty function in the field of mathematical analysis, used to represent the behavior intensity and timeliness of the personnel in the current business cycle; second, the data offset measure which is derived from the classical Euclidean distance formula, which is a basic method for measuring the difference intensity between two or more characteristics in the mathematical field; the formula inputs the timestamp offset D between personnel records in different monitoring terminals and the number of abnormal interruptions E of the personnel's behavior at the equipment end as two-dimensional deviation characteristics, to measure the overall synchronization deviation intensity by the Euclidean distance;
[0113] In order to make the synchronization deviation have a business rule-compliant impact on the data credibility, the present application further constructs a reciprocal decay term outside the Euclidean deviation structure; this form is derived from the common decay model in mathematics, which has the characteristics that the greater the deviation, the smaller the decay factor, and the lower the credibility; the smaller the deviation, the decay factor tends to 1, so that high-quality data is not excessively weakened; the above behavior value score V1 is combined with the synchronization deviation decay factor to construct the formula;
[0114] All participating quantities in the above formula are dimensionless values, so the personnel data reliability enhancement score V2 is also a dimensionless result, which is consistent with the requirement of dimensional consistency in patent examination; the logical relationship of the formula is clear: the personnel data reliability enhancement score V2 takes the V1 obtained in the previous step as the basis input, calculates the synchronous deviation intensity through the Euclidean deviation structure, and suppresses the deviation through the inverse proportion enhancement model, thereby forming a comprehensive reliability judgment model of "behavior value x data consistency", realizing the two-dimensional fusion of behavior value and data consistency.
[0115] S3 further comprises S32;
[0116] S32, by determining the reliability judgment threshold Vth2 according to the statistical result of the historical multi-terminal synchronous deviation sample, the reliability judgment threshold Vth2 is the minimum acceptable reliability standard value of the personnel data reliability enhancement score V2;
[0117] After the personnel data reliability enhancement score V2 is calculated, the personnel data reliability enhancement score V2 is compared with the reliability judgment threshold Vth2 for the second time by the personnel data reliability calibration module, and the judgment logic of the second comparison processing is:
[0118] When the personnel data reliability enhancement score V2 is greater than or equal to the reliability judgment threshold Vth2, it is determined that the personnel related data is reliable historical data, and the current personnel related data is written into the data refinement reservation area;
[0119] When the personnel data reliability enhancement score V2 is less than the reliability judgment threshold Vth2, it is determined that the personnel related data is unreliable historical data, and the current personnel is added to the data cleaning list and the corresponding cleaning action record is generated.
[0120] In this embodiment, S3 of the method is to calculate the personnel data credibility enhancement score V2 and perform secondary comparison processing, thereby realizing deep calibration of whether the "low behavior value data" has the necessity to continue to be retained. The reason why the credibility calibration process is still needed after the behavior density score V1 does not reach the threshold is that the personnel behavior record often has record deviations such as time stamp drift, log truncation or behavior interruption in a multi-terminal device environment. If the data value is determined only according to V1, the case of "low behavior density but actually incomplete record caused by device deviation" will occur, so that the data that should be retained is misjudged as worthless data. Therefore, after triggering S31, the system independently quantifies the data synchronization consistency by extracting the time stamp offset D and the number of abnormal interruptions E, and the real purpose is to identify "pseudo-low value data" caused by device record abnormalities. The combination of the behavior density score V1 and the synchronization deviation features D and E into the personnel data credibility enhancement score V2 can measure the deviation intensity in the Euclidean deviation, and then suppress the deviation influence in the inverse ratio decay model, so that the system can clearly distinguish between "weak behavior and weak reality" and "weak behavior but device abnormality". For example, if a person switches between two terminals quickly recently, the behavior time stamp difference may be very large, and without the V2 calibration mechanism it will be mistaken for behavior abnormality; and V2 will correctly identify it as a device synchronization error because of the combination structure of behavior intensity and deviation, thereby avoiding incorrect data cleaning. The purpose of setting the credibility judgment threshold Vth2 in S32 is to establish the minimum credibility standard for V2 score, so that the system has a clear "credible and non-credible" demarcation point. This threshold is based on a large number of synchronization deviation sample statistics to avoid errors caused by subjective setting. When V2 is higher than Vth2, even if the behavior density is low, it means that the personnel data has high synchronization consistency in the multi-terminal environment, and should be considered as credible historical data written to the refined retention area; when V2 is lower than Vth2, it means that the synchronization deviation is too high to cause the record to be distorted, and such data should be preferentially entered into the data cleaning list if it continues to be retained, which will cause business analysis errors and behavior model deviation. Through the above mechanism, the present embodiment introduces synchronization deviation calibration when the behavior value is insufficient, so that the system can comprehensively evaluate the effectiveness of the data, and greatly improves the data judgment accuracy in the low behavior density scene. The direct effects include: avoiding data deletion caused by device deviation, enhancing the credibility of historical data, improving the refinement of cleaning strategies, and significantly improving the stability and reliability of the system in processing cross-device data.
[0121] Embodiment 5; in particular: S4 includes S41;
[0122] S41, after receiving the personnel data final cleaning determination result that the personnel related data is untrusted historical data, automatically calling the data cleaning scheduling module, executing the cleaning rules of the system level cleaning control mechanism according to the scheduling instructions of the data cleaning scheduling module on the target personnel data; the cleaning rules include deletion rules, merging rules and freezing rules;
[0123] The deletion rule is determined to be a deletable record when the following conditions are met: the personnel data is not referenced by any system module in the last business cycle; and the behavior density score V1 and the personnel data trustworthiness enhancement score V2 corresponding to the data are both lower than the behavior density threshold Vth1 and the trustworthiness determination threshold Vth2, and the current state is maintained for two consecutive cycles, then the current record is deleted;
[0124] The merging rule is that when multiple data records have the same personnel identifier, the same time window and the content repetition ratio reaches 80%, the repeated records are merged into a single storage unit to reduce redundancy;
[0125] The freezing rule is that when the data record does not meet any of the deletion rule and the merging rule, but the synchronous deviation intensity is greater than 0.8, the current data record is marked as a frozen state, and the data in the frozen state reenters the trustworthiness calibration process in the next cycle and is not deleted temporarily;
[0126] After the system level cleaning control mechanism ends the execution of the cleaning rules, it automatically generates a cleaning control record for each processed data record, including the processing type, processing time, processing result and corresponding personnel identifier, and writes it into the data management log module for audit and backtracking query;
[0127] It should be noted that the data record here represents the summary of the personnel interaction behavior original data set and the personnel multi-source synchronous deviation data set.
[0128] S4 also includes S42;
[0129] S42, when the system level cleaning control mechanism determines that the personnel data is "highly trusted important historical data", it sends a recovery instruction to the data recovery management module, and the data recovery management module executes the data recovery action according to the recovery instruction, removes the data record of the corresponding personnel from the cleaning list and restores it to the data refinement retention area;
[0130] The system level cleaning control mechanism performs consistency verification before executing the data recovery action, and the consistency verification includes:
[0131] The behavior density score V1 is greater than or equal to the behavior density threshold Vth1;
[0132] The personnel data credibility enhancement score V2 is greater than or equal to a credibility determination threshold Vth2;
[0133] Synchronization deviation intensity ≤0.8;
[0134] When all the three conditions are met, the recovery operation can be executed, and the data can be recovered after the verification is passed, so as to ensure that the recovered data does not contain cross-terminal synchronization exception records;
[0135] After the system-level cleaning control mechanism completes the data recovery action, a recovery rule is executed, and the recovery rule is specifically as follows:
[0136] If the target record is in a frozen state, the system will be recovered from the frozen area to the main data area.
[0137] If it has been in the data cleaning list, the cleaning mark is cancelled and recovered to the refined retention area.
[0138] In this embodiment, S4 of the method is realized by a system-level cleaning control mechanism to differentially process personnel data, and the core purpose is not simply to delete or retain data, but to ensure that historical data in the human resource system has rationality in the dual dimensions of "business value + data credibility". First, the reason why the deletion rule requires that the behavior density score VI and the personnel data credibility enhancement score V2 are simultaneously lower than the threshold value and remain unchanged in two consecutive periods is to avoid data being mistakenly deleted due to period fluctuations, short-term device abnormalities or task changes. For example, a short-term off-duty of a certain personnel may cause the behavior score to decrease, but it will recover in the next period. If periodic determination is not increased, there will be a large number of "early deletion" problems, so the multi-period stability is the key to ensure that the deletion action is executed prudently. The setting of the merging rule is for another type of problem. Repetitive data will cause the behavior model to produce false high features, making the system mistakenly believe that the personnel behavior is abnormally active. By identifying the repetitive content of the same personnel in the same time window, records with a repetition ratio of 80% are automatically merged, which can effectively eliminate the "data inflation" problem and ensure that the behavior analysis result is not disturbed by repetitive data. The introduction of the freezing rule has a clear technical significance: a large amount of data is in a critical state of "cannot be deleted and cannot be directly retained", for example, the synchronization deviation intensity is too high but the overall behavior value is not low. Without the freezing strategy, these records will be mistakenly deleted or continue to occupy the system main data area, causing storage redundancy. The freezing mechanism temporarily isolates it from the main data area and recalibrates it in the next period, so that the system can avoid data pollution of the main data space without the risk of deletion, thereby ensuring the sustainable stability of data quality. In the data recovery level, the pre-recovery consistency check in S42 can avoid data pollution caused by "false recovery". The behavior value, credibility and synchronization deviation are indispensable, for example, although a certain data has high behavior intensity, it has serious synchronization deviation (such as cross-terminal time jump of more than 0.8), if directly recovered, it will cause business trajectory analysis error. Therefore, the pre-recovery check is essentially to prevent the system from re-adding abnormal data to the main data area due to recovery operation. The recovery rule further ensures that data in different states can "return to the correct position", such as freezing data returning to the main data area and cleaning queue data being removed from the cleaning list, so that the data life cycle management has reversible and controllable ability. Through the above mechanism, the S4 stage not only completes the traditional cleaning action, but also realizes the intelligent cleaning system of "controllable false deletion, false merging avoidance, and data state traceability", so that the data quality of the whole system is improved, the redundancy is reduced, the model accuracy is enhanced, and the historical data management is changed from static rules to dynamic scheduling, significantly improving the stability and fineness of human resource data governance.
[0139] While embodiments of the application have been shown and described, it is to be understood that the embodiments described are merely exemplary of the principles and application of the present application. Numerous modifications and changes can be made by those skilled in the art without departing from the spirit and scope of the application.
Claims
1. A big data-based human resource data management method, characterized in that: Comprising the following steps: S1, by accessing the data source of a plurality of device terminals and human resource system data acquisition interface configuration, collecting personnel interaction behavior class data and personnel multi-source synchronization deviation class data, forming personnel interaction behavior original data set and personnel multi-source synchronization deviation data set, and preprocessing the personnel interaction behavior original data set in the personnel data correlation analysis module to obtain the behavior density feature vector; S1 includes S11; S11, deploying a device end data sending plug-in in the system communication layer of the monitoring terminal of the plurality of industrial devices, the device end data sending plug-in being connected with the log output interface, clock recording interface and inspection task recording interface of the industrial device monitoring terminal through a device communication protocol adaptation module, and converting terminal operation data, terminal timestamp data and inspection trajectory data generated by the industrial device monitoring terminal into a transmittable format; A data access management plug-in is deployed in the data access layer of the human resource system side, and the data access management plug-in is connected with the device end data sending plug-in through a message transmission channel; And setting two collection points in each monitoring terminal, the collection points including personnel behavior parameter collection points and multi-source synchronization parameter collection points; At the same time, a data structure analysis plug-in is arranged in each collection point; The data structure analysis plug-in includes a behavior time analysis plug-in and a synchronization deviation detection plug-in; The personnel behavior parameter collection point is set in the software application layer of the monitoring terminal, and the behavior event analysis plug-in is installed, the behavior event analysis plug-in is connected with the operation interface event log module of the industrial device monitoring terminal, and the operation frequency information, operation confirmation behavior information and inspection range identification information of the personnel on the operation interface of the industrial device monitoring terminal are collected; The multi-source synchronization parameter collection point is set in the system bottom layer running environment of the monitoring terminal, and the synchronization deviation detection plug-in is installed, the synchronization deviation detection plug-in is connected with the system clock management module and the running state recording module of the monitoring terminal, and the operation timestamp offset information and behavior interruption recording information of the personnel between different industrial device monitoring terminals are collected; S1 also includes S12; S12, based on the data structure analysis plug-in arranged in the collection point, real-time extraction of personnel interaction behavior original data set and personnel multi-source synchronization deviation data set; And the personnel interaction behavior original data set and the personnel multi-source synchronization deviation data set are packaged into standardized data frames by the device end data sending plug-in, and the standardized data frames are sent to the data access layer of the human resource system side by the device end data sending plug-in Data access management plug-in is arranged; The personnel interaction behavior original data set includes the operation frequency A of the personnel in the device terminal, the number B of operation confirmation events submitted by the personnel, and the number C of independent areas of the personnel participating in regional inspection; The personnel multi-source synchronization deviation data set includes the timestamp offset D recorded by the personnel between different monitoring terminals and the number E of abnormal interruptions of the personnel in the device end behavior; S1 also includes S13; S13, after the personnel interaction behavior original data set is received by the human resource data correlation analysis module, the personnel interaction behavior original data set is preprocessed to obtain a behavior density feature vector; the preprocessing includes data integrity preprocessing, time alignment processing and data standardization processing; The data integrity preprocessing detects missing values based on the original sampling frequency through the sequence of each parameter in the personnel interaction behavior original data set, and repairs the missing values by mean interpolation of adjacent time periods; The time alignment processing aligns different time records of the personnel interaction behavior original data set obtained at the same sampling frequency to a unified time reference point through the personnel interaction behavior original data set after data integrity preprocessing; The data standardization processing performs data standardization processing on the personnel interaction behavior original data set after time alignment processing by a minimum maximum standardization processing method, eliminates the unit dimension influence between all parameters in the personnel interaction behavior original data set, and then performs secondary summarization on the processed personnel interaction behavior original data set to obtain a behavior density feature vector; S2, the behavior density feature vector is processed by structural calculation in the personnel data correlation analysis module, a behavior density score V1 is constructed, and the behavior density score V1 is compared with a preset behavior density threshold Vth1 to obtain a preliminary comparison result; S3, when the preliminary comparison result automatically triggers the personnel data credibility calibration module, a personnel multi-source synchronous deviation data set is analyzed and processed to calculate a personnel data credibility enhancement score V2, and the personnel data credibility enhancement score V2 is compared with a preset credibility judgment threshold Vth2 to obtain a final personnel data cleaning judgment result; The S3 includes S31; S31, after triggering the personnel data credibility calibration module to enter the next stage of the deep credibility calibration process, the personnel data credibility calibration module extracts the timestamp offset D between personnel records in different monitoring terminals and the number of abnormal interruptions E of personnel behavior in the device end from the personnel multi-source synchronous deviation data set; Then, the personnel data credibility calibration module calculates the behavior density score V1 and the timestamp offset D between personnel records in different monitoring terminals and the number of abnormal interruptions E of personnel behavior in the device end to generate the personnel data credibility enhancement score V2, and quantitatively analyzes the synchronization consistency level and data credibility of personnel related data in a multi-terminal environment; The S3 also includes S32; S32, the credibility judgment threshold Vth2 is determined according to the statistical result of the historical multi-terminal synchronous deviation sample, and the credibility judgment threshold Vth2 is the minimum acceptable credibility standard value of the personnel data credibility enhancement score V2; After the personnel data credibility enhancement score V2 is calculated by the personnel data credibility calibration module, the personnel data credibility enhancement score V2 is compared with the credibility judgment threshold Vth2, and the judgment logic of the secondary comparison processing is: When the personnel data credibility enhancement score V2 is greater than or equal to the credibility determination threshold Vth2, the personnel-related data is determined to be credible historical data, and the current personnel-related data is written into the data refinement retention area; When the personnel data credibility enhancement score V2 is less than the credibility determination threshold Vth2, the personnel-related data is determined to be untrustworthy historical data, the current personnel is added to the data cleaning list, and a corresponding cleaning action record is generated; S4, based on the personnel data final cleaning determination result, starting the system-level cleaning control mechanism to execute the cleaning rules and data recovery actions, and outputting the cleaning control records and recovery records to the data management log module; The S4 comprises S41; S41, after receiving the personnel data final cleaning determination result that the personnel-related data is untrustworthy historical data, automatically calling the data cleaning scheduling module, and executing the cleaning rules of the system-level cleaning control mechanism on the target personnel data according to the scheduling instructions of the data cleaning scheduling module; the cleaning rules comprise deletion rules, merging rules and freezing rules; The deletion rules are determined to be deletable records when the personnel data meets the following conditions: the personnel data is not referenced by any system module in the latest business cycle; and the behavior density score V1 and the personnel data credibility enhancement score V2 of the data are both lower than the behavior density threshold Vth1 and the credibility determination threshold Vth2, and the current state is maintained for two consecutive cycles, then the current record is deleted; The merging rules are that when multiple data records have the same personnel identifier, the same time window and the content repetition ratio reaches 80%, the repeated records are merged into a single storage unit; The freezing rules are that when the data record does not meet any one of the deletion rules and the merging rules, but the synchronous deviation intensity is greater than 0.8, the current data record is marked as a frozen state, and the data in the frozen state reenters the credibility calibration process in the next cycle and is not deleted temporarily; After the system-level cleaning control mechanism ends the execution of the cleaning rules, it automatically generates a cleaning control record for each processed data record, including the processing type, processing time, processing result and corresponding personnel identifier, and writes it into the data management log module; The S4 further comprises S42; S42, when the personnel data final cleaning determination result is "highly credible important historical data", the system-level cleaning control mechanism sends a recovery instruction to the data recovery management module, and the data recovery management module executes the data recovery action according to the recovery instruction, removes the data record of the corresponding personnel from the cleaning list and restores it to the data refinement retention area; The system-level cleaning control mechanism performs consistency verification before executing the data recovery action, and the consistency verification comprises: The behavior density score V1 is greater than or equal to the behavior density threshold Vth1; The personnel data credibility enhancement score V2 is greater than or equal to the credibility determination threshold Vth2; The synchronous deviation intensity is less than or equal to 0.8; When all three conditions are met, the recovery operation can be executed; After the system-level cleaning control mechanism completes the data recovery action, the recovery rules are executed, and the recovery rules are as follows: If the target record is in the frozen state, the system will restore it from the frozen area to the main data area; If it has been in the data cleaning list, cancel the cleaning mark and restore to the refined retention area. 2.The big data-based human resource data management method of claim 1, wherein: The S2 comprises S21; S21, based on the behavior density feature vector obtained by preprocessing, the personnel data correlation analysis module adopts the calculation method combining improved Euclidean structure and time sequence sparse cancellation factor to calculate and output the behavior density score V1 as the measurement index of personnel interaction behavior activity. 3.The big data-based human resource data management method of claim 1, wherein: The S2 further comprises S22; S22, the human resource system presets the behavior density threshold Vth1 according to the mean value of the behavior density score V1 in the full behavior value condition according to the distribution of historical personnel interaction behavior data; And based on the behavior density score V1 and the behavior density threshold Vth1 obtained in real time, preliminary comparison is carried out, and the specific comparison content is as follows: When the behavior density score V1 is greater than or equal to the behavior density threshold Vth1, the preliminary comparison result is determined as sufficient behavior value, and the personnel related data is marked as normal behavior value data; When the behavior density score V1 is less than the behavior density threshold Vth1, the preliminary comparison result is determined as insufficient behavior value, and the personnel data credibility calibration module enters the next stage of deep credibility calibration process.
Citation Information
Patent Citations
Post competency evaluation system based on virtual reality technology
CN119026971A
Enterprise-level intelligent risk control decision-making system combined with real-time data flow
CN121146507A