A human resources information management system based on big data
Through a human resource information management system based on big data, the feature vectors of applicants and employees are clustered and matching feature values are calculated, which solves the problem of insufficient fairness caused by excessive dependence on hard indicators in the existing technology, and achieves more fair and accurate candidate screening.
Patent Information
- Application Number
- CN202411222743.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-09-02
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2044-09-02
AI Technical Summary
The screening model of excessive reliance on hard indicators in the prior art has led to insufficient fairness and it is difficult to effectively screen the work ability of a large number of candidates.
A human resource information management system based on big data is adopted. By obtaining the information of applicants and employees, the applicant's feature vector and employee feature vector are established, and clustered to obtain a matching cluster cluster. According to the position and proportion of the applicant's feature vector in the cluster cluster, its matching feature value is calculated for human resource information management.
By evaluating the similarity between candidates and existing employees in the company, providing screening indicators from dimensions outside of educational background, improving the fairness and accuracy of screening, solving the shortcomings of excessive dependence on hard indicators in the existing technology.
Smart Images

Figure CN119205051B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of digital management of human resources, and in particular to a human resources information management system based on big data. Background Art
[0002] In today's competitive business environment, larger companies often face the challenge of screening a large number of applicants' resumes during the recruitment process. As the scale of the company expands, the human resources department needs to process a large number of job applicants, which is not only a time-consuming task but also requires a high degree of accuracy and fairness.
[0003] In the field of modern human resource management, facing a large group of job seekers, companies usually adopt a preliminary screening mechanism based on quantitative indicators such as educational background. Although this method is simple in operation, its rationality in terms of fairness is questionable. The screening model that relies too much on hard indicators may ignore those candidates who do not have the advantages of traditional educational background but have excellent actual work ability.
[0004] Therefore, people need a digital human resources solution that can screen massive applicants more fairly and reasonably. Summary of the invention
[0005] The purpose of this invention is to provide a human resources information management system based on big data to solve the following technical problems:
[0006] The fairness of the screening model in the existing technology that relies too much on rigid indicators is relatively insufficient.
[0007] The purpose of the present invention can be achieved through the following technical solutions:
[0008] A human resources information management system based on big data, comprising:
[0009] A first encoding module, used for acquiring information of each applicant and establishing an applicant feature vector for each applicant;
[0010] The second encoding module is used to obtain the information of each employee and establish an employee feature vector for each employee;
[0011] The data clustering module is used to cluster the candidate feature vector and the employee feature vector to obtain multiple matching clusters;
[0012] A matching analysis module, used to obtain a matching feature value of each applicant based on the total number of vectors in the matching clusters where each applicant's feature vector is located and the proportion of employee feature vectors in the matching clusters where each applicant's feature vector is located;
[0013] The management operation module is used to manage the human resource information of multiple applicants based on matching feature values.
[0014] As a further solution of the present invention: clustering the candidate feature vector and the employee feature vector together to obtain multiple matching clusters, including:
[0015] Perform initial clustering on employee feature vectors to obtain multiple first clustering clusters;
[0016] According to the distribution of the number of employee feature vectors in the first cluster, the K value is obtained;
[0017] According to the K value, the candidate feature vector and the employee feature vector are clustered together based on the K-means clustering algorithm to obtain multiple matching clusters.
[0018] As a further solution of the present invention: performing initial clustering on the employee feature vectors to obtain a plurality of first clustering clusters, including:
[0019] Based on the preset neighborhood radius and the preset minimum number of neighborhood points, the employee characteristics are clustered by the DBSCAN clustering algorithm to obtain multiple initial clusters;
[0020] Compare the number of initial clusters with a preset cluster number threshold, and if the number of initial clusters is lower than the first preset cluster number threshold, modify the preset neighborhood radius and the preset minimum number of neighborhood points according to the difference between the number of initial clusters and the first preset cluster number threshold;
[0021] Repeat the steps of clustering employee features by the DBSCAN clustering algorithm based on a preset neighborhood radius and a preset minimum number of neighborhood points, measuring multiple initial clustering clusters, comparing the number of initial clustering clusters with a preset cluster number threshold, and if the number of initial clustering clusters is lower than a first preset cluster number threshold, correcting the preset neighborhood radius and the preset minimum number of neighborhood points according to the difference between the number of initial clustering clusters and the first preset cluster number threshold, until the number of initial clustering clusters exceeds the first preset cluster number threshold, and taking the initial clustering clusters obtained for the last time as multiple first clustering clusters.
[0022] As a further solution of the present invention: according to the difference between the number of initial clusters and the first preset cluster number threshold, the preset neighborhood radius and the preset minimum number of neighborhood points are corrected, including:
[0023] Modify the preset neighborhood radius by the following formula:
[0024] ε'=ε+a×Δn;
[0025] Wherein, ε' is the preset neighborhood radius after correction, ε is the preset neighborhood radius before correction, a is the first preset unit alignment coefficient, and Δn is the difference between the number of initial clustering clusters and the first preset cluster number threshold.
[0026] As a further solution of the present invention: according to the difference between the number of initial clusters and the first preset cluster number threshold, the preset neighborhood radius and the preset minimum number of neighborhood points are corrected, and further comprising:
[0027] The preset minimum number of neighborhood points is modified by the following formula:
[0028] MinPts'=MinPts-b×e Δn ;
[0029] Wherein, MinPts' is the preset minimum number of neighborhood points after correction, MinPts is the preset minimum number of neighborhood points before correction, b is the second preset unit alignment coefficient, and e is a natural constant.
[0030] As a further solution of the present invention: according to the distribution of the number of employee feature vectors in the first cluster, the K value is obtained, including:
[0031] Arrange the plurality of first clustering clusters from large to small based on the number of employee feature vectors included, to obtain a first clustering cluster sequence;
[0032] Calculate the change value of the number of employee feature vectors in each first cluster and the number of employee feature vectors in the previous cluster in the first cluster sequence;
[0033] The sequential position of the first cluster with the largest change value in the first cluster is taken as the K value.
[0034] As a further solution of the present invention: according to the total number of vectors of the matching cluster where each applicant feature vector is located, and the proportion of employee feature vectors in the matching cluster where each applicant feature vector is located, the matching feature value of each applicant is obtained, including:
[0035] The matching feature value of each applicant is calculated by the following formula:
[0036] f=e α×t ×e β×r ;
[0037] Among them, f is the matching feature value of an applicant, t is the total number of vectors in the matching cluster where the applicant's feature vector is located, r is the proportion of employee feature vectors in the matching cluster where the applicant's feature vector is located, α and β are different weight coefficients respectively, and e is a natural constant.
[0038] As a further solution of the present invention: human resource information management of multiple applicants based on matching feature values includes:
[0039] Count the matching feature values of all applicants, and sort multiple applicants from high to low based on the matching feature values to obtain an applicant sequence;
[0040] Eliminate multiple applicants who are at the end of the applicant sequence at a preset ratio.
[0041] Beneficial effects of the present invention:
[0042] The present invention provides a human resource information management system based on big data, which obtains the information of each applicant through a first encoding module and establishes an applicant feature vector for each applicant, obtains the information of each employee through a second encoding module and establishes an employee feature vector for each employee, and then clusters the applicant feature vector and the employee feature vector together through a data clustering module to obtain a plurality of matching clustering clusters, obtains the matching feature value of each applicant according to the total number of vectors of the matching clustering cluster where each applicant feature vector is located and the proportion of employee feature vectors in the matching clustering cluster where each applicant feature vector is located through a matching analysis module, and finally performs human resource information management on a plurality of applicants based on the matching feature values through a management operation module. Compared with the prior art, the present invention evaluates the similarity between an applicant and existing employees in the enterprise by clustering the applicant feature vectors and the employee feature vectors together, thereby digitally expressing the degree of match between an applicant and the enterprise from another dimension other than educational background by matching feature values and taking existing employees as the measurement standard. This provides another reference screening indicator for relevant staff to use clustering as a technical means when dealing with the screening problem of massive applicants, and solves the problem of insufficient fairness of the screening model that relies too much on hard indicators in the prior art. BRIEF DESCRIPTION OF THE DRAWINGS
[0043] The present invention will be further described below in conjunction with the accompanying drawings.
[0044] Figure 1 It is a system structure diagram of the human resources information management system based on big data of the present invention;
[0045] Figure 2 yes Figure 1 Detailed flowchart of the steps performed by the data clustering module in . DETAILED DESCRIPTION
[0046] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0047] See also Figure 1 As shown, the present invention is a human resources information management system based on big data, comprising:
[0048] The first encoding module 110 is used to obtain information of each applicant and establish an applicant feature vector for each applicant;
[0049] A second encoding module 120, for acquiring information of each employee and establishing an employee feature vector for each employee;
[0050] A data clustering module 130 is used to cluster the candidate feature vector and the employee feature vector to obtain a plurality of matching clusters;
[0051] The matching analysis module 140 is used to obtain the matching feature value of each applicant according to the total number of vectors in the matching cluster where each applicant feature vector is located and the proportion of employee feature vectors in the matching cluster where each applicant feature vector is located;
[0052] The management operation module 150 is used to manage the human resource information of multiple applicants based on matching feature values.
[0053] In the above process, the applicant's information can be any information such as age, education, records of activities such as competitions, academic achievements, scores of examination questions preset by the enterprise, etc. Similarly, the employee's information needs to be consistent with the applicant's information, that is, the employee's information is actually the above information when the employee applied for the job in the past. In practice, in order to achieve better recruitment results, the relevant information of more outstanding employees in the enterprise can be selected for use in the present invention.
[0054] Compared with the prior art, the present invention evaluates the similarity between an applicant and existing employees in the enterprise by clustering the applicant feature vectors and the employee feature vectors together, thereby digitally expressing the degree of match between an applicant and the enterprise from another dimension other than educational background by matching feature values and taking existing employees as the measurement standard. This provides another reference screening indicator for relevant staff to use clustering as a technical means when dealing with the screening problem of massive applicants, and solves the problem of insufficient fairness of the screening model that relies too much on hard indicators in the prior art.
[0055] Specifically, combined Figure 2 As shown, in a preferred embodiment, the steps performed by the data clustering module 130 are: clustering the candidate feature vector and the employee feature vector together to obtain multiple matching clusters, specifically including:
[0056] S201, performing initial clustering on employee feature vectors to obtain multiple first clustering clusters;
[0057] S202, obtaining a K value according to the distribution of the number of employee feature vectors in the first cluster;
[0058] S203. Cluster the candidate feature vector and the employee feature vector together based on the K value and the K-means clustering algorithm to obtain a plurality of matching clusters.
[0059] In practice, you can choose to directly cluster the applicant feature vectors and employee feature vectors together, but this method may not get a reasonable number of matching clusters. For the application scenario of this application, too many or too few matching clusters will reduce the matching feature value's ability to distinguish applicants. For example, too few matching clusters means that most applicants will be assigned the same matching feature value and cannot be distinguished. If there are too many matching clusters, it means that the classification is too detailed. This overly detailed division method will make the matching feature value lose its meaning of representing the matching degree between employees and enterprises. That is, at this time, each applicant may only be similar to a few existing employees, rather than similar to the overall level of the company's employees.
[0060] In order to eliminate the above defects, in this embodiment, before clustering the candidate feature vectors and employee feature vectors together, a series of preliminary analyses are performed to obtain a more reasonable number of clusters (i.e., K value in this embodiment). Specifically, in this embodiment, the employees themselves are clustered first to evaluate the distribution of the employees themselves, and then a reasonable number of clusters is obtained on this basis.
[0061] Specifically, the above step S201, performing initial clustering on the employee feature vectors to obtain a plurality of first clusters, specifically includes:
[0062] Based on the preset neighborhood radius and the preset minimum number of neighborhood points, the employee characteristics are clustered by the DBSCAN clustering algorithm to obtain multiple initial clusters;
[0063] Compare the number of initial clusters with a preset cluster number threshold, and if the number of initial clusters is lower than the first preset cluster number threshold, modify the preset neighborhood radius and the preset minimum number of neighborhood points according to the difference between the number of initial clusters and the first preset cluster number threshold;
[0064] Repeat the steps of clustering employee features by the DBSCAN clustering algorithm based on a preset neighborhood radius and a preset minimum number of neighborhood points, measuring multiple initial clustering clusters, comparing the number of initial clustering clusters with a preset cluster number threshold, and if the number of initial clustering clusters is lower than a first preset cluster number threshold, correcting the preset neighborhood radius and the preset minimum number of neighborhood points according to the difference between the number of initial clustering clusters and the first preset cluster number threshold, until the number of initial clustering clusters exceeds the first preset cluster number threshold, and taking the initial clustering clusters obtained for the last time as multiple first clustering clusters.
[0065] Since the goal of this embodiment is to find a reasonable number of clusters, the DBSCAN clustering algorithm is used for clustering during the initial clustering in this embodiment. The DBSCAN clustering algorithm does not need to specify the number of clusters in advance, and is suitable for the application scenario of this embodiment. However, the DBSCAN clustering algorithm needs to specify the neighborhood radius and the minimum number of neighborhood points in advance, and the effects of the two will affect the clustering results. Therefore, this embodiment uses a loop iteration method to cluster multiple times, and corrects the preset neighborhood radius and the preset minimum number of neighborhood points according to the previous clustering results each time the clustering is re-clustered, so as to ensure that the number of the first clustering clusters is not too small, thereby ensuring that the number of K values is not too small.
[0066] Specifically, in a preferred embodiment, the step in the above process: correcting the preset neighborhood radius and the preset minimum number of neighborhood points according to the difference between the number of initial clusters and the first preset cluster number threshold, specifically includes:
[0067] Modify the preset neighborhood radius by the following formula:
[0068] ε'=ε+a×Δn;
[0069] Wherein, ε' is the preset neighborhood radius after correction, ε is the preset neighborhood radius before correction, a is the first preset unit alignment coefficient, and Δn is the difference between the number of initial clustering clusters and the first preset cluster number threshold.
[0070] Furthermore, in a preferred embodiment, the step of the above process: correcting the preset neighborhood radius and the preset minimum number of neighborhood points according to the difference between the number of initial clusters and the first preset cluster number threshold value, further includes:
[0071] The preset minimum number of neighborhood points is modified by the following formula:
[0072] MinPts'=MinPts-b×e Δn ;
[0073] Wherein, MinPts' is the preset minimum number of neighborhood points after correction, MinPts is the preset minimum number of neighborhood points before correction, b is the second preset unit alignment coefficient, and e is a natural constant.
[0074] In the DBSCAN algorithm, if the number of clusters is found to be insufficient, this usually means that the current parameter settings cause the algorithm to treat many points corresponding to the vectors as noise, or to merge multiple actual clusters into a larger cluster. In this embodiment, the preset neighborhood radius and the preset minimum number of neighborhood points are adjusted in the manner corresponding to the above formula to solve this problem. Specifically, the significance of the above two formulas is:
[0075] Increasing the value of the preset neighborhood radius can expand the neighborhood range of each point so that more points are included in the neighborhood of the core point. This may result in more points being identified as boundary points or core points, thus helping to form more clusters.
[0076] Reducing the preset minimum number of neighborhood points can make it easier for the algorithm to identify points as core points, thus helping to form more clusters. This is because a lower preset minimum number of neighborhood points means a smaller density requirement, and the algorithm will consider points to be core points within a smaller neighborhood, making it easier to form new clusters.
[0077] In addition, the most important thing is that, when the difference between the number of initial clusters and the first preset cluster number threshold is constant, the adjustment of the preset neighborhood radius and the preset minimum number of neighborhood points based on the above formula can make the change amplitude of the preset minimum number of neighborhood points faster than the change amplitude of the preset neighborhood radius. This will reduce the possibility of overlap between clusters and reduce the occurrence of blurred cluster boundaries, thereby not affecting the calculation of subsequent matching feature values.
[0078] Furthermore, in a preferred embodiment, the above step S202, obtaining the K value according to the distribution of the number of employee feature vectors in the first cluster, specifically includes:
[0079] Arrange the plurality of first clustering clusters from large to small based on the number of employee feature vectors included, to obtain a first clustering cluster sequence;
[0080] Calculate the change value of the number of employee feature vectors in each first cluster and the number of employee feature vectors in the previous cluster in the first cluster sequence;
[0081] The sequential position of the first cluster with the largest change value in the first cluster is taken as the K value.
[0082] The significance of the above process is actually to select an effective first cluster based on the elbow rule to ensure the rationality of the number of clusters during co-clustering. This embodiment is mainly used to ensure that the number of first clusters is not too large, thereby ensuring that the number of K values is not too large.
[0083] Further, in a preferred embodiment, the steps performed by the matching analysis module 140 include: obtaining the matching feature value of each applicant according to the total number of vectors in the matching cluster where each applicant feature vector is located and the proportion of employee feature vectors in the matching cluster where each applicant feature vector is located, specifically including:
[0084] The matching feature value of each applicant is calculated by the following formula:
[0085] f=e α×t ×e β×r ;
[0086] Among them, f is the matching feature value of an applicant, t is the total number of vectors in the matching cluster where the applicant's feature vector is located, r is the proportion of employee feature vectors in the matching cluster where the applicant's feature vector is located, α and β are different weight coefficients respectively, and e is a natural constant.
[0087] The significance of this embodiment is to evaluate the matching degree of the applicant from two aspects: the total number of vectors in the matching clusters where each applicant's feature vector is located, and the proportion of employee feature vectors in the matching clusters where each applicant's feature vector is located, so as to obtain a more objective evaluation result. For example, if the total number of vectors in the matching clusters where an applicant's feature vector is located is high, it means that there are many applicants or employees similar to the applicant, but it does not mean that the applicant is more matched with the company. At this time, the proportion of employee feature vectors in the matching clusters where the applicant's feature vector is located should be further measured to determine whether there are more applicants or employees who are similar to the applicant, so as to obtain a reasonable and accurate matching feature value.
[0088] After obtaining the matching feature values, human resource management can be performed in any manner according to the actual situation. For example, in a preferred embodiment, the steps performed by the management operation module 150 include: performing human resource information management on multiple applicants based on the matching feature values, specifically including:
[0089] Count the matching feature values of all applicants, and sort multiple applicants from high to low based on the matching feature values to obtain an applicant sequence;
[0090] Eliminate multiple applicants who are at the end of the applicant sequence at a preset ratio.
[0091] The present invention provides a human resource information management system based on big data, which obtains the information of each applicant through a first encoding module and establishes an applicant feature vector for each applicant, obtains the information of each employee through a second encoding module and establishes an employee feature vector for each employee, and then clusters the applicant feature vector and the employee feature vector together through a data clustering module to obtain a plurality of matching clustering clusters, obtains the matching feature value of each applicant according to the total number of vectors of the matching clustering cluster where each applicant feature vector is located and the proportion of employee feature vectors in the matching clustering cluster where each applicant feature vector is located through a matching analysis module, and finally performs human resource information management on a plurality of applicants based on the matching feature values through a management operation module. Compared with the prior art, the present invention evaluates the similarity between an applicant and existing employees in the enterprise by clustering the applicant feature vectors and the employee feature vectors together, thereby digitally expressing the degree of match between an applicant and the enterprise from another dimension other than educational background by matching feature values and taking existing employees as the measurement standard. This provides another reference screening indicator for relevant staff to use clustering as a technical means when dealing with the screening problem of massive applicants, and solves the problem of insufficient fairness of the screening model that relies too much on hard indicators in the prior art.
[0092] The above is a detailed description of an embodiment of the present invention, but the content is only a preferred embodiment of the present invention and cannot be considered to limit the scope of implementation of the present invention. All equivalent changes and improvements made within the scope of the present invention should still fall within the scope of the patent coverage of the present invention.
Claims
1. A human resources information management system based on big data, characterized in that: include: A first encoding module, used for acquiring information of each applicant and establishing an applicant feature vector for each applicant; The second encoding module is used to obtain the information of each employee and establish an employee feature vector for each employee; The data clustering module is used to cluster the candidate feature vector and the employee feature vector to obtain multiple matching clusters; A matching analysis module, used to obtain a matching feature value of each applicant based on the total number of vectors in the matching clusters where each applicant's feature vector is located and the proportion of employee feature vectors in the matching clusters where each applicant's feature vector is located; A management operation module is used to manage the human resource information of multiple applicants based on matching feature values; Cluster the candidate feature vector and employee feature vector together to obtain multiple matching clusters, including: Perform initial clustering on employee feature vectors to obtain multiple first clustering clusters; According to the distribution of the number of employee feature vectors in the first cluster, the K value is obtained; According to the K value, the candidate feature vector and the employee feature vector are clustered together based on the K-means clustering algorithm to obtain multiple matching clusters; Perform the initial clustering on the employee feature vectors to obtain multiple first clusters, including: Based on the preset neighborhood radius and the preset minimum number of neighborhood points, the employee characteristics are clustered by the DBSCAN clustering algorithm to obtain multiple initial clusters; Compare the number of initial clusters with a preset cluster number threshold, and if the number of initial clusters is lower than the first preset cluster number threshold, modify the preset neighborhood radius and the preset minimum number of neighborhood points according to the difference between the number of initial clusters and the first preset cluster number threshold; Repeat the steps of clustering employee features by the DBSCAN clustering algorithm based on a preset neighborhood radius and a preset minimum number of neighborhood points, measuring multiple initial clustering clusters, comparing the number of initial clustering clusters with a preset cluster number threshold, and if the number of initial clustering clusters is lower than a first preset cluster number threshold, correcting the preset neighborhood radius and the preset minimum number of neighborhood points according to the difference between the number of initial clustering clusters and the first preset cluster number threshold, until the number of initial clustering clusters exceeds the first preset cluster number threshold, and taking the initial clustering clusters obtained for the last time as multiple first clustering clusters.
2. The human resources information management system based on big data according to claim 1 is characterized in that: According to the difference between the number of initial clusters and the first preset cluster number threshold, the preset neighborhood radius and the preset minimum number of neighborhood points are modified, including: Modify the preset neighborhood radius by the following formula: ; Wherein, ε' is the preset neighborhood radius after correction, ε is the preset neighborhood radius before correction, a is the first preset unit alignment coefficient, and Δn is the difference between the number of initial clustering clusters and the first preset cluster number threshold.
3. The human resources information management system based on big data according to claim 2 is characterized in that: According to the difference between the number of initial clusters and the first preset cluster number threshold, the preset neighborhood radius and the preset minimum number of neighborhood points are modified, and the method further includes: The preset minimum number of neighborhood points is modified by the following formula: ; Wherein, MinPts' is the preset minimum number of neighborhood points after correction, MinPts is the preset minimum number of neighborhood points before correction, b is the second preset unit alignment coefficient, and e is a natural constant.
4. The human resources information management system based on big data according to claim 1 is characterized in that: According to the distribution of the number of employee feature vectors in the first cluster, the K value is obtained, including: Arrange the plurality of first clustering clusters from large to small based on the number of employee feature vectors included, to obtain a first clustering cluster sequence; Calculate the change value of the number of employee feature vectors in each first cluster and the number of employee feature vectors in the previous cluster in the first cluster sequence; The sequential position of the first cluster with the largest change value in the first cluster is taken as the K value.
5. The human resources information management system based on big data according to claim 1 is characterized in that: According to the total number of vectors in the matching clusters where each candidate feature vector is located, and the proportion of employee feature vectors in the matching clusters where each candidate feature vector is located, the matching feature value of each candidate is obtained, including: The matching feature value of each applicant is calculated by the following formula: ; Among them, f is the matching feature value of an applicant, t is the total number of vectors in the matching cluster where the applicant's feature vector is located, r is the proportion of employee feature vectors in the matching cluster where the applicant's feature vector is located, α and β are different weight coefficients respectively, and e is a natural constant.
6. The human resources information management system based on big data according to claim 1 is characterized in that: Human resource information management for multiple applicants based on matching feature values, including: Count the matching feature values of all applicants, and sort multiple applicants from high to low based on the matching feature values to obtain an applicant sequence; Eliminate multiple applicants who are at the end of the applicant sequence at a preset ratio.
Citation Information
Patent Citations
Manipulator and post matching method and system based on knowledge graph deep learning
CN116596494A
Project planning management method and system based on big data
CN118313620A