A talent information analysis method based on deep learning

Through the knowledge graph enhancement method based on large models and graph neural networks, the limitations of existing talent information analysis systems in processing unstructured text and complex skill associations are solved, efficient talent data analysis and recommendation are achieved, and the accuracy and explainability of recruitment decisions are improved.

CN120278688BActive Publication Date: 2025-10-14SICHUAN TALENT DEVELOPMENT GROUP CO LTD +2
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510349965.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-24
Publication Date
2025-10-14
Estimated Expiration
2045-03-24

AI Technical Summary

Technical Problem

Existing talent information analysis systems have limitations in processing unstructured text, complex skill associations, and dynamic corporate needs, making it difficult to provide companies with sufficient and accurate recruitment decision support.

Method used

A knowledge graph enhancement method based on large models and graph neural networks is used to perform in-depth feature extraction and cluster analysis on massive talent information, and to make efficient matching recommendations based on corporate needs.

Benefits of technology

It improves the representation ability and analysis accuracy of talent data, can accurately divide talent groups in high-dimensional feature space, and provide efficient recruitment decision support.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120278688B_ABST
    Figure CN120278688B_ABST
Patent Text Reader

Abstract

The application discloses a talent information analysis method based on deep learning, and relates to the technical field of human resource management, comprising the following steps: candidate data collection; generating a high-dimensional feature vector of the candidate; constructing a knowledge graph of the field to generate an enhanced feature vector of the candidate; dividing the candidate into multiple groups by using a clustering algorithm, and defining the category name of clustering for each group; generating a demand feature vector according to the demand of the enterprise and selecting a group; and sorting and recommending the candidate in the group based on similarity calculation. The application is based on the organic combination of a large model, a knowledge graph and clustering analysis, can fully excavate the skill connection and development potential hidden in massive data, has high scalability in each link of feature representation and matching degree calculation, and expands the application potential of talent data analysis in intelligent recruitment, career planning and human resource management and the like.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of human resource management, and particularly relates to a talent information analysis method based on deep learning. BACKGROUND

[0002] In recent years, with the continuous acceleration of enterprise digitization and the diversification of recruitment needs, how to efficiently and accurately mine and match massive talent data has become an important challenge for enterprise talent management. Deep semantic analysis and knowledge graph enhancement based on large models have shown great potential in intelligent recommendation, data mining and other fields. However, existing talent information analysis systems often have certain limitations in processing unstructured text, complex skill association and dynamic enterprise needs, making it difficult to provide sufficient and accurate recruitment decision support for enterprises.

[0003] Graph Feature Enhancement Method, also known as Knowledge Graph Augmented Representation, is a method that uses knowledge graphs to enhance data representation and understanding. It combines traditional representation learning techniques (such as word embeddings, graph neural networks, etc.) with the structured information of knowledge graphs, aiming to enhance semantic understanding of data. SUMMARY

[0004] The purpose of the present application is to provide a knowledge graph enhancement method based on large models and graph neural networks, which can extract deep features and perform clustering analysis on massive talent information, and realize efficient matching recommendation by combining enterprise demand vectorization, thus providing a talent information analysis method based on deep learning.

[0005] To achieve the above purpose, the technical solution adopted by the present application is as follows: a talent information analysis method based on deep learning, comprising the following steps:

[0006] S1, candidate data collection;

[0007] N candidates related to the recruitment position are marked in turn as ~ M types of data of each candidate are collected, where the nth candidate The mth type of data is , and the data set of is obtained , 1≤n≤N, 1≤m≤M,

[0008] S2, generating a high-dimensional feature vector of the candidate ;

[0009] using a large model to obtain Mapping to d-dimensional space to obtain a mapping feature and generate a high-dimensional feature vector of according to the following formula ;

[0010] ,

[0011] wherein, is a corresponding fusion weight, and ;

[0012] S3, constructing a knowledge graph of the field to generate an enhanced feature vector, including steps S31-S34;

[0013] S31, collecting various public data related to the field of the position to constitute a knowledge graph dataset, defining |V| entities related to the position ~ , defining |R| entity relationships ~ , constituting an entity set , a relationship set ;

[0014] S32, identifying entities and extracting entity relationships from the knowledge graph dataset to generate a knowledge graph K=(V,E), E being an entity relationship set , , being any two entities in V, being , an entity relationship of

[0015] S33, generating a knowledge graph enhanced feature of the candidate using a graph feature enhancement method ;

[0016] S34, calculating the enhanced feature vector of according to the formula , constituting an enhanced feature vector set of the candidate , wherein, , are respectively trainable weights of , ;

[0017] S4, processing using a clustering algorithm to divide the candidate into Z groups: defining a cluster category name for each group;

[0018] S5, the enterprise generates a job description according to the demand for employing people, and maps it to a d-dimensional space to obtain a demand feature vector , and selects a group according to the demand for employing people;

[0019] S6, calculate the cosine similarity between each enhanced feature vector in the group and , and arrange the candidates in descending order of the corresponding cosine similarity and recommend them to the enterprise.

[0020] As a preferred: in S1, the m types of data include structured data and unstructured data;

[0021] The structured data includes educational background, work experience, professional skills, language ability, training experience, internship experience, awards and honors obtained, job-seeking intention and expected salary;

[0022] The unstructured data includes skill description, project experience description, personal statement;

[0023] The large model is a pre-trained BERT model.

[0024] As a preferred: in S2, According to the following formula;

[0025] ,

[0026] In the formula, is The mapping feature of the kth data in , 1≤k≤M, is a pre-trained scoring function for calculating attention score, is the exp function.

[0027] As a preferred: the public data is derived from professional websites, industry reports, company announcements and professional forums.

[0028] As a preferred: in S31, the entities include candidates, skills, positions, industries and educational institutions, and the entity relationships include needs, belongs and provides;

[0029] In S32, entity recognition is performed by a named entity recognition method, and the relationship between two entities is identified by a relationship extraction model.

[0030] As a preferred: step S4 includes S41-S42;

[0031] S41, divide into Z clusters by a clustering algorithm, and form a group of candidates corresponding to the enhanced feature vectors in each cluster to obtain Z groups ~ ;

[0032] S42, analyze each group and manually define the category name of each group.

[0033] As preferred: in S6, introduce dimension weight when calculating cosine similarity, enhance feature vector The cosine similarity of is obtained according to the following formula:

[0034] ,

[0035] ,

[0036] In the formula, is the dimension number of , is the component in the pth dimension, , , is the L2 distance, is the dimension weight of the p dimensions, which is obtained by presetting or training.

[0037] Compared with the prior art, the advantages of the present application are:

[0038] (1) After obtaining various data of the candidate, multi-dimensional feature extraction of large model and knowledge graph is introduced, wherein the large model obtains a high-dimensional feature vector of the candidate, and the knowledge graph based on the field can generate knowledge graph enhanced features of the candidate, and the two are fused to obtain an enhanced feature vector of the candidate. This method not only can efficiently process and analyze massive talent data, but also can improve the representation ability of talent data, ensure the accuracy and explainability in complex semantic scenarios.

[0039] (2) Based on the enhanced feature vector, clustering analysis is performed to accurately divide the talent group in the high-dimensional feature space, helping enterprises to quickly locate the concerned group. For example, if an enterprise needs to find talents with "machine learning" background, clustering can help the enterprise to focus on those groups in the clustering that have the feature.

[0040] (3) On the basis of clustering, the enterprise demand is vectorized to , and the cosine similarity between the strong feature vector of the candidate in the concerned group and is calculated to accurately sort and screen the candidate.

[0041] In summary, the present application based on the organic combination of large models, knowledge graphs and cluster analysis significantly improves the quality and practicality of talent information analysis and recommendation, and provides an efficient and reliable solution for enterprise talent management, job allocation and strategic planning. Compared with the prior art, the method not only can fully tap the skill connection and development potential implied in massive data, but also has high scalability in each link of feature representation and matching degree calculation, expanding the application potential of talent data analysis in intelligent recruitment, career planning and human resource management and other aspects. BRIEF DESCRIPTION OF DRAWINGS

[0042] Figure 1 The flowchart of the present application. DETAILED DESCRIPTION

[0043] The present application will be further described below in conjunction with examples and drawings.

[0044] Example 1: see Figure 1 A talent information analysis method based on deep learning, comprising the following steps:

[0045] S1, candidate data collection;

[0046] N candidates related to the recruitment position are marked in turn as ~ M types of data of each candidate are collected, wherein the nth candidate The mth type of data is , and the data set of is obtained , 1≤n≤N, 1≤m≤M,

[0047] S2, generating a high-dimensional feature vector of the candidate ;

[0048] Map to d-dimensional space by using a large model to obtain mapping features , and generate a high-dimensional feature vector of according to the following formula ;

[0049] ,

[0050] In the formula, is the corresponding fusion weight, and ;

[0051] S3, constructing a knowledge graph of the field to generate an enhanced feature vector, comprising steps S31-S34;

[0052] ​​S31, collect various public data related to the position to form a knowledge graph dataset, and define |V| entities related to the position ~ , define |R| entity relationships ~ , form an entity set , a relationship set ;

[0053] S32, identify entities and extract entity relationships from the knowledge graph dataset to generate a knowledge graph K=(V,E), E is an entity relationship set , 、 , for any two entities in V, is the entity relationship of 、 ;

[0054] S33, generate the knowledge graph enhancement features of the candidate using a graph feature enhancement method ;

[0055] S34, calculate the enhanced feature vector of according to the formula , to form an enhanced feature vector set of the candidate , wherein , 、 are trainable weights of 、 respectively;

[0056] S4, process using a clustering algorithm to divide the candidates into Z groups: , define the class name of the cluster for each group;

[0057] S5, the enterprise generates a job description according to the demand for hiring, maps it to a d-dimensional space using a large model to obtain a demand feature vector , and selects a group according to the demand for hiring;

[0058] S6, calculate the cosine similarity between each enhanced feature vector in the group and , and arrange the candidates in descending order of the corresponding cosine similarity and recommend them to the enterprise.

[0059] In this embodiment, the m types of data described in S1 include structured data and unstructured data;

[0060] The structured data includes educational background, work experience, professional skills, language ability, training experience, internship experience, awards and honors obtained, job-seeking intention and expected salary;

[0061] The unstructured data includes skill descriptions, project experience descriptions, and personal statements.

[0062] The large model is a pre-trained BERT model.

[0063] In S2, is calculated according to the following formula;

[0064] ,

[0065] In the formula, is the mapping feature of the kth type of data in the middle, 1≤k≤M, is a pre-trained scoring function for calculating attention scores, is an exp function.

[0066] In S31, the source of the public data includes professional websites, industry reports, company announcements, and professional forums; the entities include candidates, skills, positions, industries, and educational institutions, and the entity relationships include needs, belong to, and provide.

[0067] In S32, entity recognition is performed by a named entity recognition method, and the relationship between two entities is identified by a relationship extraction model.

[0068] Step S4 includes S41-S42;

[0069] In S41, the is divided into Z clusters by a clustering algorithm, and the candidates corresponding to the enhanced feature vectors in each cluster form a group, obtaining Z groups ~ ;

[0070] In S42, each group is analyzed, and the class name of each group is manually defined.

[0071] In S6, the dimension weight is introduced when calculating the cosine similarity, and the cosine similarity between the enhanced feature vector and is obtained according to the following formula;

[0072] ,

[0073] ,

[0074] In the formula, is the dimension number of is the component in the pth dimension, , is calculated as the L2 distance,​ dimensional dimension weight, by pre-setting or training.

[0075] Embodiment 2: see Figure 1 There are the following application scenarios:

[0076] A company is recruiting a data scientist, and the human demand generates a job description: the candidate has the following skills and experience: proficient in Python programming, familiar with machine learning and deep learning frameworks (such as TensorFlow, PyTorch), with more than 3 years of data analysis experience, good team cooperation ability and innovative thinking. At the same time, the company has a high requirement for the cultural adaptation of the candidate, hoping that the candidate can integrate into the rapidly changing work environment and play an active role in cross-department cooperation.

[0077] A talent information analysis method based on deep learning, comprising the following steps:

[0078] S1, candidate data acquisition;

[0079] Various channels to obtain the resume of the candidate, such as recruitment platform recommendation, job recruitment, internal recommendation, etc. For each candidate, obtain their M-class data, including structured data, unstructured data, etc.

[0080] S2, use a pre-trained large model such as BERT model, QWEN2 model to generate a high-dimensional feature vector for each candidate. The large model can be regarded as a nonlinear mapping function, which maps the data to a d-dimensional space; for example, for Map the features to the mapping features of the encoder layer in the hidden space using the encoder of the large model This process can be represented by the following formula:

[0081] ,

[0082] where is the encoder mapping process, which is generally implemented by multiple layers of Transformer. After multiple layers of Transformer, we get . Since there are multiple sources of data for a talent, it is necessary to fuse the data representations into a unified feature vector. Weighted average, attention weighting or splicing mapping methods can be used. For example, use weighted average:

[0083] ,

[0084] After this step, the candidate obtains a high-dimensional feature vector :

[0085] S3, build the knowledge graph of the field, generate the enhanced feature vector, including steps S31-S34;

[0086] S31, collect various public data related to the field of the position to constitute the knowledge graph dataset, define |V|=5 entities related to the position ~ , which are respectively the candidate, the skill, the position, the industry and the education institution, define |R|=3 kinds of entity relationships ~ , which are respectively need, belong and provide;

[0087] S32, identify entities and extract entity relationships from the knowledge graph dataset to generate the knowledge graph K=(V,E), wherein the entity recognition adopts the named entity recognition (NER) method, and the relationship extraction adopts the relationship extraction model. For example, a skill is Python, a position is data scientist, and the relationship between them is need, then an element in the entity relationship set can be expressed as (Python, need, data scientist);

[0088] S33, generate the knowledge graph enhanced features of the candidate by the graph feature enhancement method ;

[0089] S34, calculate the enhanced feature vector of according to the formula , constitute the enhanced feature vector set of the candidate , wherein , , are respectively , trainable weights of

[0090] S4, process by using the clustering algorithm to divide the candidates into Z groups: , define the category name of the cluster for each group;

[0091] S41, in this embodiment, the clustering algorithm is K-Means clustering, and more specifically, it is divided into Sa1-Sa4;

[0092] Sa1, select Z initial cluster centers ;

[0093] Sa2, calculate the distance of each candidate to each cluster center, and assign it to the nearest cluster;

[0094] Sa3, process Then, Z new clusters are obtained, and for each cluster, the cluster center is updated as the mean of all members of the cluster;

[0095] Sa4, repeating Sa2-Sa3 until the cluster allocation is no longer significantly changed or the objective function converges; Z clusters are obtained, and the candidates in the same cluster corresponding to the enhanced feature vectors constitute the same group, and Z groups are obtained ~ ;

[0096] S42, analyze each group, and manually define the category name of each group, such as machine learning, data analysis, etc. The clustering result can help the enterprise to identify which group has stronger machine learning skills and which group performs more outstanding in the field of data analysis;

[0097] S5, the job description at the beginning of the embodiment is vectorized by a large model to obtain a demand feature vector , and this time the recruitment focuses on the candidate's "machine learning" skill, so the group with the category name "machine learning" is selected;

[0098] S6, calculate the cosine similarity of each enhanced feature vector in the group and , and arrange the candidates in descending order according to the corresponding cosine similarity and recommend them to the enterprise.

[0099] The above only describes the preferred embodiments of the present application and is not intended to limit the present application. Any modification, equivalent replacement and improvement made within the spirit and principle of the present application shall be included in the protection scope of the present application.

Claims

1. A talent information analysis method based on deep learning, characterized by: The following steps are included: S1, candidate data collection; Mark the N candidates related to the recruitment position as T1~T N , collect M types of data for each candidate, where the nth candidate T n The mth type of data is , get T n Data collection , 1≤n≤N, 1≤m≤M, S2, generate candidate T n The high-dimensional feature vector of ; Use a large model Mapping to d-dimensional space to obtain mapping features , and generate T according to the following formula n The high-dimensional feature vector of ; , Where, for The corresponding fusion weight, and ; S3, constructing a knowledge graph of the domain and generating an enhanced feature vector, including steps S31 to S34; S31, collect various public data related to the job to form a knowledge graph dataset, and define |V| entities related to the job ~ , define |R| entity relationships ~ , forming an entity set and relationship sets ; S32, identify entities and extract entity relationships from the knowledge graph dataset to generate a knowledge graph K=(V,E), where E is the entity relationship set , v i and v j For any two entities in V, r ij v i and v j Entity relationships; S33, generating candidates using graph feature enhancement methods Knowledge graph enhancement features ; S34, according to the formula ,calculate The enhanced feature vector , which constitutes the set of enhanced feature vectors of the candidate , where 、 They are 、 The trainable weights of S4, processed using clustering algorithm , divide the candidates into Z groups: , define the category name of the cluster for each group; S5, the enterprise generates job descriptions based on employment needs, and uses the large model to map it to the d-dimensional space to obtain the demand feature vector , and select a group based on employment needs; S6, calculate the enhanced feature vector of each group and The candidates are ranked in descending order according to their cosine similarity and recommended to the companies. In S1, the m types of data include structured data and unstructured data; The structured data includes educational background, work experience, professional skills, language proficiency, training experience, internship experience, awards and honors received, job search intentions, and expected salary; The unstructured data includes skill descriptions, project experience descriptions, and personal statements; The large model is a pre-trained BERT model.

2. The talent information analysis method based on deep learning according to claim 1, characterized in that: In S2, Calculate according to the following formula; , Where, D n The kth type of data Mapping characteristics, 1≤k≤M, is the pre-trained scoring function used to calculate the attention score, and exp(⋅) is the exp function.

3. The talent information analysis method based on deep learning according to claim 1, characterized in that: The public data mentioned above are sourced from career websites, industry reports, company announcements and professional forums.

4. The talent information analysis method based on deep learning according to claim 1, characterized in that: In S31, the entities include candidates, skills, positions, industries, and educational institutions, and the entity relationships include need, belong to, and provide; In S32, entity recognition is performed using a named entity recognition method, and the relationship between two entities is identified using a relationship extraction model.

5. The talent information analysis method based on deep learning according to claim 1, characterized in that: Step S4 includes S41-S42; S41, through clustering algorithm Divide into Z clusters, and form a group of candidates corresponding to the enhanced feature vectors in each cluster, and obtain Z groups C1~C Z ; S42, analyze each group and manually define the category name of each group.

6. The talent information analysis method based on deep learning according to claim 1, characterized in that: In S6, dimension weights are introduced when calculating cosine similarity to enhance the feature vector With F R The cosine similarity is obtained according to the following formula; , , Where, for The number of dimensions, for The component in the pth dimension, , To calculate the L2 distance, w p is the dimension weight of p dimensions, w p Obtained by presetting or training.

Citation Information

Patent Citations

  • Occupational development planning method based on time sequence knowledge graph

    CN115455205A

  • Recruitment talent portrait generation method and system based on multi-dimensional information

    CN119205054A