People and post matching analysis method in combination with graph neural network, market factors and large model

By combining graph neural network, market factors and large-scale models, the problem of fuzziness and dynamic change capture difficulties in natural language expression in the existing person-position matching algorithm is solved, and more accurate and explainable person-position matching and salary prediction are achieved.

CN120218875APending Publication Date: 2025-06-27SHANTOU UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510259994.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-06
Publication Date
2025-06-27

AI Technical Summary

Technical Problem

The existing personnel job matching algorithm has ambiguity and ambiguity in natural language expression, the results are uninterpreted by the black box process, data quality problems affect analysis results, and it is difficult to capture dynamic changes in job seekers or positions.

Method used

Combining graph neural network, market factors and large models, the self-trained word vector model is used for human-job matching, combining individual, company and market factors to predict, and fine-tuning it through large-scale models to make it more suitable for resume analysis.

Benefits of technology

It improves the accuracy and interpretability of human-position matching, can consider market and individual dynamic changes more comprehensively, and alleviates a large number of human resources problems in human resource management.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120218875A_ABST
    Figure CN120218875A_ABST
Patent Text Reader

Abstract

The invention discloses a person and post matching intelligent analysis method combining a graph neural network, market factors and a large model, and designs a learning method based on machine learning and deep learning by taking job seeker resume information and post requirement information as initial data. And analyzing the post matching degree between the job seeker and the job seeking post, recommending a salary and performing overall evaluation. The invention belongs to the field of artificial intelligence, and relates to a post evaluation intelligent analysis system which can be applied to the fields of job hunting of job seekers, recruitment of companies, resume matching, analysis and application under related situations and the like.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence, and in particular to an intelligent analysis method for person-job matching that combines graph neural networks, market factors, and large models. Background Art

[0002] In today's increasingly complex and volatile business environment, human resource management has become one of the key factors for organizational success. Among them, person-job matching, as an important part of human resource management, plays an irreplaceable role in improving organizational performance, enhancing employee satisfaction, and achieving organizational strategic goals. Traditional job matching often relies on manual resume screening and interviews, which is not only inefficient but also easily affected by subjective factors, resulting in inaccurate recruitment results.

[0003] With the rapid development of artificial intelligence technology, especially the breakthroughs achieved in the fields of natural language processing, machine learning, and deep learning, the application of artificial intelligence in the Internet industry is becoming more and more extensive. In the field of human resources, artificial intelligence technology is used to improve recruitment efficiency and achieve more accurate job matching. How to utilize and develop information and intelligent technology to revolutionize the traditional recruitment method has become a scientific problem with great research value at present.

[0004] To implement an artificial intelligence-based resume analysis system, the most important problem to be solved currently is to effectively extract information from resumes and job information and accurately analyze the extracted information. Currently, the person-job matching algorithm on the Internet mainly relies on technologies such as NLP, machine learning, and data mining. NLP technology is widely used to analyze the text information in job seeker resumes and job descriptions, extract key skills, experiences, and job requirements. The methods include bag-of-words model, TF-IDF, Word2Vec, CNN, RNN, and Transformer, etc. These methods can convert text into numerical vectors for easy calculation and matching; machine learning algorithms are used to analyze the complex relationship between job seekers and positions and predict the accuracy of matching. Common algorithms include decision trees, random forests, support vector machines (SVMs), neural networks, etc.; by mining and analyzing a large amount of recruitment data and job seeker data, potential patterns and trends are revealed to provide data support for the person-job matching algorithm. Through the application of these technologies, the person-job matching algorithm can automatically match the skills and experiences of job seekers with the requirements of job positions, improving the efficiency and accuracy of recruitment. With the continuous development of technology, the person-job matching algorithm will play a more important role in the future.

[0005] However, in the existing technology, there are the following deficiencies: 1. The meaning expressed in natural language often has ambiguity and vagueness, which may lead to misunderstandings when NLP technology is used to parse job seekers' resumes and job descriptions. For example, the same job description may have different interpretations in different companies or contexts.

[0006] 2. The internal mechanisms of many machine learning algorithms, especially deep learning models, are opaque to the outside world, which is the so-called "black box process". This makes it difficult for the present invention to explain the specific reasons for the matching results, affecting the credibility and acceptability of the results.

[0007] 3. The results of data mining and analysis highly depend on the quality of the input data. If there are problems such as errors, omissions, or inconsistencies in the data, then the results of mining and analysis will be severely affected.

[0008] 4. The current person-job matching algorithms mainly analyze and match based on static data, and it is difficult to capture the dynamic changes of job seekers or positions. For example, the skills and experience of job seekers may grow over time, while the requirements of positions may also be adjusted due to market changes. Summary of the Invention

[0009] The technical problem to be solved by the embodiments of the present invention is to provide an intelligent person-job matching analysis method combining graph neural network, market factors and large models, which can self-train a word vector model in the industry field for person-job matching tasks, predict the job level and starting salary of job seekers by combining personal, company and market factors, and cooperate with the large model method to perform industry-specific and domain-specific fine-tuning on the large model to make it more suitable for resume analysis.

[0010] To solve the above technical problems, the embodiments of the present invention provide an intelligent person-job matching analysis method combining graph neural network, market factors and large models, including the following steps: S1: Calculate the matching degree between job seekers and positions, including establishing an Internet job description information library and a resume information library, and establishing the calculation of their matching degree; S2: Obtain industry information data and combine it with job position information and resume information for salary prediction; S3: Extract positions, the matched resumes and their evaluation information from the Internet job description information library to adjust the ChatGLM2 large model; S4: Use the ChatGLM2 large model to comprehensively evaluate the resume and the position.

[0011] Furthermore, the S1 further includes the steps: S110: Clean the Internet job description information library and perform word segmentation to obtain the domain word library required for word vector training; S120: Perform word vector training; S130: Detect the trained word vector model; S140: Obtain resume information and corresponding job description information from the Internet job description information library and the resume information library, and perform data cleaning; S150: Vectorize the cleaned resume information and job information; S160: Calculate the matching degree of the vectorized information.

[0012] Furthermore, the S150 further includes the steps of: Extract skill-specific nouns and project-related words from the resume information and extract skill-specific nouns and project-related words from the job description information , and obtain N word vector representations of the resume and M word vector representations of the job description through iteration, and use a neural network for text feature extraction.

[0013] Furthermore, the S160 further includes the steps of: S1601: Set preconditions, and define the set of required skills in the job requirements as , where represents each ability requirement, and use to represent a resume with n segments of experience, . Use to represent the total set of a job application, , and the label Y represents the recruitment result of the job application; S1602: After generating the feature vectors, calculate the matching degree between each job requirement and each experience, and use the attention-based relationship score to quantify the matching contribution of each ability requirement to each job seeker. The calculation method is:

[0014] where, is the resume vector after RNN learning and the job vector

[0015] added together and activated by tanh to obtain a vector with resume information and job information; where, is the job matching contribution weight;

[0016] Among them, is the matching contribution parameter after ability weighting; In the above, , , are trainable parameters, is the semantic feature representation of the learned ability requirements; Quantify the matching contribution of each candidate's experience to each ability requirement through the following formula:

[0017] Among them, is the resume vector after RNN learning and the job vector , the vector with resume information and job information after adding and passing through the tanh activation

[0018] Among them, is the candidate matching contribution weight;

[0019] Among them, is the matching contribution parameter after candidate ability weighting; Among them, , , are trainable parameters, is the semantic feature representation of the learned ability requirements Add another attention layer to the importance of each representation of the job requirements and experience : :

[0020] Among them, is the job vector after the job passes through the activation layer function;

[0021] Among them, is the calculation weight of the job vector for each ability;

[0022] Among them, is the local job semantic vector representing the job vector; Among them, , are trainable parameters, is the local semantic vector of the job; Calculate each resume experience Importance to generate the final local resume vector , the formula is as follows:

[0023] where is the candidate vector after the resume passes through the activation layer function;

[0024] where is the calculation weight of the candidate vector for each ability;

[0025] where is the local candidate semantic vector representing the resume vector; where , are trainable parameters, is the local semantic vector of the resume; S1603: For each recruitment, establish two undirected graphs, named and , , in Graph J-J, historical positions and current position postings serve as nodes, and the edge set represents their relationship; in Graph J-J, historical employment resumes and current resumes serve as nodes, and the edge set represents their relationship; S1604: The graph matrix learns nodes, calculates the edge weights between every two nodes in the graph, and obtains the matrix about the positions and the matrix about the candidates, and learns the node representation through the update function of the GNN unit:

[0026] where : the aggregated information of node i at the t-th layer;

[0027] where : the update gate, controlling the importance of new information;

[0028] where : the reset gate, controlling the importance of old information;

[0029] Among them, : The candidate node representation combines the current aggregated information and the old information controlled by the reset gate;

[0030] Among them, is the final representation of node i at layer t, is a list of node vectors of the current label graph matrix at layer t, is the i-th row of the matrix corresponding to node i. And As learning parameters, And are the reset and update gate mechanisms respectively; At the same time, a trainable job embedding matrix and a resume embedding matrix are constructed. By looking up the embedding matrices and , the initial representation of the node is obtained. After inputting the two matrices into the corresponding gated graph neural network respectively, the representations of all nodes in Graph J-J and Graph R-R are obtained, denoted as and respectively; S1605: Modeling the experience of recruiters. Define the experience of matching the capabilities presented in the resume with the job posting as the relationship J-R, and define the experience of matching the requirements in the job posting with the resume as the relationship R-J; For the relationship J-R, apply the soft-attention mechanism to map the job embedding vector to the vector space of Graph R-R, and use the attention mechanism to estimate the matching degree between each job and . Associate the job representation in Graph J-J with the experience representation and take the output as the global vector of the resume; For the relationship R-J, apply the soft-attention mechanism to map the historical successful recruitment records of the current job posting in the resume to the Graph J-J vector space, concatenate the resume representation from in Graph R-R with the experience representation from of the relationship R-J, and regard the output as the global information vector of the current resume; S1606: After obtaining the local semantic vectors , and the global information vectors , After that, first, the local semantic vector and the global semantic vector are concatenated, and a comparison mechanism based on a fully connected network is adopted to measure the matching degree between the recruitment information and the resume, and the output of the fully connected network is converted into a Logistic function to obtain the predicted matching degree .

[0031] Furthermore, step S2 further includes steps: S210: Clean the data in the industry database to obtain structured industry information data; S220: Analyze the current salary levels of each position and each level based on the structured industry information data; S230: Analyze the current market saturation of each position and each level based on the structured industry information data; S240: Analyze the current proportions of employment, non-employment, and withdrawal of offers for each position and each level based on the structured industry information data; S250: Integrate the salary level, market saturation, and recruitment proportion into market factors; S260: Clean the job description information and resume information to obtain structured job and resume information; S270: Combine market factors, job information, and resume information for salary prediction.

[0032] Furthermore, step S4 further includes steps: S410: Extract the positions and corresponding resume evaluation information from the existing Internet job description information database, and perform data cleaning to form structured fine-tuning data; S420: Use the structured fine-tuning data to fine-tune the large model to generate a domain fine-tuned large model suitable for the recruitment industry; S430: Extract the information from the job description information and resume information and perform structured processing; S440: Combine the fine-tuned large model to generate evaluation opinions on the positions and corresponding resumes.

[0033] Furthermore, step S420 further includes steps: A method of fine-tuning using the ChatGLM model and the LoRA method; The method of fine-tuning using the LoRA method includes: Select a pre-trained model; design a LoRA layer to introduce additional low-rank matrix parameters on the basis of the pre-trained model; freeze the parameters of the original model, and during the fine-tuning process, the parameters of the original pre-trained model are frozen; train the parameters of the LoRA layer, and only train the introduced low-rank matrix parameters; combine the LoRA layer with the original model: after the fine-tuning is completed, combine the trained LoRA layer parameters with the original pre-trained model to form a new model; During the training process of the ChatGLM model, matrix A is randomly initialized with a Gaussian distribution, and matrix B is initialized with 0, and the parameters are frozen during the training process , only train the parameters in A and B, and after the training is completed, use the parameters in AB and combine them with the parameters in the original model to obtain .

[0034] Furthermore, the S440 further includes the steps of: endowing the large model with basic role situations, specifying the tasks it processes and the output format, putting the resume that needs to generate evaluation opinions and the corresponding positions into the model, and obtaining evaluation opinions.

[0035] Implementing the embodiments of the present invention has the following beneficial effects: The present invention is more professional and comprehensive for the human-post analysis task, and can fully utilize historical recruitment information and self-trained word vectors to improve the accuracy of matching degree calculation; at the same time, combining important market factors into the salary prediction link makes the model more interpretable; finally, using large model technology can alleviate a large number of manpower problems. Description of the Drawings

[0036] Figure 1 Is the overall flowchart of the system corresponding to the intelligent resume analysis method; Figure 2 Is the flowchart of the matching degree calculation module between the position and the job seeker's resume in the intelligent resume analysis method; Figure 3 Is a schematic diagram of the matching degree calculation process; Figure 4 Is the flowchart of the job seeker's salary prediction module in the intelligent resume analysis method; Figure 5 Is the flowchart of the human-post evaluation generation module in the intelligent resume analysis method. Detailed Embodiments

[0037] To make the objectives, technical solutions, and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the accompanying drawings.

[0038] Combined with Figure 1 As shown, the embodiments of the present invention mainly include the following steps: S1: Calculate the matching degree between job seekers and positions, including establishing an Internet job description information database and a resume information database, and establishing the matching degree calculation between the two. S2: Obtain industry information data, combine web position information and resume information for salary prediction. S3: Extract positions, the matching resumes and their evaluation information from the Internet job description information database to adjust the ChatGLM2 large model. S4: Use the ChatGLM2 large model to comprehensively evaluate resumes and positions.

[0039] As Figure 2 shown, the calculation of the matching degree between job seekers and positions should include the following steps: S110: Data cleaning; S120: Word vector training; S130: Domain word vectors; S140: Data cleaning; S150: Text vectorization; S160: Matching degree calculation.

[0040] In the above process, in S110, by cleaning the information in the information database, meaningless symbols and characters are removed, and the cleaned text is tokenized to obtain the domain word library required for word vector training; in S120, the method selected for word vector training is Multi-word context in word2vec, considering multiple words in the context; in S130, the trained word vector model is detected to check its accuracy; in S140, resume information and the corresponding job description information are obtained from the resume and job information databases, and meaningless characters are removed and tokenization is performed on these two pieces of information; in S150, the cleaned resume information and job information are vectorized for the next word vector calculation; in S160, the matching degree of the two vectorized pieces of information is calculated.

[0041] The specific implementation process of S150 is as follows: S150 performs vectorization processing on resume information and job description information. For example, technical specialty nouns and project-related words are extracted from resume information , and technical specialty nouns and project-related words are extracted from job description information . Through iteration, N word vector representations of the resume and the job description The M word vectors represent that the purpose of self-semantic representation is to calculate the semantic similarity between recruitment information and resumes. It is unreasonable to simply regard recruitment information and resumes as simple texts. After obtaining resumes and job requirements, it is very necessary to use neural networks to extract text features from documents. RNN mainly focuses on short sentences when processing long-format documents. It learns document representations from multiple abstract layers of the document structure and can handle long-format documents.

[0042]

[0043]

[0044] Among them, 、 represent the resume vector and the job vector respectively, , ; in which r represents resume, and i represents the i-th resume; in which j represents job, and i represents the i-th job description; represents the dimension of this vector is ; represents that the resume is obtained through feature vectorization to get vector, and this vector is input into the RNN network for learning to obtain the output hidden layer vector; It's the same for

[0045] Existing research on person-job matching mainly focuses on calculating the similarity between job seekers' resumes and recruitment information, without considering the recruiter's experience (i.e., historical successful recruitment records). The method of the present invention estimates the matching between personnel and jobs from the resumes of candidates and relevant recruitment histories. Given a target resume-job pair, local semantic representations are generated through a co-attention neural network and global experience representations are generated through a graph neural network. The final matching degree is calculated by combining these two representations. In this way, historical successful recruitment records are introduced to enrich the features of resumes and job requirements and strengthen the current matching process.

[0046] The matching degree calculation process in S160 is as Figure 3 shown.

[0047] Specifically, by searching for all relevant historical success records, relevant resumes and job descriptions are first obtained. Then, two graphs are constructed for each recruitment, where the obtained resumes and positions are regarded as nodes of the graph. Based on the established graph, the GNN can capture the relationship between the current project node and relevant historical project nodes and generate corresponding node embedding vectors. Finally, the latent node representation is used to generate a global embedding, which is considered as the hidden function of the experience obtained from historical recruitment records.

[0048] S1601 Prerequisites Take a job requirement, and define the set of required skills in the job requirement as , where represents each ability requirement. Use to represent a resume with n segments of experience (e.g., educational experience, competition experience, research papers, and work experience), . Use to represent the total set of job applications, . The label Y represents the recruitment result of the job application, that is, Y = 1 means the resume successfully meets the job description requirements and the candidate enters the interview stage, while Y = 0 means failure. Therefore can be expressed as . However, there is exactly one pair of the same R in P. Next, use to represent the set of historically successful recruitment records and use it as experience to guide new job applications. is 's subset, each label in it is equal to 1.

[0049] S1602 Self-semantic representation After generating the feature vectors, calculate the matching degree between each job requirement and each experience. To capture the semantic similarity between the job and the resume, the co-attention neural network is used to quantify the matching contribution of each candidate's experience to specific ability requirements.

[0050] Here, the present invention uses the attention-based relationship score to quantify the matching contribution of each ability requirement to each job seeker, which can be calculated as follows:

[0051] where, is the resume vector after being learned by the RNN and the job vector , the vector with resume information and job information after addition and being activated by tanh, which is used for subsequent calculations;

[0052] Among them, is the contribution weight to the position matching;

[0053] Among them is the matching contribution parameter after ability weighting; Among them , , are trainable parameters, is the semantic feature representation of the learned ability requirements.

[0054] Similarly, similar to the relationship score , the present invention obtains the relationship score to quantify the matching contribution of the experience of each candidate to each ability requirement.

[0055]

[0056] Among them, is the resume vector after RNN learning and the position vector , the vector with resume information and position information after addition and tanh activation, for subsequent calculation;

[0057] Among them, is the candidate matching contribution weight;

[0058] Among them, is the matching contribution parameter after candidate ability weighting; Among them , , are trainable parameters, is the semantic feature representation of the learned ability requirements.

[0059] Then, the present invention adds another attention layer to the importance of each representation of the position requirement and experience . Specifically, is the semantic local vector of the position, generated by the importance a of each position requirement .

[0060]

[0061] Among them, is the position vector of the position after passing through the activation layer function;

[0062] Among them, is the calculation weight of each ability for the position vector;

[0063] Among them, is the local position semantic vector representing the position vector; Among them and are trainable parameters, is the local semantic vector of the position.

[0064] Similarly, the present invention can calculate the importance of each resume experience to generate the final local resume vector , and the formula is as follows:

[0065] Among them, is the candidate vector after the resume passes through the activation layer function;

[0066] Among them, is the calculation weight of each ability for the candidate vector;

[0067] Among them, is the local candidate semantic vector representing the resume vector; Among them and are trainable parameters, is the local semantic vector of the resume.

[0068] S1603 Construct the graph matrix For each recruitment, two undirected graphs are established, named and , . In Graph J-J, the historical positions and the current position postings serve as nodes, and the edge set represents their relationship; in Graph J-J, the historical employment resumes and the current resume serve as nodes, and the edge set represents their relationship. The present invention describes the construction details of Graph J-J as follows, and the construction process of Graph R-R is similar.

[0069] For the current recruitment pair , the first step in constructing GraphJ-J is to find all the positions related to . The search strategy is to find , which is a set of recruitment information that meets the conditions , constituting , where . After the search, each is marked as a position related to the current position .

[0070] The second step is to calculate the edge weights and construct an adjacency matrix for each graph. First, the discovered by the search strategy and the related positions are used as different single nodes. First, the representation based on BiLSTM is obtained using the attention mechanism, and then the cosine similarity function is used to calculate the similarity between the two nodes to construct the adjacency matrix .

[0071]

[0072] Cosine similarity represents the similarity of vectors using the cosine value of the angle between two vectors in the vector space. The closer the cosine value is to 1, the closer the angle is to 0 degrees, and the more similar the vectors are.

[0073] The cosine similarity between vectors a and b is defined as follows:

[0074] where represent an N-dimensional vector respectively.

[0075] S1604 Graphic Matrix Learning Node embedding Calculate the edge weights between every two nodes in the graph, which means the constructed graph is a complete graph.

[0076] After establishing the graph matrix, two matrices are obtained (positions) and (candidates), which regard each position in GraphJ-J and each resume in Graph R-R as a separate node. The node representation is learned through the update function of the GNN unit as follows:

[0077] where, : the aggregated information of node i at the t-th layer;

[0078] where, : the update gate, controlling the importance of new information;

[0079] Among them, : Reset gate, controlling the importance of old information;

[0080] Among them, : Candidate node representation, combining the current aggregated information and the old information controlled by the reset gate;

[0081] Among them, is the final representation of node i at the t-th layer, is a list of node vectors of the current label graph matrix at the t-th layer, is the i-th row of the matrix corresponding to node i. and As learning parameters, and are the reset and update gate mechanisms respectively. At the same time, the present invention constructs a trainable job embedding matrix and resume embedding matrix , by looking up the embedding matrices and , the initial representation of the node is obtained (such as ). After inputting the two matrices into the corresponding gated graph neural networks respectively, the present invention obtains the representations of all nodes in Graph J-J and Graph R-R respectively, denoted as and .

[0082] S1605 Modeling the experience of recruiters After obtaining the final node vectors in Graph-JJ and R-R, the present invention further extracts higher-level representations for modeling the experience of recruiters. Recruiters accurately know which capabilities in the resumes they have successfully passed highly match this job posting, and which requirements in the job posting highly match this resume. The present invention defines these two kinds of experiences as relationship J-R and relationship R-J respectively.

[0083] Relationship J-R focuses on enhancing the modeling of candidate capabilities. First, the present invention applies the soft-attention mechanism to map the job embedding vector to the vector space of Graph R-R.

[0084]

[0085] Among them, : The attention weight of the i-th job embedding vector;

[0086] Among them, : The representation of the position in the Graph R-R vector space, which reflects the skills that recruiters valued in past successful recruitments; Among them represents each node vector of Graph J-J, and are trainable parameters. Represents the representation of the position in the Graph R-R vector space. It estimates the skills that recruiters valued in past successful recruitments. The present invention uses another attention mechanism to estimate the matching degree between each position and between.

[0087]

[0088] Among them, : The attention weight of the t-th position and the candidate resume;

[0089] Among them, : The attention weight of the t-th position, indicating the matching degree between the job requirements and the current resume capabilities.

[0090]

[0091] Among them, : The experience obtained from the J-R relationship, indicating the experience of the position-resume matching; , and are training parameters, and the attention score can be regarded as the matching degree between the job requirements and the current resume capabilities. Indicates the experience obtained from the J-R relationship.

[0092] Finally, the present invention associates the position representation in Graph J-J with the experience representation and takes the output as the global vector of the resume.

[0093]

[0094] Among them, : The global vector of the resume; and are learning parameters.

[0095] Relation RJ focuses on simulating the personal preferences of recruiters. First, the present invention applies a soft-attention mechanism to map the resume, the current job posting, the historical successful recruitment records) into the Graph JJ vector space.

[0096]

[0097] in, : The attention weight of the i-th resume embedding vector;

[0098] in, : The representation of resumes in Graph JJ vector space reflects the skills that recruiters valued in previous successful recruitments; in Represents each node vector of Graph RR, and It is a trainable parameter. The present invention uses another attention mechanism to estimate the skills that recruiters value in previous successful recruitments. The degree of match between them.

[0099]

[0100] in, : The matching score between the ttth resume and the position;

[0101] in, : The attention weight of the tth resume, indicating the matching degree between the job requirements and the current resume capabilities.

[0102]

[0103] in, : The attention weight of the tth resume, indicating the matching degree between the job requirements and the current resume capabilities.

[0104] , and is the training parameter. It can be seen that the matching degree between the job requirements and the capabilities in the current resume. represents the experience from relation RJ. Finally, the present invention takes The resume shows that the relationship with RJ The experience representation of is concatenated, and the output is regarded as the global information vector of the current resume.

[0105]

[0106] Among them, : The global vector representation of the resume, which synthesizes resume information and matching experience and are the parameters to be learned.

[0107] S1606 Matching Degree Prediction After obtaining the local semantic vector , and the global information vector , first, the local semantic vector and the global semantic vector are concatenated. Then, a comparison mechanism based on a fully connected network is adopted to measure the matching degree between the recruitment information and the resume. Finally, the present invention converts the output of the fully connected network into a Logistic function to obtain the predicted matching degree .

[0108]

[0109] Among them, : Concatenate the local semantic vector and the global semantic vector of the position to form a comprehensive position representation;

[0110] Among them, : Concatenate the local semantic vector and the global semantic vector of the position to form a comprehensive position representation;

[0111] Among them, D: The intermediate vector representing the matching degree between the recruitment information and the resume.

[0112]

[0113] Among them, : The predicted matching degree, with a value range of [0, 1].

[0114] , and are the parameters to be learned. Among them .

[0115] As Figure 4 shown, the specific salary prediction implementation steps in S2 are as follows: S210: Clean the data in the industry database to obtain structured industry information data; S220: Analyze the current salary levels of each position at each level based on structured industry information data; S230: Analyze the current market saturation of each position at each level based on structured industry information data; S240: Analyze the current proportions of employment, non-employment, and withdrawal from employment offers for each position at each level based on structured industry information data; S250: Integrate the salary levels, market saturation, and recruitment proportion into market factors; S260: Clean the data of the job description information and resume information to obtain structured job and resume information; S270: Combine market factors, job information, and resume information for salary prediction.

[0116] The following describes the entire salary prediction process: Establish a salary prediction system Salary prediction is an important issue for many organizations as it directly impacts employees' financial well-being and the organization's competitiveness. Accurately predicting employees' salaries enables organizations to make informed decisions regarding compensation and benefits packages, leading to more equitable salary distribution, improved organizational performance, and increased employee satisfaction. In today's highly competitive job market, accurate salary prediction is becoming increasingly crucial as organizations strive to attract and retain top talent, maintain a positive and productive workforce, and stay ahead in the competition. Predicting employees' salaries can be challenging as it requires considering a wide range of factors, including years of experience, education, skills, job responsibilities, industry trends, etc.

[0117] Salary prediction systems can also help identify and eliminate any potential salary disparities among different employee groups. By providing a more objective and data-driven approach to determining salaries, these systems can help promote fairness and equality in the workplace. Additionally, salary prediction systems are useful for companies in budgeting and financial planning. By having a better understanding of the potential salary ranges for specific jobs, companies can budget accordingly. The practical value of this work is the salary prediction system, which can provide valuable insights and help make more informed decisions regarding compensation, benefiting both employers and employees. Furthermore, with the emergence of machine learning and artificial intelligence, it has become possible to develop sophisticated algorithms and models to analyze vast amounts of data and make accurate salary predictions. These systems use deep learning techniques such as neural networks to analyze large volumes of data and predict salaries based on a wide range of factors. This technology is becoming increasingly prevalent and is poised to play an important role in future human resources and compensation decision-making.

[0118] Salary levels are not only related to an individual's capabilities but also have a close relationship with the enterprise and even society. Affected by many subjective and objective factors, salary prediction is a rather complex calculation process. Moreover, since recruitment information is used here to predict the possible salary for admission, it is also related to the completeness of the recruitment information. Therefore, the quality of feature extraction is crucial for the model. Here, a manual selection method is adopted to extract features, and the feature information can be divided into the following types: 1. Personal factors Among the recruitment information, personal factors include age, gender, region, contact phone number, and Email. Among them, name, contact phone number, and Email belong to unique identification attributes and have no association with salary prediction; age and gender are basic personal attributes, and certain positions may have certain requirements for them. Therefore, they have a certain implicit impact on salary prediction; the region can, to a certain extent, reflect the possible length of stay of a person in a certain place and has an important impact on salary prediction.

[0119] 2. Job factors Among the recruitment information, school factors include educational background and major. Educational background and major can reflect a person's learning ability and basic skills and have an important impact on salary prediction. They are key feature factors in the salary prediction model. The information applied to salary prediction in the job description includes many aspects, such as the salary level that the job can offer, the educational level required in the job, and the matching degree of the required working years and work experience of the candidate in the job.

[0120] 3. Social factors Among the recruitment information, social factors include work experience, work company, and job position. Among them, work experience can reflect an individual's adaptability and proficiency in skills. It is a key feature factor for salary prediction, so it is also used as one of the data features in the salary prediction model. Work company and job position are random factors and have nothing to do with salary prediction.

[0121] Finally, the present invention models the selected attribute items and separately models the resume information, job information, and market factors to form:

[0122]

[0123]

[0124] Among them, a is the personal information of the job seeker, b is the job information, and c is the market factor. The present invention combines them to obtain the combined , and then constructs a deep learning model of the transformer encoder part after the meeting to perform the salary prediction task.

[0125] The specific implementation is as follows: The designed Transformer adopts the classic Encoder-Decoder structure. Since the present invention does not deal with sequence problems and the lengths of the input and output vectors are given, there is no need for position embedding and output embedding. The Encoder and Decoder adopt the same basic structure, and the difference is that the last layer of the Decoder directly generates the prediction result. For each input vector, it is linearly mapped into three vectors: q, k, and v. q and k are used to calculate its Self-Attention map. After passing through Softmax, k is weighted and then input into a three-layer perceptron for fitting.

[0126]

[0127]

[0128]

[0129] Explanation: The input vector xx is respectively mapped into three vectors: Query (Q), Key (K), and Value (V). Where are trainable parameters.

[0130]

[0131] Explanation: The input vector xx is respectively mapped into three vectors: Query (Q), Key (K), and Value (V).

[0132] Explanation: Layer normalization is used to stabilize the training process.

[0133]

[0134] Explanation: Obtain the final salary prediction result.

[0135] Where are trainable parameters.

[0136] The above part is collectively referred to as the Encoder part. There are N Encoders in the transformer. The specific implementation is as follows: The above parts are collectively referred to as the Encoder part. There are N Encoders in the transformer. In the present invention, the outputs of these N Encoders are concatenated to obtain an output with a multi-layer perceptron. Then, a linear layer is used to reduce the length of the vector to 1 to obtain the final salary prediction result.

[0137] As Figure 5 shown, the specific implementation steps of the person-job evaluation module in S4 are as follows: S410: Extract the job and the corresponding resume evaluation information from the existing Internet job description information library, and perform data cleaning to form structured fine-tuning data; S420: Use the structured fine-tuning data to fine-tune the large model to generate a domain fine-tuned large model applicable to the recruitment industry; S430: Extract the information in the job description information and resume information and perform structured processing; S440: Combine the fine-tuned large model to generate evaluation opinions on the job and the corresponding resume.

[0138] The specific implementation of S420 is as follows: Fine-tuning of large models is a common technique in the field of artificial intelligence. It further trains a pre-trained model on a dataset of a specific task to adapt to a new task or domain. In the present invention, the ChatGLM model and the LoRA method are selected to perform the fine-tuning task of the present invention. Fine-tuning based on LoRA (Low-Rank Adaptation) is a fine-tuning method for large language models (such as GPT-3). Its core idea is to introduce additional low-rank matrix parameters on the basis of the pre-trained model and only fine-tune these parameters, so as to achieve the purpose of reducing the fine-tuning cost and computational resource consumption while maintaining the model performance.

[0139] The basic steps of the LoRA fine-tuning method are as follows: Select a pre-trained model: First, select a pre-trained model suitable for the target task. This model has been trained on a large amount of data and has learned rich language representations and knowledge.

[0140] Design the LoRA layer: On the basis of the pre-trained model, introduce additional low-rank matrix parameters. The number of parameters of these low-rank matrices is much smaller than the number of parameters of the original model, so the fine-tuning cost and computational resource consumption can be greatly reduced.

[0141] Freeze the parameters of the original model: During the fine-tuning process, the parameters of the original pre-trained model will be frozen, that is, they will not participate in the fine-tuning process. This can ensure that the knowledge and representation ability of the original model are retained during the fine-tuning process.

[0142] Train the LoRA layer parameters: Only train the parameters of the introduced low-rank matrix. These parameters will be trained on the dataset of the target task to learn the features and patterns suitable for the target task.

[0143] Combine the LoRA layer with the original model: After fine-tuning, combine the trained LoRA layer parameters with the original pre-trained model to form a new model. This new model not only retains the knowledge and representation ability of the original model but also has specific performance and optimization on the target task.

[0144] The present invention uses the data of the resume position evaluation opinions based on the company's HR human resources, the self-written position evaluation data, and the evaluation data combined with the output of the large model as the original fine-tuning data.

[0145] Select ChatGLM as the base model of the present invention when selecting the pre-trained large model. During the training process, matrix A is randomly initialized with a Gaussian distribution, and matrix B is initialized with 0.

[0146]

[0147] Freeze parameters during the training process , only train the parameters in A and B. After training, use the parameters in AB and combine them with the parameters in the original model to obtain h.

[0148]

[0149] Thus, the present invention obtains the large model after fine-tuning.

[0150] S140 is specifically implemented as follows: In the specific implementation of S140, the present invention uses the model after fine-tuning the large model as the basic model and uses the "Prompt" technology to make the model output the analysis results required by the present invention.

[0151] The "Prompt" technology, in large models (such as large models in the field of natural language processing), usually refers to a method of guiding or instructing the model to perform specific tasks. It does not directly modify the parameters of the model but guides the model to generate the desired output by designing appropriate inputs (Prompts).

[0152] In the field of NLP, Prompt is usually used in applications such as text generation, dialogue systems, machine translation, and text summarization. The design of Prompt depends on the understanding of the task and the capabilities of the model. An effective Prompt can significantly improve the performance of the model on specific tasks.

[0153] For large models, the importance of the Prompt technology lies in: Improve the generalization ability of the model: By designing appropriate Prompts, the model can better understand and process unseen data, thereby enhancing its generalization ability.

[0154] Reduce data requirements: In some cases, the use of Prompt technology can reduce the need for large-scale training data, thereby lowering the costs of data collection and processing.

[0155] Simplify model adjustment: Compared with directly modifying the model's parameters, the use of Prompt technology can more flexibly adjust the model's behavior without the need for complex modifications and training of the model.

[0156] First, the present invention endows the large model with basic role scenarios, stipulating the tasks it processes and the output format. This serves as the basic term for the promot of the present invention.

[0157] Then, the resume for which the present invention needs to generate evaluation opinions and the corresponding position are input into the model to obtain the evaluation opinions.

[0158] The present invention is based on a self-trained Word2vec word vector model in a professional field. Through training with a deep learning model on a large amount of text data in the professional field, it generates word vectors that can accurately express the semantics of the professional field, improving the accuracy of matching. At the same time, the job requirement information is transformed into structured data, and a job requirement description graph is constructed. By making full use of the historical information of recruitment, a graph neural network is built. Combining word vectors and graph algorithms, the matching degree between candidates and positions is calculated from both local and global levels, enhancing the comprehensiveness and accuracy of matching.

[0159] The present invention not only considers job requirements and job seeker information but also comprehensively analyzes the dynamic changes in the recruitment market, such as market salary levels, industry development trends, and the urgent need for positions in the market, thereby improving the accuracy of salary prediction. At the same time, the Transformer model has powerful feature extraction and sequence processing capabilities, which can capture the complex relationships in job requirements and job seeker information and consider the influence of market factors, thus enhancing the precision of salary prediction.

[0160] The present invention uses large model technology as the basic model. These models have undergone large-scale pre-training and can capture rich text information and semantic features. By fine-tuning the large model, it can be adapted to data and tasks in specific fields, improving the accuracy of evaluation. Through fine-tuning, the model can learn the vocabulary, grammar, and semantic rules of specific fields, thereby more accurately parsing and understanding resume information and job information.

[0161] The advantages of the present invention are concentrated in three aspects: comprehensiveness, professionalism, and intelligence. Around these three aspects, the present invention will exhibit the characteristic of robustness in various scenarios.

[0162] The present invention first proposes a graph neural network model with historical recruitment information, and trains a professional domain word vector model by specially collecting vocabulary text information in the recruitment field for constructing word vectors of resume information and job information, and fully utilizes this information to analyze local and global information to perform the task of calculating job matching degree; secondly, various factors (such as market factors, personal factors, job factors) are more comprehensively considered, and a deep learning network is constructed to perform the task of salary prediction; finally, combined with large model technology, through fine-tuning technology, fine-tuning information is constructed to fine-tune the large model, and the situation of resumes and jobs is analyzed more comprehensively and intelligently to generate evaluations.

[0163] Compared with traditional human-job analysis systems, the system proposed in this paper is more professional and comprehensive in human-job analysis tasks, and can fully utilize historical recruitment information and self-trained word vectors to improve the accuracy of matching degree calculation; at the same time, important market factors are combined and incorporated into the salary prediction link to make the model more interpretable; finally, the use of large model technology can alleviate a large amount of manpower problems.

[0164] The above-disclosed is only a preferred embodiment of the present invention, and of course it cannot be used to limit the scope of the rights of the present invention. Therefore, equivalent changes made according to the claims of the present invention still fall within the scope covered by the present invention.

Claims

1. An intelligent analysis method for job matching combining graph neural network, market factors and large models, characterized in that: The following steps are involved: S1: Calculate the matching degree between job seekers and positions, including establishing an Internet job description information database and a resume information database, and establishing a matching degree calculation between the two; S2: Obtain industry information data and combine it with network position information and resume information to make salary predictions; S3: extracting positions and matching resumes and evaluation information from the Internet position description information database to adjust the ChatGLM2 model; S4: Use the ChatGLM2 model to conduct a comprehensive evaluation of resumes and positions.

2. The method according to claim 1, characterized in that: The S1 further comprises the steps of: S110: Clean the Internet job description information database and perform word segmentation processing to obtain a domain vocabulary required for word vector training; S120: perform word vector training; S130: Testing the trained word vector model; S140: Obtaining resume information and corresponding job description information from the Internet job description information database and the resume information database, and performing data cleaning; S150: vectorizing the cleaned resume information and job information; S160: Calculate the matching degree of the vectorized information.

3. The method according to claim 2, characterized in that The S150 further comprises the steps of: Extracting skill-related proper nouns and project-related terms from the resume information And extracting skill-related terms and project-related words from the job description information , obtain the resume by iteration N word vector representations and job descriptions M words Vector representation, using neural network for text feature extraction.

4. The method according to claim 3, characterized in that The S160 further comprises the steps of: S1601: Set preconditions and define the required skill set in the job requirements as ,in Each capability requirement is represented by Represents a resume with n experiences. .use To represent the total set of job applications, , label Y represents the recruitment result of the job application; S1602: After generating the feature vector, calculate the matching degree between each job requirement and each experience, using the attention-based relationship score To quantify the matching contribution of each capability requirement to each job seeker, the calculation method is: in, is the resume vector after RNN learning and position vector The vector containing resume information and job information after addition and tanh activation; in, Contribute weights to job matching; in, Contribution parameters for ability-weighted matching; The above , , is a trainable parameter, is the semantic feature representation of the learned capability requirement; The matching contribution of each candidate's experience to each competency requirement is quantified by the following formula: in, is the resume vector after RNN learning and position vector , the vector containing resume information and job information after addition and tanh activation in, Contribute weights to candidate matching; in, Contribute parameters to the matching after weighting the candidates’ abilities; in, , , is a trainable parameter, is the semantic feature representation of the learned capability requirement Add another layer of attention to job requirements and experience The importance of each representation : in, is the position vector after the position passes through the activation layer function; in, The calculation weight of each capability for the position vector; in, is the local job semantic vector representing the job vector; in, , is a trainable parameter, is the local semantic vector of the position; Count each resume experience Importance To generate the final local resume vector , the formula is as follows: in, is the candidate vector after the resume passes through the activation layer function; in, The calculated weights for each capability of the candidate vector; in, is the local candidate semantic vector representing the resume vector; in , is a trainable parameter, is the local semantic vector of the resume; S1603: For each recruitment, create two undirected graphs, named and , , historical and current job postings in Graph JJ As a node, the edge set Represents their relationship; historical employment resume and current resume in Graph JJ As a node, the edge set represents their relationship; S1604: Graph matrix learning node, calculate the edge weights between every two nodes in the graph, and obtain the matrix for the position and the matrix about candidates , the node representation is learned through the update function of the GNN unit: in, : Aggregate information of node i at layer t; in, : Update gate, controlling the importance of new information; in, : Reset gate to control the importance of old information; in, : Candidate node representation, combining the current aggregation information and the old information controlled by the reset gate; in, is the final representation of node i at layer t, is a list of node vectors at level t of the current label graph matrix, is the i-th row of the matrix corresponding to node i. as well as As learning parameters, as well as They are reset and update gating mechanisms; At the same time, build a trainable job embedding matrix and the resume embedding matrix , by finding the embedding matrix and , we get the initial representation of the nodes. After inputting the two matrices into the corresponding gated graph neural networks, we get the representations of all nodes in Graph JJ and Graph RR, which are recorded as and ; S1605: Experience modeling of recruiters, defining the experience of matching the capabilities presented in the resume with the job posting as the relationship JR, ​​and defining the experience of matching the requirements in the job posting with the resume as the relationship RJ; For relation JR, a soft-attention mechanism is applied to map the job embedding vector to the vector space of Graph RR, and an attention mechanism is used to estimate the relationship between each job and The degree of match between them is expressed as With experience Connect them together and use the output as the global vector of resume; For relation RJ, a soft-attention mechanism is applied to map the historical successful recruitment records of the current job posting of the resume into the Graph JJ vector space, and the The resume shows that the relationship with RJ The experience representation of is concatenated, and the output is regarded as the global information vector of the current resume; S1606: After obtaining the local semantic vector , and the global information vector , After that, the local semantic vector and the global semantic vector are first connected, and a comparison mechanism based on a fully connected network is used to measure the matching degree between the recruitment information and the resume. Convert to Logistic function to get prediction match .

5. The method according to claim 1, characterized in that The S2 further comprises the steps of: S210: Clean the industry database data to obtain structured industry information data; S220: Analyze the current salary levels of various positions and levels based on structured industry information data; S230: Analyze the current market saturation of each position and level based on structured industry information data; S240: Analyze the current proportion of those who are hired, not hired, and those who give up recruitment at each level of each position based on structured industry information data; S250: Market factors are integrated based on salary levels, market saturation and recruitment ratio; S260: Perform data cleaning on the job description information and resume information to obtain structured job and resume information; S270: Combine market factors, job information and resume information to make salary forecasts.

6. The method according to claim 1, characterized in that The S4 further comprises the steps of: S410: extracting positions and corresponding resume evaluation information from an existing Internet position description information database, and performing data cleaning to form structured fine-tuning data; S420: Use structured fine-tuning data to fine-tune the large model and generate a field-fine-tuned large model applicable to the recruitment industry; S430: extracting information from the job description and resume information and performing structured processing; S440: Combine the fine-tuned large model to generate evaluation opinions on the positions and corresponding resumes.

7. The method according to claim 6, characterized in that The S420 further includes the steps of: Fine-tuning method using ChatGLM model and LoRA method; Methods for fine-tuning the LoRA method include: Select a pre-trained model; design the LoRA layer, and introduce additional low-rank matrix parameters based on the pre-trained model; freeze the original model parameters. During the fine-tuning process, the parameters of the original pre-trained model are frozen; train the LoRA layer parameters, and only train the introduced low-rank matrix parameters; combine the LoRA layer with the original model: after the fine-tuning is completed, combine the trained LoRA layer parameters with the original pre-trained model to form a new model; During the training of the ChatGLM model, matrix A is randomly initialized with a Gaussian distribution, and matrix B is initialized with 0. The parameters are frozen during training. , only the parameters in A and B are trained. After the training is completed, the parameters in AB are combined with the parameters in the original model to obtain hS.

8. The method according to claim 7, characterized in that The S440 also includes the steps of: assigning basic role information to the large model, specifying the tasks it handles and the output format, placing the resumes and corresponding positions for which evaluation opinions need to be generated into the model, and obtaining evaluation opinions.