Online recruitment reciprocity bilateral recommendation method and system and medium
Through heterogeneous two-part graph neural network and adversarial learning technology, the problem that job seekers and recruiters' preferences in online recruitment are not fully considered, and more accurate matching prediction is achieved, which improves the effectiveness of the recruitment system.
Patent Information
- Application Number
- CN202510392295.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-31
- Publication Date
- 2025-07-01
AI Technical Summary
The existing online recruitment recommendation system fails to fully consider the preferences of both job seekers and recruiters, resulting in poor matching results and traditional methods cannot effectively handle complex interactive information between job seekers and recruiters.
Using heterogeneous two-part graph neural network and adversarial learning technology, a convolutional neural network is used to encode the job seeker's resume and job recruitment position text to construct an interactive graph between job seeker and job recruitment position, using generator and discriminator to optimize the embedded representation, and combining multi-layer perceptron to calculate the matching probability.
Two-way modeling between job seekers and recruiters is realized, the accuracy of matching and recommendation relevance is improved, and more accurate matching predictions between job seekers and positions are provided.
Smart Images

Figure CN120235598A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of online recruitment, and specifically relates to an online recruitment reciprocal bilateral recommendation method, system and medium. Background Art
[0002] With the rapid development of artificial intelligence technology, online recruitment platforms have been widely used globally. Taking Zhaopin, the largest recruitment platform in China, as an example, the platform has more than 360 million registered users and more than 120 million monthly active users. According to the prediction of Insight Partners, the scale of the global online recruitment market will grow from 29.29 billion US dollars in 2021 to 47.31 billion US dollars in 2028. The rise of online recruitment has made job seekers and recruiters face a huge amount of resume and job information in their daily work, resulting in the deficiencies of the traditional manual matching mode. In the vast amount of information, it is often difficult for job seekers and recruiters to quickly find suitable matching objects that meet their own needs.
[0003] Existing online recruitment systems mainly rely on traditional single recommendation methods, including collaborative filtering-based recommendation systems, content-based recommendations, and hybrid recommendation methods. These methods mostly focus on one-way matching according to the resume data of job seekers or job requirements. For example, recommending jobs through resumes, or recommending resumes according to job requirements. However, these methods do not fully consider the preferences of both job seekers and recruiters, and usually ignore the bilateral interaction and demand matching in the recruitment process.
[0004] In the actual application of online recruitment, the matching between job seekers and recruiters not only needs to consider the attribute characteristics of jobs and resumes, but also should pay attention to the needs and preferences of both parties. Recruitment positions and resumes are two different objects, and they show heterogeneous characteristics in the actual interaction process. In this case, traditional models based on single nodes or one-way matching cannot effectively process the complex interaction information between job seekers and recruiters.
[0005] To address this challenge, in recent years, graph neural networks (GNNs) have received extensive attention because of their ability to effectively process graph-structured data. However, traditional graph neural network methods are mainly applied to graphs with single-type nodes, ignoring the characteristics of heterogeneous graphs. In the recruitment task, job seekers and recruitment positions represent different types of nodes respectively, and the relationship between them constitutes a typical bipartite graph structure. Although existing research applies graph neural network models to recommend jobs or resumes, most of the work only considers one-sided needs and fails to achieve two-way reciprocal matching, resulting in a significant reduction in the effectiveness and accuracy of the recommendation results.
[0006] In addition, although there have been advancements in graph-based recommendation models for handling heterogeneous graph structures, challenges such as the cold start problem, data sparsity, and insufficient information fusion still exist. To overcome these issues, some current studies attempt to introduce adversarial learning methods, but most of these methods focus on learning between single types of nodes or edges and fail to effectively combine the diverse preferences of job seekers and recruiters.
[0007] Therefore, there is an urgent need for a new online recruitment recommendation model that can not only comprehensively consider the needs of both job seekers and recruiters but also effectively capture the complex interaction relationships between them, thereby providing accurate decision-making support for job seekers and recruiters. Summary of the Invention
[0008] The present invention provides an online recruitment reciprocal bilateral recommendation method, system, and medium, aiming to solve the problem that the existing recruitment recommendation system fails to fully consider the preferences of both job seekers and recruiters, resulting in poor matching effects.
[0009] To achieve the above objective, the first aspect of the present invention provides an online recruitment reciprocal bilateral recommendation method, including the following steps:
[0010] Obtain job seeker resume data and recruitment position data;
[0011] Preprocess the job seeker resume data to obtain the job seeker resume text, and use a convolutional neural network to design a resume text encoder to encode the job seeker resume text into a resume text vector;
[0012] Preprocess the recruitment position data to obtain the recruitment position text, and use a convolutional neural network to design a position text encoder to encode the recruitment position text into a position text vector;
[0013] Construct a heterogeneous bipartite recruitment graph, where the nodes represent job seeker resumes and recruitment positions respectively, and the edges represent the interaction relationships between job seeker resumes and recruitment positions;
[0014] According to the obtained resume text vector, position text vector, and their interaction relationships, calculate the initial graph embedding representations of job seeker resume nodes and recruitment position nodes to obtain the initial resume graph embedding and the initial position graph embedding;
[0015] Use the generator to perform information transfer on the obtained initial resume graph embedding and initial position graph embedding, and generate the embedding representation from the job seeker resume to the recruitment position and the embedding representation from the recruitment position to the job seeker resume respectively through the generator;
[0016] Use a discriminator to perform adversarial learning on the generated embeddings obtained, distinguish the information between the embedding representations from the job seeker's resume to the job posting and from the job posting to the job seeker's resume, and optimize the embedding representation of the generator;
[0017] Based on the optimized embedding representation, update the initial resume graph embedding and the initial job graph embedding to obtain the updated resume graph embedding and the updated job graph embedding;
[0018] Concatenate the obtained resume text vector, job text vector, updated resume graph embedding, and updated job graph embedding to generate a comprehensive matching vector between the job seeker and the job;
[0019] Use a multi-layer perceptron to process the comprehensive matching vector, calculate the matching probability between the resume and the job; output the matching probability value between the resume and the job as the prediction result of the match between the job seeker and the job posting.
[0020] Furthermore, the method for encoding the job seeker's resume text into a resume text vector includes:
[0021] Perform word segmentation on the job seeker's resume text to obtain each word in the resume;
[0022] Map each word after word segmentation to the corresponding word embedding vector;
[0023] Use a one-dimensional convolutional neural network to perform a convolution operation on the obtained word embedding vectors to extract local features in the resume text;
[0024] Perform pooling on the convolution result through a max pooling layer to retain the most important feature information;
[0025] Input the result after pooling into an average pooling layer to perform an aggregation process on various skill information in the resume text to obtain a resume text vector.
[0026] Furthermore, the method for encoding the job posting text into a job text vector includes:
[0027] Perform word segmentation on the job posting text to obtain each word in the job description;
[0028] Map each word after word segmentation to the corresponding word embedding vector;
[0029] Use a one-dimensional convolutional neural network to perform a convolution operation on the obtained word embedding vectors to extract local features in the job text;
[0030] Perform pooling on the convolution result through a max pooling layer to retain the most important feature information in the job requirements description;
[0031] Input the result after pooling into the average pooling layer to comprehensively process the requirements of the job text and obtain the job text vector.
[0032] Furthermore, the method for constructing the heterogeneous bipartite recruitment graph includes:
[0033] Take the job seeker's resume and the recruitment position as two different types of node sets in the graph, which are respectively represented as the job seeker node set and the position node set;
[0034] Model the interaction between each job seeker's resume and the recruitment position. If there is a historical interaction between the job seeker and the position, establish an edge between the job seeker node and the position node;
[0035] Generate the adjacency matrix of the heterogeneous bipartite recruitment graph. Among them, the adjacency matrix represents the interaction relationship between the recruitment position and the job seeker's resume, as well as between the job seeker's resume and the recruitment position. And only cross-node type edges are allowed between nodes, and connections between the same type of nodes are not allowed;
[0036] Fill the adjacency matrix according to the obtained interaction information between the job seeker and the position, and finally construct the heterogeneous bipartite recruitment graph.
[0037] Furthermore, the initial resume graph embedding and the initial position graph embedding are obtained through the following method:
[0038] According to the obtained resume text vector and the obtained job text vector, use them as the initial feature representations of the job seeker nodes and the position nodes;
[0039] Use the constructed heterogeneous bipartite recruitment graph to combine the obtained resume text vector and position text vector with the graph structure information of the job seeker nodes and the position nodes;
[0040] Use the graph embedding method to process the node features and the adjacency matrix, and recursively update the embedding representations of the job seeker nodes and the position nodes;
[0041] Through graph convolution operations, obtain the initial resume graph embedding and the initial position graph embedding.
[0042] Furthermore, the method for the generator to generate the embedding representation from the job seeker's resume to the recruitment position and the embedding representation from the recruitment position to the job seeker's resume includes:
[0043] Combine the initial resume graph embedding with the adjacency matrix of the position graph, and obtain the embedding representation from the resume to the position through graph convolution operations;
[0044] Take the embedding representation from the job seeker's resume to the recruitment position as the input, and use the propagation mechanism of the generator to transfer it to the recruitment position node to generate the embedding representation of the recruitment position;
[0045] Embed the initial job graph into the adjacency matrix of the resume graph, and obtain the embedding representation from the job to the resume through graph convolution operations;
[0046] Take the embedding representation from the job to the job seeker's resume as input, and use the propagation mechanism of the generator to pass it to the job seeker's resume node to generate the embedding representation of the job seeker's resume;
[0047] Obtain the embedding representation from the job seeker's resume to the recruitment position and the embedding representation from the recruitment position to the job seeker's resume.
[0048] Furthermore, the method for optimizing the embedding representation of the generator includes:
[0049] Use the discriminator to classify the embedding representation generated by the generator. The discriminator receives the embedding representation from the job seeker's resume to the recruitment position and the embedding representation from the recruitment position to the job seeker's resume generated by the generator, and judges its authenticity;
[0050] The discriminator calculates the discriminant loss by maximizing its ability to distinguish the generated embedding from the job seeker's resume to the recruitment position from the real job node embedding, and the ability to distinguish the generated embedding from the recruitment position to the job seeker's resume from the real resume node embedding;
[0051] According to the feedback of the discriminator, calculate the loss function of the generator, and update the parameters of the generator according to the loss function;
[0052] Update the parameters of the generator through the backpropagation algorithm;
[0053] Repeat the process of the above steps to continuously optimize the embedding representation of the generator.
[0054] Furthermore, the method for using a multi-layer perceptron to process the comprehensive matching vector, calculate the matching probability between the resume and the position, and output the matching probability value between the resume and the position includes:
[0055] Input the obtained comprehensive matching vector into the multi-layer perceptron. The multi-layer perceptron includes at least one hidden layer, and each hidden layer is transformed through a non-linear activation function to extract high-level features in the comprehensive matching vector;
[0056] Through each layer in the multi-layer perceptron, gradually perform a linear transformation on the comprehensive matching vector, and introduce non-linearity through the activation function;
[0057] In the last layer, use the sigmoid activation function to map the output of the multi-layer perceptron to obtain the matching probability between the resume and the position;
[0058] Output The sigmoid function takes the calculated matching probability value as the prediction result of the matching between the resume and the position, and is used to evaluate the matching degree between the job seeker and the position.
[0059] To achieve the above object, a first aspect of the present invention provides an online recruitment reciprocal bilateral recommendation system, and the recommendation system includes the following modules:
[0060] A data acquisition module for acquiring job seeker resume data and recruitment position data;
[0061] A resume data preprocessing module for preprocessing the job seeker resume data to obtain a job seeker resume text, and designing a resume text encoder through a convolutional neural network to encode the job seeker resume text into a resume text vector;
[0062] A position data preprocessing module for preprocessing the recruitment position data to obtain a recruitment position text, and designing a position text encoder through a convolutional neural network to encode the recruitment position text into a position text vector;
[0063] A heterogeneous bipartite graph construction module for constructing a heterogeneous bipartite recruitment graph, where the nodes respectively represent job seeker resumes and recruitment positions, and the edges represent the interaction relationships between job seeker resumes and recruitment positions;
[0064] An initial graph embedding calculation module for calculating the initial graph embedding representations of job seeker resume nodes and recruitment position nodes according to the obtained resume text vectors, position text vectors, and their interaction relationships, to obtain an initial resume graph embedding and an initial position graph embedding;
[0065] A generator module for using the generator to perform information transfer on the obtained initial resume graph embedding and initial position graph embedding, and respectively generating an embedding representation from the job seeker resume to the recruitment position and an embedding representation from the recruitment position to the job seeker resume;
[0066] A discriminator module for using the discriminator to perform adversarial learning on the obtained generated embeddings, distinguish the information of the embedding representation from the job seeker resume to the recruitment position and the embedding representation from the recruitment position to the job seeker resume, and optimize the embedding representation of the generator;
[0067] An embedding update module for updating the initial resume graph embedding and the initial position graph embedding based on the optimized embedding representation to obtain an updated resume graph embedding and an updated position graph embedding;
[0068] A matching vector generation module for splicing the obtained resume text vectors, position text vectors, updated resume graph embedding, and updated position graph embedding to generate a comprehensive matching vector between the job seeker and the position;
[0069] A matching probability calculation module, which is used to process the comprehensive matching vector using a multi-layer perceptron, calculate the matching probability between the resume and the position, and output the matching probability value between the resume and the position as the prediction result of the matching between the job seeker and the recruitment position.
[0070] To achieve the above object, a first aspect of the present invention provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is run by a processor, it executes the steps of the online recruitment reciprocal bilateral recommendation method.
[0071] Advantages of the present invention:
[0072] Compared with the prior art, an online recruitment reciprocal bilateral recommendation method, system and medium provided by the present invention innovatively realizes the two-way modeling of the preferences of both job seekers and recruiters by introducing a heterogeneous bipartite graph neural network and adversarial learning technology. Specifically, the present invention first encodes the job seeker's resume text and the recruitment position text respectively through a convolutional neural network (CNN) to obtain a resume text vector and a position text vector; then constructs a heterogeneous bipartite recruitment graph to model the interaction relationship between job seekers and positions, calculates the initial graph embedding representation, and further optimizes the embedding representation through the adversarial learning of the generator and the discriminator. In this way, the model can not only capture the feature information of the job seeker's resume and the recruitment position, but also fully consider the preferences of both job seekers and recruiters, and optimize each other's matching through information transmission. This reciprocal recommendation mechanism effectively solves the problem that the traditional one-way recommendation method cannot meet the needs of both job seekers and recruiters at the same time, thereby improving the accuracy of matching and the relevance of recommendations. Finally, by combining a multi-layer perceptron to process the comprehensive matching vector, the matching probability between the job seeker and the position can be predicted more accurately, providing a more accurate recommendation result. Description of the Drawings
[0073] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required to be used in the description of the embodiments.
[0074] Figure 1 It is a flowchart of an online recruitment reciprocal bilateral recommendation method disclosed in an embodiment of the present invention.
[0075] Figure 2 It is an overall framework diagram of an HBRBR model disclosed in an embodiment of the present invention.
[0076] Figure 3 It is a node degree distribution diagram on a job and resume data set disclosed in an embodiment of the present invention.
[0077] Figure 4It is an experimental performance graph of the HBRBR model and the comparison model disclosed in the embodiments of the present invention on different evaluation indexes.
[0078] Figure 5 It is a performance graph of an ablation study disclosed in the embodiments of the present invention.
[0079] Figure 6 It is a performance graph of the HBRBR model under different depths l disclosed in the embodiments of the present invention.
[0080] Figure 7 It is a performance graph of the HBRBR model with different learning rates disclosed in the embodiments of the present invention.
[0081] Figure 8 It is a performance graph of the HBRBR model with different weight decays disclosed in the embodiments of the present invention.
[0082] Figure 9 It is a performance graph of the HBRBR model under different batch sizes disclosed in the embodiments of the present invention. Detailed implementation manners
[0083] In order to enable those skilled in the art of the present technology to better understand the solution of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0084] According to the embodiments of the present invention, it should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions, and although the logical order is shown in the following methods, in some cases, the steps shown or described can be executed in a different order than here.
[0085] As Figure 1 shown, the present invention provides an online recruitment reciprocal bilateral recommendation method, including the following steps:
[0086] Step S100, obtain job seeker resume data and recruitment position data;
[0087] The job seeker resume data and recruitment position data come from a large online recruitment platform. The content in the dataset includes the basic information, educational background, work experience, skills, expected positions, etc. of job seekers, and the recruitment position data includes the position name, company information, position requirements, salary and benefits, etc.
[0088] In actual operation, these data can be obtained through cooperation with well-known recruitment platforms. For example, Zhaopin is a common recruitment platform that provides a large amount of job seeker resumes and recruitment position data. In addition to Zhaopin, other well-known domestic and foreign recruitment websites such as Lieyunwang, 51job, LinkedIn, etc. can also be considered. These platforms can also provide rich job seeker and position data, and the data format is usually a structured data set or crawled through an API interface. These data include but are not limited to:
[0089] Job seeker resume data: expected working industry, expected job type, work experience, current working industry, personal skills, etc.;
[0090] Recruitment position data: position name, company name, position category, position requirement description, work location, salary, etc.;
[0091] Interaction data between job seekers and positions: records the matching situation between job seeker resumes and recruitment positions (for example, which job seekers have applied for which positions, and which positions have been viewed by which job seekers).
[0092] Step S200: Preprocess the job seeker resume data to obtain the job seeker resume text, and use a convolutional neural network to design a resume text encoder to encode the job seeker resume text into a resume text vector;
[0093] Step S300: Preprocess the recruitment position data to obtain the recruitment position text, and use a convolutional neural network to design a position text encoder to encode the recruitment position text into a position text vector;
[0094] In this embodiment, as described in step S200, whether it is a job seeker's resume or a recruitment advertisement, each attribute feature contains multiple semantic information. Therefore, further preprocessing of the text features is required. Specifically, first filter out sentences with a length less than 2. Then, clean the text in the original resume to remove useless punctuation marks, spaces, special characters, etc. The cleaned resume data will be used as input for the subsequent encoding stage. Next, the text data in the resume is usually in units of sentences and needs to be tokenized through a tokenization tool (such as the Jieba tokenizer) to convert each resume into a set of several words; step S300 is similar to step S200, but it processes recruitment position data. Recruitment position data usually includes content such as position name, company information, position requirements, work experience requirements, etc. In order to input these data into the recommendation system, similar preprocessing steps are also required.
[0095] After completing the text cleaning and word segmentation processing and in the process of encoding the job text, a convolutional neural network (CNN) is used to design the resume text encoder and the job text encoder, and the adversarial learning of the cascade structure is used to capture the potential features of the job seeker resume and the job advertisement respectively. First, the job seeker resume data is input, and the resume text vector is obtained by using the resume text encoder. At the same time, the job advertisement is input into the job text encoder to obtain the job text vector. Finally, the adversarial learning of the cascade structure is designed, and the generator and discriminator are used to capture the potential features of the job seeker resume and the job requirements to accelerate the training process. The following will describe the resume text encoder in detail, then briefly describe the job text encoder, and finally introduce the process of adversarial learning:
[0096] As shown in step S200, in order to capture the semantic information and hierarchical relationship in the resume, a resume text encoder based on a convolutional neural network (CNN) is designed. The encoder consists of two one-dimensional convolutional layers, a maximum pooling layer in the middle, and an average pooling layer at the end.
[0097] Specifically, the first convolution layer contains a convolution layer, a batch normalization layer (BatchNormalization), a rectified linear unit (ReLU) activation layer, and a maximum pooling layer. Next, the second convolution layer is followed by an adaptive maximum pooling layer. Each resume description is represented by multiple words, for example, the resume description r j It can be expressed as a sentence consisting of a group of words. The formula is as follows:
[0098]
[0099] S=|w1,w2,....,w s | (2)
[0100] Among them, r j Represents a job applicant's resume description, which consists of multiple sentences. S is the sentence in the resume description, n is the number of sentences in the resume, j is the index of the resume, T is the transposition, and S∈R d×|S| Representative nr j Resume j Sentences used to describe historical work experience and desired positions, w i ∈R d represents the word embedding of the i-th word in the sentence, and s represents the number of words in the sentence; it can be understood that formula (1) describes the representation of resumes, where each resume consists of multiple sentences, and formula (2) describes how each sentence is represented as a collection of words, and each word is represented by a word embedding vector.
[0101] The input of the resume text encoder is the resume matrix, which is then processed through two one-dimensional convolutional layers. In each convolutional layer, the weight parameter r j The dot product with the n-gram sequence of each word is used to generate the output sequence. To reduce the training cost and accelerate the training of the deep network, batch normalization is performed on the output of the convolutional layer, followed by a non-linear transformation through the ReLU activation function and one-dimensional max pooling.
[0102] Specifically, the size of the second max pooling layer is set to the length of its input, aiming to map the current job description and the next job expectation in the resume to a vector. When analyzing resume data, it is found that candidates usually show their skills in detail in their work experience, and each resume description usually contains multiple skill information. To comprehensively reflect the candidate's ability, this method designs an average pooling layer to summarize this information. The formula of this layer is defined as:
[0103]
[0104] where is the final embedded representation of the resume text, represents the embedded vector of each word or sub-word in the resume description, and the embedded vector of each word reflects the meaning of the word in the semantic space; n represents the number of words in the resume description (i.e., the number of words or sub-words in the resume); represents the number of words in the resume description r j in. It can be understood from this formula that the vectors of each word in the resume are averaged to obtain a single vector X rj , which can effectively represent the semantic information of the entire resume.
[0105] As in step S300, to capture the semantic information and hierarchical relationship in the job requirements, a job text encoder based on a convolutional neural network (CNN) is designed. This job text encoder consists of two one-dimensional convolutional layers, and the convolutional layer is concatenated with the max pooling layer, followed by another max pooling layer. The design of the convolutional layer is similar to that of the resume text encoder.
[0106] When processing job requirements, the job description p i will be transformed into a sequence of word vectors where each word represents the embedded vector of a word in the job description. The job description usually contains multiple requirements, which respectively represent the expectations of the job for professional knowledge in different aspects. Therefore, this method designs a max pooling layer to extract the key information in the job description.
[0107] Specifically, the purpose of the max pooling layer is to obtain an embedding vector containing the most significant features by pooling each dimension of each word vector. The formula is as follows:
[0108]
[0109] Among them, represents the final embedding vector of the job text; l represents the length of the latent vector, that is, the size of each dimension after pooling; represents the maximum value of all job description word vectors on each dimension k, where k belongs to [0, 1, 2,..., l];
[0110] Through this method, multiple requirements in the job description can be aggregated into a vector representation containing the key information of the job as the embedding of the job text.
[0111] Step S400: Construct a heterogeneous bipartite recruitment graph, where the nodes represent job seeker resumes and recruitment positions respectively, and the edges represent the interaction relationships between job seeker resumes and recruitment positions;
[0112] The purpose of step S400 is to construct a heterogeneous bipartite recruitment graph. In this graph, the nodes represent job seeker resumes and recruitment positions respectively, and the edges represent the interaction relationships between job seeker resumes and recruitment positions. Specifically, the heterogeneous bipartite recruitment graph consists of two types of nodes: one is the node set representing job seeker resumes, and the other is the node set representing recruitment positions. In this way, the matching situation between each job seeker and position can be established. The specific construction process is as follows:
[0113] Step S401: First, regard job seeker resumes and recruitment positions as two different types of node sets in the graph, which are represented as the job seeker node set R and the position node set P respectively. Each node represents the resume of a job seeker or the information of a recruitment position. For example, P = {P1, P2,..., P m} represents all recruitment positions, and R = {r1, r2,..., r n} represents all job seeker resumes.
[0114] Step S402: Next, it is necessary to model the interaction between each job seeker resume and recruitment position. If there is a historical interaction between a certain job seeker and a certain recruitment position (for example, the job seeker has applied for the position or the position has viewed the job seeker's resume), an edge will be established between the job seeker node and the position node. In this way, the historical interaction can be recorded in the graph, thereby capturing the relationship between the job seeker and the position.
[0115] Step S403: On this basis, generate the adjacency matrix of the heterogeneous bipartite recruitment graph. The role of the adjacency matrix is to record the relationships between different nodes in the graph. The elements in this matrix are 1 or 0, indicating whether there is an edge connection between nodes. In this study, the relationships within the applicant nodes and within the position nodes are not considered. Therefore, the adjacency matrix only contains edges across node types, and nodes can only be connected through the edges connecting applicant resumes and recruitment positions. There are no allowed edges between nodes of the same type (i.e., between applicant nodes or between position nodes).
[0116] Step S404: According to the interaction information between applicants and positions obtained in the above steps, fill the corresponding positions in the adjacency matrix, and finally construct a complete heterogeneous bipartite recruitment graph.
[0117] In this graph, the feature vectors of the nodes represent the detailed information of each applicant resume and recruitment position, such as the applicant's work experience, skills, expected position, etc., and the requirements of the recruitment position, position name, company name, etc. The edges represent the historical interactions (such as application, viewing, etc.) between the applicant and the position. Through the combination of these nodes and edges, the heterogeneous bipartite graph can effectively capture the relationships between applicants and positions and help with accurate recommendations.
[0118] Step S500: According to the obtained resume text vectors, position text vectors, and their interaction relationships, calculate the initial graph embedding representations of the applicant resume nodes and recruitment position nodes to obtain the initial resume graph embedding and the initial position graph embedding;
[0119] By combining the applicant resume text vectors, position text vectors, and their interaction relationships, obtain the initial representation of each node (i.e., each applicant and position) in the heterogeneous bipartite recruitment graph. This representation will be used in the subsequent graph embedding process so that the model can capture the potential connections between applicants and positions.
[0120] Step S501: According to the obtained resume text vectors and position text vectors, use them as the initial feature representations of the applicant nodes and position nodes.
[0121] In this stage, directly use the resume text and position text vectors as the initial feature representations of the applicant nodes and position nodes in the graph. These vectors are the resume and position text vectors generated through the previous steps (such as steps S200 and S300), which respectively capture the basic information of applicants and positions. These features will be used as the inputs of the nodes and passed to the graph embedding model to further learn the mutual relationships between nodes.
[0122] Step S502: Use the constructed heterogeneous bipartite recruitment graph to combine the obtained resume text vectors and job text vectors with the graph structure information of job seeker nodes and job nodes.
[0123] In this step, the resume text vectors and job text vectors are combined with the graph structure information of the heterogeneous bipartite recruitment graph. The heterogeneous bipartite graph consists of two types of nodes (job seekers and jobs) and the interaction relationships (edges) between them. In this process, the node features (i.e., the resume text vectors and job text vectors) are combined with the edges (interaction relationships) between the nodes to form an overall input for subsequent graph embedding calculations.
[0124] Step S503: Use a graph embedding method to process the node features and the adjacency matrix, and recursively update the embedding representations of job seeker nodes and job nodes.
[0125] At this stage, a graph embedding method (such as a graph convolutional network or other graph neural network methods) is used to process the node features and the adjacency matrix. The adjacency matrix records the relationship information between nodes, that is, whether there is an interaction between a job seeker node and a job node. Through the graph embedding method, the node features will be combined with the adjacency matrix, and the embedding representation of each node will be recursively updated. This means that the representation of each node will be affected not only by its own features but also by the features of its neighbor nodes, thus better capturing the relationships between nodes.
[0126] Step S504: Through graph convolution operations, obtain the initial resume graph embedding and the initial job graph embedding.
[0127] In this step, the graph convolution operation (Graph Convolution Operation) is used to finally obtain the initial resume graph embedding and job graph embedding. Graph convolution updates the representation of a node by aggregating the information of its neighbor nodes. The embedding representations of job seeker nodes and job nodes will be updated through graph convolution operations, and finally the initial resume graph embedding and job graph embedding will be obtained.
[0128] It can be understood that during the graph embedding process, the above-mentioned constructed heterogeneous bipartite graph embedding model is used. Suppose there is a parameterized graph embedding function f emb , and its formula is:
[0129] H p , H r = f emb (X p , B p , X r , B r ; θ) (5)
[0130] Where:
[0131] is the feature matrix of the job nodes, which contains the feature information of each job, with a dimension of m×k, where m is the number of jobs and k is the dimension of the feature vector of each job;
[0132] is the feature matrix of the resume nodes, which contains the feature information of each resume, with a dimension of n×q, where n is the number of resumes and q is the dimension of the feature vector of each resume;
[0133] is the adjacency matrix between jobs and resumes, representing the interaction relationship between job nodes and resume nodes;
[0134] is the adjacency matrix between resumes and jobs, representing the interaction relationship between resume nodes and job nodes;
[0135] and are the initial graph embedding representations of jobs and resumes respectively, obtained by calculating through the graph embedding model;
[0136] θ represents the parameters of the graph embedding model.
[0137] It can be understood from this formula that based on the feature matrices of jobs and resumes and their interaction matrix, through the graph embedding function f emb , the embedding representations H p and H r of jobs and resumes are obtained. These embedding representations can capture the potential relationships between job seekers and jobs and are used for subsequent matching and recommendation tasks.
[0138] Step S600: Use the generator to perform information transfer on the obtained initial resume graph embedding and initial job graph embedding, and respectively generate the embedding representations from the job seeker's resume to the recruitment position and from the recruitment position to the job seeker's resume through the generator;
[0139] To better facilitate the training process and capture potential features, in the online recruitment reciprocal bilateral recommendation model (HBRBR), the purpose of the generator is to achieve two-way information transfer, that is, from the job seeker's resume to the job information and from the job information to the job seeker's resume. The purpose of the discriminator is to distinguish the information transferred from the job to the resume (from the resume to the job) and the information from the resume (job) itself.
[0140] The inputs of the generator include the job feature matrix H p , the resume feature matrix H r , the job adjacency matrix B p and the resume adjacency matrix B r. Since there are no self - connecting edges within the node domains in the heterogeneous bipartite recruitment graph, this process can be defined by the following formula:
[0141]
[0142] To ensure the stability of the generator's calculation, the adjacency matrix is normalized. The specific normalization process is as follows:
[0143]
[0144] Among them, represents the normalized job adjacency matrix at position p, I represents the identity matrix, D p and D r are the degree matrices of the job adjacency matrix B p and the resume adjacency matrix B r respectively, represents the normalized resume adjacency matrix of resume node r.
[0145] At the first layer of the generator, when calculating the embedding from the job to the resume, the generator will generate a new embedding according to the job feature matrix X p and the adjacency matrix B p . Similarly, when calculating the embedding from the resume to the job, the generator will generate a new embedding according to the resume feature matrix and the adjacency matrix. As the network deepens, in the subsequent layers, the generator will continue to update the features based on the embeddings of the previous layer. The specific generator process can be defined as:
[0146]
[0147] Among them, c is the index of the layer, representing the calculation of the generated embedding of the current layer. Specifically: is the hidden feature R of the resume set aggregated from the features of the position set P. And, is defined in a similar way. When the layer index c = 0, and are the input features respectively. W r c and are parameters. Note that the generator does not consider self - loop calculations, which means it only aggregates the information of neighbor nodes. In addition, the propagation of the generator is only single - hop neighbor - aware. The output of the generator is the embedding, representing the transfer of information from the position to the resume and from the resume to the position
[0148] The meanings of the parameters in the formula: H p and H rFeature matrices representing job nodes and resume nodes respectively; B p and B r represent the adjacency matrices between job and resume nodes respectively; and are the normalized adjacency matrices of jobs and resumes respectively; D p and D r are the degree matrices of the job and resume adjacency matrices respectively, used for normalization; σ: activation function; and are the weight matrices of the c-th layer of the generator; and are the job and resume embedding representations of the c-th layer of the generator.
[0149] The above process realizes the two-way information transfer from jobs to resumes and from resumes to jobs through a multi-layer generator architecture, and captures the potential associations between job seekers and jobs by updating the embedding representations layer by layer.
[0150] Step S700: Use the discriminator to perform adversarial learning on the generated embeddings, distinguish the information of the embedding representation from the job seeker's resume to the recruitment position and the embedding representation from the recruitment position to the job seeker's resume, and optimize the embedding representation of the generator;
[0151] To achieve the goal of adversarial learning, the task of the discriminator is to distinguish two different sources of information: one is the embedding passed from the job node to the resume node, and the other is the embedding from the resume node itself.
[0152] Specifically, the discriminator has two sets of inputs:
[0153] The first set of inputs is the embedding from the job to the resume, which represents the information passed from the job node to the resume node by the generator, and also includes the feature embedding of the resume set itself.
[0154] The second set of inputs is the embedding from the resume to the job, which represents the information passed from the resume node to the job node by the generator, and also includes the feature embedding of the job set itself.
[0155] In adversarial learning, a two-player adversarial maximization-minimization game is formed between the generator and the discriminator. The goal of the discriminator is to maximize the ability to identify these two different feature representations, while the generator hopes to prevent the discriminator from making an accurate distinction by aligning the source domain (such as the embedding from the job to the resume) and the target domain (such as the embedding of the resume).
[0156] In the first layer of adversarial training, the discriminator outputs the final feature representation of the resume by comparing the embeddings passed from the job to the resume and the embeddings of the resume itself. In the subsequent layers, the discriminator aims to maximize the difference between the embeddings from the job to the resume and the embeddings from the resume to the job that it can identify. The goal of the discriminator is to distinguish between the feature representations generated by the generator and the feature representations of the actual job or resume.
[0157] Specifically, the output of the discriminator will only contain the single-hop topological information from the adjacency matrix and the semantic information from the resume or job features. After training, through adversarial training, a Nash equilibrium is reached between the generator and the discriminator, and finally, the embedding representations H r and H P can be generated, which can fuse the feature information from the job and the resume, thus helping the recommendation system to perform better matching.
[0158] Loss functions of the discriminator and the generator:
[0159] Discriminator loss function: The goal of the discriminator is to maximize the ability to distinguish information from two different sources (job and resume), so the loss function of the discriminator is defined as follows:
[0160]
[0161] where: Prob β,α (source = 0|h r(i) ) represents the prediction of the discriminator for the resume feature h r(i) , to judge whether it comes from the resume; Prob β,α (source = 1|h p→r(i) ) represents the prediction of the discriminator for the embedding h p→r(i) from the job to the resume, to judge whether it comes from the job; β and α are the parameters of the generator and the discriminator.
[0162] Generator loss function: The goal of the generator is to prevent the discriminator from accurately identifying the difference between them by aligning the source domain (embedding from the job to the resume) and the target domain (embedding of the resume), so the loss function of the generator is defined as follows:
[0163]
[0164] where, h p→r(i) represents the embedding from the job to the resume generated by the generator; Prob β,α (source = 0|h p→r(i) ) is the loss of the generator, and the goal is to maximize the similarity between the generated embedding and the real data.
[0165] Through adversarial training, the generator and the discriminator play against each other. Eventually, the generator can generate more accurate job-to-resume or resume-to-job embedding representations, while the discriminator can effectively distinguish their sources.
[0166] As mentioned above, current embedding methods only capture the single-hop topological structure in the adjacency matrix and the feature information from resumes and jobs. However, single-hop aggregation methods cannot fully represent complex graph structures, especially when more hierarchical node relationships need to be captured. Therefore, it is necessary to introduce multi-hop mechanisms or deep network structures.
[0167] The idea of a cascaded architecture is adopted in the model. Specifically, the generator and the discriminator are regarded as two parts of a deep network in the model. In each layer, the output embedding of the generator serves as the input of the discriminator for further processing.
[0168] To improve training efficiency, during the training process of each layer, only one-hop embeddings are trained. This means that in each training epoch, only the single-hop neighbor information of the current layer is aggregated, rather than training the entire embedding representation by propagating a fixed number of hops through multiple layers. To accelerate the training process and improve memory efficiency, for each epoch, these hidden embeddings are sampled to form mini-batches for processing. At the same time, for each layer, after the embedding calculation is completed, the model instance and unused memory are released to reduce memory occupancy.
[0169] This method is different from the traditional end-to-end training method. It does not need to train the embedding of the entire graph through multi-layer propagation, but can significantly improve memory utilization efficiency, shorten the training time, and accelerate the convergence speed of the model by only training single-hop embeddings.
[0170] Step S800: Update the initial resume graph embedding and the initial job graph embedding based on the optimized embedding representation to obtain the updated resume graph embedding and the updated job graph embedding;
[0171] It can be understood that the updated embedding representation can more accurately reflect the relationship between job seekers and jobs. Through this update, the graph embedding can capture more potential features.
[0172] Step S900: Concatenate the obtained resume text vector, job text vector, updated resume graph embedding, and updated job graph embedding to generate a comprehensive matching vector between the job seeker and the job;
[0173] Step S1000: Use a multi-layer perceptron to process the comprehensive matching vector, calculate the matching probability between the resume and the job; output the matching probability value between the resume and the job as the prediction result of the matching between the job seeker and the recruitment position.
[0174] In this embodiment, as described in the above steps S900 - S1000, in order to achieve a comprehensive match between resume data and job data, a training and prediction module is designed. First, a concatenation layer is designed to connect the resume text embedding, job text embedding, resume graph embedding, and job graph embedding. Then, a classification module is designed to calculate the matching probability between the resume and the job using a multi - layer perceptron (MLP).
[0175] In the concatenation layer, the resume text embedding and job text embedding obtained from the text encoder, as well as the resume graph embedding and job graph embedding learned from the bipartite graph, are concatenated. Specifically, the resume text embedding learned by the resume text encoder, the job text embedding learned by the job text encoder, the resume graph embedding learned from the heterogeneous bipartite graph, and the job graph embedding are concatenated. The concatenation operation can be represented by the following formula:
[0176] X RP = CONCAT(X r ·X p , H r + H p ) (11)
[0177] In formula (11): X r is the resume text embedding, representing the features of the resume content; X P is the job text embedding, representing the features of the job description; H r is the resume graph embedding, representing the graph embedding features of the resume nodes; H P is the job graph embedding, representing the graph embedding features of the job nodes; CONCAT represents the concatenation operation.
[0178] Next, a classification module is designed, which consists of a multi - layer perceptron (MLP) and a sigmoid function. Specifically, the MLP performs a linear layer transformation on the concatenated vector X RP and outputs a matching probability. Then, the sigmoid function is used to calculate the matching probability between the resume and the job, and the prediction result y' is obtained. The formula is as follows:
[0179]
[0180]
[0181]
[0182] where σ is the sigmoid activation function, defined as:y' is the probability value of the match between the resume and the job, ranging from [0, 1]. MLP(X RP ) is the output after processing the concatenated vector X RP through the multi - layer perceptron.
[0183] To optimize the model, a loss function is defined, namely the Bayesian Personalized Ranking Loss. This loss function aims to minimize the gap between the prediction results and the actual matching situation, thereby improving the accuracy of the recommendation system.
[0184] Through the above design, the model can effectively combine the features of resumes and job positions, calculate the matching degree between them, and make job recommendations based on this.
[0185] The overall framework of the Online Recruitment Reciprocal Bilateral Recommendation Model (HBRBR model) is as Figure 2 shown. The HBRBR model mainly consists of three modules. The first part is to construct a heterogeneous bipartite graph of resumes and job positions. The second part is to generate embeddings based on the capabilities of job seekers and job requirements (icons G and D represent the generator and discriminator respectively). Specifically, a CNN-based resume text encoder and a job position text encoder are designed to capture the bilateral preferences of job seekers and recruiters, and obtain the initial resume text embeddings and job position text embeddings respectively. In addition, adversarial learning with a cascade architecture is introduced to capture implicit features and facilitate the training process. The third part is training and prediction. The HBRBR model can comprehensively grasp the needs of both job seekers and recruiters, help job seekers find suitable jobs, recommend potential talents, and make bilateral recommendations at the same time.
[0186] To further illustrate this method, the following experiments will be conducted for verification:
[0187] The relevant statistical information of the dataset used in this experiment is summarized in Table 1. The experimental dataset contains a total of 4,501 job seeker resume data, 19,379 recruitment information data, and 28,595 interaction data between job seeker resumes and recruitment information. The node distribution of the dataset is visualized, as Figure 3 shown. The job position and resume datasets have similar degree distributions, both with a long-tailed exponential distribution. The negative examples between job seekers and recruiters are generated at a ratio of 1:1 based on the interaction data successfully matched during the verification and training processes. Specifically, for each positive example, 20 job positions are randomly selected as candidates, and 20 job position candidates are used as negative examples. Therefore, the number of positive samples in the experiment is 28,595, and the number of negative samples is 8,526. Finally, the valid interactions are randomly divided, 70% are assigned to the training set, 15% are assigned to the verification set, and the remaining 15% are assigned to the test set. The dataset is divided according to the job seeker resume id or job position id, and the interaction data of the same job seeker or job position is divided into one dataset as much as possible. Finally, there are 20,069 training instances, 3,993 verification instances, and 4,533 test instances.
[0188] Table 1: Experimental Statistical Information Dataset
[0189]
[0190] To compare the performance, ten state-of-the-art methods were selected as baselines. The comparison models can be divided into three categories: (1) Collaborative filtering-based methods: LightGCN, LFRR, and BPR; (2) Content-based methods: PJFNN, BPJFNN, APJFN, and BERT; (3) Hybrid methods: DPGNN. In addition, since content-based methods are essentially reciprocal methods, content-based reciprocal baseline models were not included.
[0191] LightGCN: LightGCN is a simplified graph convolutional network model that only contains neighborhood aggregation in GCN. Specifically, LightGCN learns user and item embeddings through a user-item interaction graph and uses the weighted sum of the embeddings learned in all layers as the final embedding.
[0192] LFRR: Latent Factor Reciprocal Recommender is a latent factor-based model that uses an aggregation function to combine user preference scores to implement a reciprocal collaborative filtering method.
[0193] BPR: BPR is a popular general recommendation method based on implicit feedback. The learning method of BPR is based on stochastic gradient descent and bootstrap sampling.
[0194] PJFNN: Person-Job Fit Neural Network is an end-to-end CNN-based model for evaluating the matching degree between talent qualifications and job requirements. Specifically, PJFNN is a bipartite neural network that can efficiently learn the joint representation of person-job fit from historical job application data.
[0195] APJFNN: Ability-Aware Person-Job Fit Neural Network is an ability-aware person-job fit neural network based on end-to-end recurrent neural network. The core idea of APJFNN is to utilize the rich information in historical job application data. It designs four hierarchical ability-aware attention strategies to measure the different importance of job requirements for semantic representation and the different contributions of each work experience to specific ability requirements, so as to obtain the word-level semantic representation of job requirements and job seeker experience.
[0196] BPJFNN: Basic Person-Job Fit Neural Network is a simplified version of the above APJFNN model. BPJFNN treats all the abilities in the recruitment information and the experience in the candidate's resume as a whole. BPJFNN uses a bidirectional long short-term memory (BiLSTM) network to learn the resume and recruitment information representation.
[0197] BERT: Bidirectional Encoder Representations from Transformers is a pre-trained language model that can be used to generate representations for job seekers and recruiters. BERT can be used to pre-train deep bidirectional representations from unlabeled text by jointly conditioning the context of all layers.
[0198] DPGNN: The Dual-Perspective Graph Neural Network is a dual-perspective graph representation learning method used to model the directed interactions between job seekers and positions. To model preferences from the dual perspectives of job seekers and recruiters, DPGNN merges two different nodes for each candidate (or position) and represents successful and failed job matches through a unified dual-perspective interaction graph.
[0199] To comprehensively evaluate the performance of the model, five evaluation metrics were selected: AUC (Area Under the ROC Curve), Accuracy, Recall, Precision, and F1-Score. The definitions and descriptions of these metrics are as follows:
[0200] AUC represents the area under the ROC curve, which is a commonly used statistical value that can represent the prediction ability of the model. The larger the area under the ROC curve, the better the model. When AUC equals 1, it means the model is perfect.
[0201]
[0202] where M represents the number of positive items, N represents the total number of interactions between users and items, and rank i represents the descending rank of the i-th positive item.
[0203] Accuracy represents the ratio of the number of samples correctly predicted by the model to the total number of samples. The higher the accuracy, the better.
[0204]
[0205] where TP, TN, FP, and FN represent True Positive, True Negative, False Positive, and False Negative respectively.
[0206] Recall represents the proportion of samples that are correctly predicted as positive among all samples that are actually positive.
[0207]
[0208] Precision represents the proportion of samples that are actually positive among all samples predicted as positive.
[0209]
[0210] F1-Score is the harmonic mean of precision and recall, which can balance the performance of the two metrics on an imbalanced dataset.
[0211]
[0212] The experiments were implemented using PyTorch Geometric 2.2.0 on a Linux machine with a GeForce RTX 3090 GPU equipped with 24GB of memory. First, the dimension y of the feature embedding was set to 64. The Adam optimizer was used to optimize the parameters during model training. The experiments adopted an early stopping mechanism. The threshold for validation-based early stopping was 50 epochs. This means that if the validation performance of the model did not improve significantly within 50 training epochs, the training would be stopped early. The learning rate was set to 0.001, the weight decay was set to 0.1, the mini-batch size was 256, the dropout was 0.1, and the depth was 2. The baseline model was implemented using the popular open-source recommendation library RecBole. For the baseline model, either the best settings were followed or the corresponding parameters were optimized using the validation set. Additionally, positive samples were defined as cases where the application was successful and the interview was passed. However, some matching failures might not be due to the mismatch between the job description and job requirements, but other reasons, such as the salary being lower than expected. Therefore, these records of matching failures should also be classified as positive samples. To ensure effectiveness, it was decided to randomly replace the resume description or job requirements in successful matches instead of directly extracting negative samples from failed matches. To evaluate the robustness of the model, the dataset was split into 70% for training, 15% for validation, and 15% for testing.
[0213] To demonstrate the effectiveness of the model of the proposed method in the mutual two-sided recommendation in online recruitment, the HBRBR model was compared with all baseline methods, and the overall performance is shown in Table 2 and Figure 4 as follows. The best-performing method is shown in bold. Obviously, the HBRBR model proposed by the present method is significantly better than the comparison models in all evaluation metrics, demonstrating that the HBRBR model can effectively capture the resumes of job seekers and job requirements, and can not only recommend suitable jobs for job seekers, but also recommend suitable job seekers for recruiters. From the results, the following observations can be summarized:
[0214] Among the four baselines based on collaborative filtering, LightGCN performs the best, but the improvement is not significant compared with LFRR and BPR.
[0215] As for the three content-based baselines, BPJFNN, APJFNN, and BERT, they highly rely on text content and perform poorly under most metrics. In addition, the CNN-based PJFNN performs well. The possible reason is that the first three models are better at processing structured job seeker resumes and recruitment information data, while in the dataset of this solution, both recruiters and job seekers have different text organization habits. This also further validates the advantage of choosing the CNN-based structure for text learning in the HBRBR model of this method.
[0216] The hybrid-based baselines, such as DPGNN and the HBRBR model of this method, utilize information from both interactions and text. They perform better than most baselines, indicating that it is very important to utilize both text descriptions and interactions simultaneously.
[0217] Table 2: Performance of All Methods
[0218]
[0219]
[0220] The main technical contribution of the method of this invention is to use a CNN-based text encoder for text learning and introduce adversarial learning for network topology learning. Now, analyze the contribution of each component to the final performance, considering the following two HBRBR variants:
[0221] 1) HBRBR w / o txt-encoder: Remove the CNN-based text encoder, that is, remove the resume text encoder and the job text encoder.
[0222] 2) HBRBR w / o adv_learning: Eliminate the adversarial learning process of the generator and discriminator, and directly use the resume embeddings and job embeddings generated by the text encoder for training and prediction.
[0223] In Figure 5 it can be seen that the performance order can be summarized as HBRBR w / o adv_learning < HBRBR w / o txt-encoder < HBRBR. These results indicate that both components (i.e., the CNN-based text encoder and adversarial learning) contribute to improving the performance of HBRBR. In particular, the resume text encoder and the job text encoder bring more improvements to the method of this invention.
[0224] Next, the robustness of the model of this invention will be studied and a detailed analysis of the key hyperparameters will be conducted.
[0225] Change the depth of HBRBR. Introduce l to represent the depth of HBRBR. Here, the depth is changed from 1 to 5. AsFigure 6 As shown, the model of the present invention achieves optimal performance when l = 2, indicating that HBRBR can deepen the use of interactive historical data. In addition, as Figure 6 shown by the shadow in, when the depth is greater than 2, the AUC is greater than 0.7, the model performance is relatively stable, and the prospect is promising. When the depth l = 4, the model may encounter the problem of over-smoothing and obtain sub-optimal results.
[0226] Change the learning rate of HBRBR. The HBRBR model is tuned in the range of {0.00001, 0.0001, 0.001, 0.01, 0.1}. It is found that as the learning rate increases, the performance of the HBRBR model first increases and then decreases. As Figure 7 shown by the shaded part in, when the learning rate is in the range of 0.0001 to 0.01, HBRBR performs well. In addition, when the learning rate is around 0.01, HBRBR achieves the best result.
[0227] Change the weight decay of HBRBR. The HBRBR model is adjusted in the range of {0.0005, 0.001, 0.005, 0.01, 0.1}. As Figure 8 shown, it can be seen from the figure that when the weight decay increases from 0.0005 to 0.001, the performance of HBRBR improves significantly. Subsequently, it can be observed that when the weight decay is in the range of 0.001 to 0.1 (as Figure 8 shown by the shaded part in), the performance of HBRBR is already good enough.
[0228] Change the batch size of HBRBR. The batch size is adjusted in the range of {32, 64, 128, 256, 512}. As Figure 9 shown, it can be seen that the performance of the HBRBR model has a process of first rising and then falling as the batch size increases. And, the AUC values in the shaded area are all greater than 0.6, indicating that when the batch size is in this range, the model performs well. It can be seen that HBRBR performs best when batchsize = 64, which indicates that for this data set, it is more appropriate to set the batchsize to 64.
[0229] The present invention proposes a reciprocal bipartite recommendation model based on a heterogeneous bipartite graph neural network, aiming to recommend suitable positions for job seekers and suitable resumes for recruiters. In the above steps, a heterogeneous bipartite recruitment graph is constructed to simulate the interaction between recruitment information and job seekers. To learn effective node representations, a resume text encoder and a position text encoder are respectively designed. To optimize the model, adversarial learning with a cascaded structure is introduced to capture the potential preferences of positions and resumes. A large number of experiments show that, compared with eight SOTA models, the model proposed by the present invention performs better in all evaluation metrics from the perspectives of candidates and positions. The results show that the HBRBR model of the present invention can provide effective decision-making support for job seekers and recruiters.
[0230] According to another aspect of the embodiments of the present application, an online recruitment reciprocal bilateral recommendation system is further provided. The recommendation system includes the following modules:
[0231] A data acquisition module, configured to acquire job seeker resume data and recruitment position data;
[0232] A resume data preprocessing module, configured to preprocess the job seeker resume data to obtain a job seeker resume text, and design a resume text encoder through a convolutional neural network to encode the job seeker resume text into a resume text vector;
[0233] A position data preprocessing module, configured to preprocess the recruitment position data to obtain a recruitment position text, and design a position text encoder through a convolutional neural network to encode the recruitment position text into a position text vector;
[0234] A heterogeneous bipartite graph construction module, configured to construct a heterogeneous bipartite recruitment graph, where the nodes respectively represent job seeker resumes and recruitment positions, and the edges represent the interaction relationships between job seeker resumes and recruitment positions;
[0235] An initial graph embedding calculation module, configured to calculate the initial graph embedding representations of job seeker resume nodes and recruitment position nodes according to the obtained resume text vectors, position text vectors, and their interaction relationships, to obtain an initial resume graph embedding and an initial position graph embedding;
[0236] A generator module, configured to use a generator to perform information transfer on the obtained initial resume graph embedding and initial position graph embedding, and respectively generate an embedding representation from a job seeker resume to a recruitment position and an embedding representation from a recruitment position to a job seeker resume;
[0237] A discriminator module, configured to use a discriminator to perform adversarial learning on the obtained generated embeddings, distinguish the information of the embedding representation from a job seeker resume to a recruitment position and the embedding representation from a recruitment position to a job seeker resume, and optimize the embedding representation of the generator;
[0238] An embedding update module, configured to update the initial resume graph embedding and the initial job graph embedding based on the optimized embedding representation, so as to obtain the updated resume graph embedding and the updated job graph embedding;
[0239] A matching vector generation module, configured to splice the obtained resume text vector, job text vector, updated resume graph embedding and updated job graph embedding to generate a comprehensive matching vector between the job seeker and the job;
[0240] A matching probability calculation module, configured to process the comprehensive matching vector using a multi-layer perceptron, calculate the matching probability between the resume and the job, and output the matching probability value between the resume and the job as the prediction result of the matching between the job seeker and the recruitment position.
[0241] According to another aspect of the embodiments of the present application, an electronic device is further provided, including a processor and a memory, where the processor is configured to implement the steps of the method when executing a computer program stored in the memory.
[0242] In the above embodiments of the present invention, the descriptions of the respective embodiments have their own emphases. For parts not detailed in a certain embodiment, reference may be made to the relevant descriptions of other embodiments.
[0243] In several embodiments provided by the present application, it should be understood that the disclosed technical content can be implemented in other ways. Among them, the device embodiments described above are merely illustrative. For example, the division of the units can be a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces. The indirect couplings or communication connections of the units or modules can be in an electrical or other form.
[0244] In addition, each functional unit in the various embodiments of the present invention can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above integrated unit can be implemented in the form of hardware or in the form of a software functional unit.
[0245] When the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The foregoing storage medium includes: various media such as USB flash drives, read-only memories (ROMs), random access memories (RAMs), mobile hard disks, magnetic disks, or optical discs that can store program codes.
[0246] The foregoing is only the preferred embodiment of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present invention, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of the present invention.
Claims
1. A mutual bilateral recommendation method for online recruitment, characterized in that: The steps include: Obtain job applicant resume data and job position data; Preprocessing the job seeker resume data to obtain the job seeker resume text, and using a convolutional neural network to design a resume text encoder to encode the job seeker resume text into a resume text vector; Preprocessing the job position data to obtain job position texts, and using a convolutional neural network to design a job text encoder to encode the job position text into a job text vector; Construct a heterogeneous bipartite recruitment graph, where nodes represent job seekers’ resumes and job positions, and edges represent the interaction between job seekers’ resumes and job positions. Based on the obtained resume text vector, position text vector and the interaction between them, the initial graph embedding representation of the job seeker resume node and the recruitment position node is calculated to obtain the initial resume graph embedding and the initial position graph embedding; The generator is used to transfer information between the initial resume graph embedding and the initial job graph embedding, and the embedding representation of the job seeker resume to the job position and the embedding representation of the job position to the job seeker resume are generated through the generator respectively; Use the discriminator to perform adversarial learning on the generated embeddings, distinguish the information from the embedding representation of the job applicant's resume to the job position and the embedding representation from the job position to the job applicant's resume, and optimize the embedding representation of the generator; Based on the optimized embedding representation, the initial resume image embedding and the initial job image embedding are updated to obtain updated resume image embedding and updated job image embedding; The obtained resume text vector, position text vector, updated resume image embedding and updated position image embedding are concatenated to generate a comprehensive matching vector between job seekers and positions; The comprehensive matching vector is processed using a multi-layer perceptron to calculate the matching probability between the resume and the position; the matching probability value between the resume and the position is output as the prediction result of the match between the job seeker and the recruitment position.
2. The online recruitment mutual benefit bilateral recommendation method according to claim 1, characterized in that: The method of encoding the job seeker resume text into a resume text vector comprises: Perform word segmentation on the job seeker's resume text to obtain each word in the resume; Map each word after segmentation into the corresponding word embedding vector; Use a one-dimensional convolutional neural network to perform convolution operations on the obtained word embedding vectors to extract local features in the resume text; The convolution result is pooled through the maximum pooling layer to retain the most important feature information; The pooled result is input into the average pooling layer to aggregate the various skill information in the resume text to obtain the resume text vector.
3. The online recruitment mutual benefit bilateral recommendation method as claimed in claim 1, characterized in that: The method of encoding the recruitment position text into a position text vector includes: Perform word segmentation on the job posting text to obtain each word in the job description; Map each word after segmentation into the corresponding word embedding vector; Use a one-dimensional convolutional neural network to perform convolution operations on the obtained word embedding vectors to extract local features in the job text; The convolution result is pooled through the maximum pooling layer to retain the most important feature information in the job requirements description; The pooled result is input into the average pooling layer to comprehensively process the various requirements of the job text and obtain the job text vector.
4. The online recruitment mutual benefit bilateral recommendation method as claimed in claim 1, characterized in that: Methods for constructing a heterogeneous bipartite recruitment graph include: The job seeker resume and the job position are regarded as two different types of node sets in the graph, which are represented as the job seeker node set and the job position node set respectively; Model the interaction between each job seeker's resume and the job position. If the job seeker has a history of interaction with the position, create an edge between the job seeker node and the position node. Generate an adjacency matrix of a heterogeneous bipartite recruitment graph, where the adjacency matrix represents the interaction between job openings and job seekers' resumes, and between job seekers' resumes and job openings, and only allows cross-node type edge connections between nodes, and does not allow connections between nodes of the same type; According to the obtained interaction information between job seekers and positions, the adjacency matrix is filled, and finally a heterogeneous bipartite recruitment graph is constructed.
5. The online recruitment mutual benefit bilateral recommendation method as claimed in claim 4, characterized in that: The initial resume graph embedding and initial job graph embedding are obtained by the following method: According to the obtained resume text vector and the obtained position text vector, use them as the initial feature representations of the job seeker node and the position node; Using the constructed heterogeneous bipartite recruitment graph, the obtained resume text vector and position text vector are combined with the graph structure information of job seeker nodes and position nodes; Using graph embedding methods, node features and adjacency matrices are processed to recursively update the embedding representations of job seeker nodes and job position nodes. Through graph convolution operations, we get the initial resume graph embedding and the initial job graph embedding.
6. The online recruitment mutual benefit bilateral recommendation method according to claim 5, characterized in that: The method of generating the embedding representation of the job seeker resume to the job position and the embedding representation of the job position to the job seeker resume by the generator includes: Combine the initial resume graph embedding with the adjacency matrix of the job graph, and obtain the resume-to-job embedding representation through graph convolution operations; Take the embedding representation of the job seeker's resume to the job position as input, use the generator's propagation mechanism to pass it to the job position node, and generate the embedding representation of the job position; Combine the initial job graph embedding with the adjacency matrix of the resume graph, and obtain the embedding representation of the job to the resume through graph convolution operation; Take the embedding representation of the job position to the job seeker resume as input, and use the propagation mechanism of the generator to pass it to the job seeker resume node to generate the embedding representation of the job seeker resume; Get the embedding representation of job seeker resume to job position and the embedding representation of job position to job seeker resume.
7. The online recruitment mutual benefit bilateral recommendation method as claimed in claim 1, characterized in that: Methods for optimizing the embedding representation of the generator include: Use a discriminator to classify the embedding representation generated by the generator. The discriminator receives the embedding representation of the job seeker's resume to the job position and the embedding representation of the job position to the job seeker's resume from the generator and judges their authenticity. The discriminator calculates the discriminative loss by maximizing its ability to distinguish between the generated embeddings from the job seeker resume to the job position and the real job position node embedding, and the generated embeddings from the job position to the job seeker resume and the real resume node embedding; According to the feedback from the discriminator, the loss function of the generator is calculated, and the parameters of the generator are updated according to the loss function; Update the parameters of the generator through the back-propagation algorithm; Repeat the above steps to continuously optimize the embedding representation of the generator.
8. The online recruitment mutual benefit bilateral recommendation method as claimed in claim 1, characterized in that: The method of using a multilayer perceptron to process the comprehensive matching vector, calculate the matching probability between the resume and the position, and output the matching probability value between the resume and the position includes: The obtained comprehensive matching vector is input into a multi-layer perceptron. The multi-layer perceptron includes at least one hidden layer, and each hidden layer is transformed by a nonlinear activation function to extract high-level features in the comprehensive matching vector; Through each layer in the multi-layer perceptron, the integrated matching vector is linearly transformed step by step, and nonlinearity is introduced through the activation function; In the last layer, the sigmoid activation function is used to map the output of the multilayer perceptron to obtain the matching probability between the resume and the position; The output sigmoid function uses the calculated matching probability value as the prediction result of the match between the resume and the position, which is used to evaluate the degree of match between the job seeker and the position.
9. An online recruitment mutual bilateral recommendation system, characterized in that: The recommendation system includes the following modules: Data acquisition module, used to obtain job seeker resume data and recruitment position data; A resume data preprocessing module is used to preprocess the resume data of job seekers to obtain the resume text of the job seekers, and to design a resume text encoder through a convolutional neural network to encode the resume text of the job seekers into a resume text vector; A job data preprocessing module is used to preprocess the job data to obtain job texts, and to design a job text encoder through a convolutional neural network to encode the job texts into job text vectors; Heterogeneous bipartite graph construction module, used to construct heterogeneous bipartite recruitment graphs, where nodes represent job seekers' resumes and job positions, and edges represent the interactive relationships between job seekers' resumes and job positions; An initial graph embedding calculation module is used to calculate the initial graph embedding representation of the job seeker resume node and the recruitment position node based on the obtained resume text vector, position text vector and the interaction relationship between them, and obtain the initial resume graph embedding and the initial position graph embedding; A generator module is used to use the generator to transfer information between the initial resume graph embedding and the initial job graph embedding, and to generate an embedding representation from the job seeker's resume to the job position and an embedding representation from the job position to the job seeker's resume, respectively; The discriminator module is used to perform adversarial learning on the generated embedding using the discriminator, distinguish the information from the embedding representation of the job applicant's resume to the job position and from the job position to the job applicant's resume, and optimize the embedding representation of the generator; An embedding update module, used for updating the initial resume image embedding and the initial job image embedding based on the optimized embedding representation to obtain an updated resume image embedding and an updated job image embedding; A matching vector generation module is used to concatenate the obtained resume text vector, position text vector, updated resume image embedding and updated position image embedding to generate a comprehensive matching vector between the job seeker and the position; The matching probability calculation module is used to process the comprehensive matching vector using a multi-layer perceptron, calculate the matching probability between the resume and the position, and output the matching probability value between the resume and the position as the prediction result of the match between the job seeker and the recruitment position.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the online recruitment mutual benefit bilateral recommendation method according to any one of claims 1 to 8 are executed.