Online recruitment bilateral reciprocity recommendation method and device based on hypergraph and electronic equipment
Through a hypergraph-based method, the resume and position data are extracted and fused in structured and unstructured features to capture semantic correlations, solving the problem of low accuracy of resume-position matching in the existing technology, and achieving the effect of two-way recommendation.
Patent Information
- Application Number
- CN202510393297.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-31
- Publication Date
- 2025-07-01
AI Technical Summary
In the prior art, the resume-job matching method ignores the complex semantic relationship between resume and job data, resulting in poor accuracy of matching results and difficult to meet the two-way needs of job seekers and recruiters.
Using a hypergraph-based method, structured and unstructured attribute features are extracted on resume and position data, feature embedding is generated, and internal and cross-semantic associations are captured through hypergraph learning and attention mechanisms, feature representations of different types of data are fused, and matching probability prediction and recommendation are carried out.
Improve the accuracy of resume-job matching, and achieve bilateral mutually beneficial recommendations between job seekers and recruiters to meet the needs of both parties.
Smart Images

Figure CN120235599A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to a recommendation method, and more particularly, to an online recruitment bilateral reciprocal recommendation method, device and electronic device based on a hypergraph. Background Art
[0002] Now, with the rapid development of intelligent recruitment, numerous online recruitment platforms have emerged. The rise of online recruitment services has promoted the interaction between countless job seekers and recruiters, generating a large amount of resume and recruitment advertisement data. Therefore, it is often difficult for job seekers to quickly find suitable positions, while recruiters face the challenge of effectively identifying qualified candidates.
[0003] Generally, recruiters will post recruitment advertisements outlining job requirements and descriptions, including job titles, educational requirements, work locations and salaries. At the same time, job seekers submit personal resumes detailing their preferences and work experiences, such as recent job categories, educational backgrounds, expected work locations and expected salaries. Therefore, the key to an online recruitment platform lies in the degree of matching between resume preferences and job requirements, i.e., resume - job matching (or person - job matching). Different from traditional one - way user - item recommendation systems, resume - job matching is a two - way reciprocal process that needs to consider the bilateral needs of the recruitment market.
[0004] Specifically, online recruitment exhibits three main characteristics: (1) Bilateralism. Job seekers and recruiters provide information from different perspectives. Job seekers usually share personal background information, while recruiters focus on job - related requirements. However, the expectations for the same attributes of resumes and positions may vary. For example, job seekers usually seek higher salaries, while recruiters aim to minimize costs and find candidates who meet job requirements at the lowest possible salary. (2) Reciprocity. Effective resume - job matching needs to meet the expectations of both sides of the recruitment, which means satisfying both the preferences of job seekers and the requirements of the job. Traditional unilateral recommendation systems can only meet the needs of one side and thus often fail to achieve successful matching. For example, a job seeker may apply for a position but be rejected by the recruiter, or the recruiter may send an interview invitation but the job seeker rejects it. (3) Data structure diversity. Job seekers' resumes and recruitment advertisements usually consist of multiple data structures, including structured attributes (i.e., educational background, work location and salary) and unstructured text data (i.e., work experience and job description).
[0005] According to the above characteristics of online recruitment, resume-job matching is a two-sided reciprocal recommendation problem in online recruitment. Different from the traditional one-way user-item model, it requires a deep understanding of the preferences of job seekers and recruiters. Some existing research can be roughly divided into two categories: text-matching-based methods and behavior-based methods. Most text-matching-based methods generate personalized recommendations by learning the representations of job seeker resume data and job requirements. Behavior-based methods mainly use the historical interaction data between job seekers and job positions in the online recruitment platform to generate recommendations. In the prior art, resume-job matching is regarded as a supervised text-matching problem, ignoring the complex semantic relationships between resume and job data, resulting in poor accuracy of the matching results and making it difficult for the recommendation results to meet the needs of both parties. Summary of the Invention
[0006] The technical problem to be solved by the present invention is that the above methods commonly used in the prior art regard resume-job matching as a supervised text-matching problem, ignoring the complex semantic relationships between resume data and job data, resulting in poor accuracy of the matching results and making it difficult for the recommendation results to meet the needs of both parties. To solve the above problems, the present invention provides a hypergraph-based two-sided reciprocal recommendation method, device, and electronic device for online recruitment.
[0007] The content of the present invention includes:
[0008] In a first aspect, an embodiment of the present invention provides a hypergraph-based two-sided reciprocal recommendation method for online recruitment, including:
[0009] Performing feature extraction on multiple resume data and multiple job data respectively to obtain the attribute features of each resume data and the attribute features of each job data. The attribute features of the resume data include first structured attribute features and first unstructured attribute features, and the attribute features of the job data include second structured attribute features and second unstructured attribute features;
[0010] Generating resume feature embeddings based on the first structured attribute features, generating job feature embeddings based on the second structured attribute features, generating initial resume text embeddings based on the first unstructured attribute features, and generating initial job text embeddings based on the second unstructured attribute features;
[0011] Fusing the resume feature embeddings and the job feature embeddings to obtain resume structured embeddings and job structured embeddings, and fusing the initial resume text embeddings and the initial job text embeddings to obtain resume text embeddings and job text embeddings;
[0012] Predict the matching probabilities of the multiple resume data and the multiple job data based on the total feature representation, and determine the recommendation result based on the prediction result. The total feature representation is obtained by concatenating the resume structured embedding, the job structured embedding, the resume text embedding, and the job text embedding.
[0013] Optionally, performing feature extraction on the multiple resume data and the multiple job data respectively to obtain the attribute features of each resume data and the attribute features of each job data, including:
[0014] Performing feature extraction on the structured data in the multiple resume data to obtain the first structured attribute features of each resume data;
[0015] Performing feature extraction on the unstructured data in the multiple resume data to obtain the first unstructured attribute features of each resume data;
[0016] Performing feature extraction on the structured data in the multiple job data to obtain the second structured attribute features of each job data;
[0017] Performing feature extraction on the unstructured data in the multiple job data to obtain the second unstructured attribute features of each job data.
[0018] Optionally, generating a resume feature embedding based on the first structured attribute features and generating a job feature embedding based on the second structured attribute features, including:
[0019] Constructing a resume feature hypergraph based on the first structured attribute features and generating a first node embedding and a first hyperedge embedding. The first hyperedge embedding is used to represent the resume feature embedding. The nodes of the resume feature hypergraph are used to represent the first structured attribute features. The hyperedges of the resume feature hypergraph are used to represent the resume data. The number of hyperedges of the resume feature hypergraph is equal to the number of resume data;
[0020] Constructing a job feature hypergraph based on the second structured attribute features and generating a second node embedding and a second hyperedge embedding. The second hyperedge embedding is used to represent the job feature embedding. The nodes of the job feature hypergraph are used to represent the second structured attribute features. The hyperedges of the job feature hypergraph are used to represent the job data. The number of hyperedges of the job feature hypergraph is equal to the number of job data.
[0021] Optionally, generating an initial resume text embedding based on the first unstructured attribute features and generating an initial job text embedding based on the second unstructured attribute features, including:
[0022] Obtain the sub-components corresponding to the first unstructured attribute features and the sub-components corresponding to the second unstructured attribute features. The sub-components include an attribute source, an attribute feature name, and an attribute feature value. The attribute source corresponding to the first unstructured attribute feature is resume data, and the attribute source corresponding to the second unstructured attribute feature is job data;
[0023] Input the sub-components corresponding to the first unstructured attribute features and the sub-components corresponding to the second unstructured attribute features into a pre-trained language model respectively to obtain the initial resume text embedding and the initial resume text embedding. The pre-trained language model is constructed based on BERT.
[0024] Optionally, the fusing of the resume feature embedding and the job feature embedding to obtain a resume structured embedding and a job structured embedding includes:
[0025] Learn the resume feature hypergraph based on the internal semantic associations between the attribute features of the resume data to obtain an initial resume hyperedge embedding;
[0026] Learn the job feature hypergraph based on the internal semantic associations between the attribute features of the job data to obtain an initial job hyperedge embedding;
[0027] Update the initial resume hyperedge embedding and the initial job hyperedge embedding based on the cross-semantic associations between the attribute features of the resume data and the attribute features of the job data to obtain the resume structured embedding and the job structured embedding.
[0028] Optionally, the fusing of the initial resume text embedding and the initial job text embedding to obtain a resume text embedding and a job text embedding includes:
[0029] Input the initial resume text embedding into a resume text attention layer and update the initial resume text embedding based on the internal semantic associations between the attribute features of the resume data to obtain a resume interaction feature;
[0030] Input the initial job text embedding into a job text attention layer and update the initial job text embedding based on the internal semantic associations between the attribute features of the job data to obtain a job interaction feature;
[0031] Connect the resume interaction feature and the job interaction feature and input them into an internal attention layer, and perform attribute interaction based on the cross-semantic associations between the attribute features of the resume data and the attribute features of the job data to obtain an interaction feature;
[0032] After splitting the interaction features, input them into an average pooling layer to obtain the resume text embedding and the job text embedding.
[0033] Optionally, predicting the matching probabilities of multiple resume data and multiple job data based on the total feature representation and determining the recommendation result based on the prediction result includes:
[0034] Input the total feature representation into an external attention layer for interactive fusion to obtain a resume-job feature representation;
[0035] Input the resume-job feature representation into an average pooling layer and a multi-layer perceptron for processing to obtain an output feature;
[0036] Calculate the matching probability between any resume data and any job data based on the output feature to obtain a prediction result;
[0037] Based on the prediction result, perform two-way recommendation on the resume data and the job data whose matching probability is greater than or equal to the threshold to obtain a recommendation result.
[0038] In a second aspect, an online recruitment two-sided reciprocal recommendation device based on a hypergraph according to an embodiment of the present invention further includes:
[0039] A feature extraction module, configured to respectively extract features from multiple resume data and multiple job data to obtain the attribute features of each resume data and the attribute features of each job data, where the attribute features of the resume data include first structured attribute features and first unstructured attribute features, and the attribute features of the job data include second structured attribute features and second unstructured attribute features;
[0040] An embedding generation module, configured to generate a resume feature embedding based on the first structured attribute feature, generate a job feature embedding based on the second structured attribute feature, generate an initial resume text embedding based on the first unstructured attribute feature, and generate an initial job text embedding based on the second unstructured attribute feature;
[0041] A feature fusion module, configured to fuse the resume feature embedding and the job feature embedding to obtain a resume structured embedding and a job structured embedding, and fuse the initial resume text embedding and the initial job text embedding to obtain a resume text embedding and a job text embedding;
[0042] A prediction and recommendation module, configured to predict the matching probabilities of multiple resume data and multiple job data based on the total feature representation and determine the recommendation result based on the prediction result, where the total feature representation is obtained by splicing the resume structured embedding, the job structured embedding, the resume text embedding, and the job text embedding.
[0043] In a third aspect, an embodiment of the present invention provides an electronic device, including: a memory, a processor, and a program stored on the memory and executable on the processor; the processor is configured to read the program in the memory to implement the steps in the hypergraph-based online recruitment bilateral reciprocal recommendation method as described in the first aspect.
[0044] In a fourth aspect, an embodiment of the present invention provides a readable storage medium for storing a program, which when executed by a processor implements the steps in the hypergraph-based online recruitment bilateral reciprocal recommendation method as described in the first aspect.
[0045] The beneficial effects of the present invention are as follows. In the embodiments of the present invention, first, the attribute features of resume data and job data are respectively obtained to obtain the structured and unstructured attribute features of resume data, and the structured and unstructured attribute features of job data. Through the above method, the information contained in different types of data in resume data and job data is fully considered, improving the accuracy and comprehensiveness of feature extraction. Then, for the structured and unstructured attribute features, corresponding methods are respectively used to generate their corresponding feature embeddings to obtain a more accurate embedding representation. Then, the embeddings of the structured and unstructured attribute features are fused, fully considering the needs of job seekers and recruiters, and predictions are made based on the fused total feature representation to obtain the final recommendation result, realizing the bilateral reciprocal recommendation for job seekers and recruiters. BRIEF DESCRIPTION OF THE DRAWINGS
[0046] FIG Figure 1 is a flowchart of the hypergraph-based online recruitment bilateral reciprocal recommendation method provided by an embodiment of the present invention;
[0047] FIG Figure 2a is a schematic diagram of the hypergraph-based online recruitment bilateral reciprocal recommendation model framework provided by an embodiment of the present invention;
[0048] FIG Figure 2b is Figure 2a a schematic diagram of the feature extraction module in FIG
[0049] FIG Figure 2c is Figure 2a a schematic diagram of the embedding generation module in FIG
[0050] FIG Figure 2d is Figure 2a a schematic diagram of the feature fusion module in FIG
[0051] FIG Figure 2e is Figure 2a a schematic diagram of the training and prediction module in FIG
[0052] FIG Figure 3Schematic diagrams of the experimental performance of each model on different evaluation metrics;
[0053] Appendix Figure 4a Schematic diagrams of the influence results of structured data and text data;
[0054] Appendix Figure 4b Schematic diagrams of the influence results of the internal attention layer and the external attention layer;
[0055] Appendix Figure 5 Schematic diagram of the online recruitment bilateral reciprocal recommendation device based on a hypergraph provided by an embodiment of the present invention;
[0056] Appendix Figure 6 Schematic diagram of the structure of an electronic device provided by an embodiment of the present invention. Detailed implementation manners
[0057] In the embodiments of the present application, the term "and / or" describes the association relationship of associated objects, indicating that three relationships may exist. For example, A and / or B may represent: A exists alone, A and B exist simultaneously, and B exists alone. The character " / " generally represents an "or" relationship between the associated objects before and after. In the embodiments of the present application, the term "plurality" refers to two or more, and other quantifiers are similar. The terms "first", "second", etc. in the description and claims of the present application are used to distinguish similar objects, rather than to describe a specific order or sequence. It should be understood that such terms can be interchanged under appropriate circumstances so that the embodiments of the present application can be implemented in an order other than those illustrated or described herein, and the objects distinguished by "first" and "second" are usually of the same category, and the number of objects is not limited. For example, the first object may be one or more.
[0058] Next, the technical solutions in the embodiments of the present application will be clearly and completely described with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present application without creative efforts shall fall within the protection scope of the present application.
[0059] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the technical field to which this application belongs. The terms used herein are only for the purpose of describing the embodiments of this application and are not intended to limit this application.
[0060] The embodiments of the present application provide a hypergraph-based online recruitment bilateral reciprocal recommendation method, device and electronic device, aiming to obtain better recommendation results and achieve bilateral reciprocity between job seekers and recruiters.
[0061] Please refer to Figure 1 , Figure 1 , which is a schematic flowchart of the online recruitment bilateral reciprocal recommendation method based on hypergraph provided by the embodiment of the present invention. The method specifically includes the following steps:
[0062] Step 101: Extract features from multiple resume data and multiple job data respectively to obtain the attribute features of each resume data and the attribute features of each job data. The attribute features of the resume data include the first structured attribute features and the first unstructured attribute features, and the attribute features of the job data include the second structured attribute features and the second unstructured attribute features.
[0063] Step 102: Generate resume feature embeddings based on the first structured attribute features, generate job feature embeddings based on the second structured attribute features, generate initial resume text embeddings based on the first unstructured attribute features, and generate initial job text embeddings based on the second unstructured attribute features.
[0064] Step 103: Fuse the resume feature embeddings and the job feature embeddings to obtain resume structured embeddings and job structured embeddings, and fuse the initial resume text embeddings and the initial job text embeddings to obtain resume text embeddings and job text embeddings.
[0065] Step 104: Predict the matching probabilities of multiple resume data and multiple job data based on the total feature representation and determine the recommendation results based on the prediction results. The total feature representation is obtained by splicing the resume structured embeddings, the job structured embeddings, the resume text embeddings, and the job text embeddings.
[0066] In this embodiment, multiple resume data and multiple job data are obtained, and the multiple resume data and multiple job data are matched by the online recruitment bilateral reciprocal recommendation method provided by the embodiment of the present invention, so as to obtain the recommendation results that are mutually beneficial to both job seekers and recruiters.
[0067] As a specific embodiment, assume that R is the set of resume data of job seekers, J is the set of job data of recruiters, and given a resume data r∈R of a job seeker and a job data j∈J of a recruiter. First, the attribute features of the resume data r can be represented as where f i r is the i-th attribute feature of the resume data. Similarly, the attribute features of the job data j can be represented as where f i jIt is the i-th feature of job position data i. The binary label indicating the matching result of R and J is y<R,J> ∈ {0, 1}. Among them, 0 indicates that the resume data does not match the job position data, and 1 indicates that the resume data matches the job position data.
[0068] The online recruitment two-sided reciprocal recommendation task based on hypergraph in this embodiment can be described as learning a prediction function from the preferences of resume data and the requirements of job position data. Specifically, first, learn the latent representations of resume data and job position data, as well as their corresponding latent semantic associations. Then, train a classification function to predict the matching degree between resume data and job position data. The above process can be described as: Y′<R,J> = F(R,J), where Y′<R,J> is the predicted matching degree between resume data and job position data, and F(R,J) is the prediction function.
[0069] Please refer to Figures 2a - 2e , the embodiment of the present invention also provides a hypergraph-based online recruitment two-sided reciprocal recommendation model framework, which can be used to train a hypergraph-based online recruitment two-sided reciprocal recommendation model, and the hypergraph-based online recruitment two-sided reciprocal recommendation model can be used to implement the hypergraph-based online recruitment two-sided reciprocal recommendation method provided by the embodiment of the present invention. As Figure 2a shown, this framework consists of four key modules, namely feature extraction, embedding generation, feature fusion, and training and prediction. The four key modules will be introduced separately below.
[0070] In step 101, the content in resume data and job position data can be divided into structured data and unstructured data according to the data type. Feature extraction based on the structured data and unstructured data in resume data and job position data can obtain the attribute features of resume data and the attribute features of job position data.
[0071] Optionally, in some embodiments, step 101 includes:
[0072] Perform feature extraction on the structured data in multiple pieces of the resume data to obtain the first structured attribute feature of each piece of the resume data;
[0073] Perform feature extraction on the unstructured data in multiple pieces of the resume data to obtain the first unstructured attribute feature of each piece of the resume data;
[0074] Perform feature extraction on the structured data in multiple pieces of the job position data to obtain the second structured attribute feature of each piece of the job position data;
[0075] Perform feature extraction on the unstructured data in multiple pieces of the job position data to obtain the second unstructured attribute feature of each piece of the job position data.
[0076] It should be understood that there is no limitation on the order of feature extraction for structured and unstructured data in resume data and job data. Exemplarily, as Figure 2b shown, feature extraction can be performed simultaneously on the structured data in the resume data, the unstructured data in the resume data, the structured data in the job data, and the unstructured data in the job data.
[0077] In specific implementation, the resume data and the job data further include multi-attribute data. The multi-attribute data usually contains several phrases, which have both the hierarchical semantic attributes of structured attribute features and the text semantic attributes of unstructured attribute features. The multi-attribute data can be regarded as both structured data and text data at the same time.
[0078] The attribute features include structured attribute features and unstructured attribute features. The structured attribute features are mainly used to describe hard skills, and specifically can include categorical attribute features, numerical attribute features, and multi-attribute features. The unstructured attribute features are mainly used to describe soft skills, including text attribute features and multi-attribute features. As shown in Table 1 below, in specific implementation, the multi-attribute features are regarded as a type of unstructured attribute features and are treated in the way of unstructured attribute embedding generation.
[0079] As a specific embodiment, as shown in Table 1, a total of 12 attribute features of the resume data are extracted from the resume data. Among them, 7 are structured attribute features and 5 are unstructured attribute features. The structured attribute features include current residence, expected work city, previous salary, expected salary, educational background, age, and start work date. The unstructured attribute features include work experience, previous work industry, expected work industry, previous job category, and expected job category.
[0080] Table 1 Description of Attribute Feature Nodes of Resume Data
[0081]
[0082] As a specific embodiment, as shown in Table 2, a total of 9 attribute features of the job data are extracted from the job data. Among them, 7 are structured attribute features and 2 are unstructured attribute features. The structured attribute features include city, minimum work experience, minimum education level, whether business trips are required, number of people required, maximum monthly salary, and minimum monthly salary. The unstructured attribute feature is job information, which is a combination of the job name attribute (job subcategory) and the job description attribute information (job information).
[0083] Table 2 Description of Attribute Feature Nodes of Job Data
[0084]
[0085] For structured data and unstructured data, the specific extraction methods may be different when extracting different types of attribute features. As a specific example, for categorical attribute features, one-hot encoding is used to obtain the initial feature vector corresponding to the categorical attribute features. For numerical attribute features, node serialization is utilized and represented using indices. Specifically, the data of each numerical attribute feature is de-duplicated and sorted to ensure that the feature values of each numerical attribute feature correspond to unique identifiers. For unstructured attribute features, further data preprocessing can be performed. Specifically, data cleaning is first carried out to delete missing value data, such as null values. Then, special symbols in the text, such as the "|" symbol, are deleted. Finally, the HIT stop word list and TF-IDF are used to filter out noise data. The specific methods for feature extraction of resume data and job data can be referred to the descriptions in the related technologies and will not be elaborated here.
[0086] In step 101, feature extraction is performed on each resume data and each job data, and their corresponding attribute features are obtained. According to the different data types of the attribute features, for both resume data and job data, their corresponding attribute features are divided into structured attribute features and unstructured attribute features. In subsequent processing, different methods can be adopted for various processes for structured attribute features and unstructured attribute features.
[0087] Optionally, in some embodiments, generating a resume feature embedding based on the first structured attribute feature and generating a job feature embedding based on the second structured attribute feature includes:
[0088] Constructing a resume feature hypergraph based on the first structured attribute feature and generating a first node embedding and a first hyperedge embedding, where the first hyperedge embedding is used to represent the resume feature embedding, the nodes of the resume feature hypergraph are used to represent the first structured attribute features, the hyperedges of the resume feature hypergraph are used to represent the resume data, and the number of hyperedges of the resume feature hypergraph is equal to the number of resume data;
[0089] Constructing a job feature hypergraph based on the second structured attribute feature and generating a second node embedding and a second hyperedge embedding, where the second hyperedge embedding is used to represent the job feature embedding, the nodes of the job feature hypergraph are used to represent the second structured attribute features, the hyperedges of the job feature hypergraph are used to represent the job data, and the number of hyperedges of the job feature hypergraph is equal to the number of job data.
[0090] As Figure 2c shown, in this embodiment, a resume feature hypergraph is constructed based on the first structured attribute feature, and at the same time, a job feature hypergraph is constructed based on the second structured attribute feature, and resume feature embedding and job feature embedding are respectively generated. Specifically as follows:
[0091] First, according to the first structured attribute features of the resume data and the second structured attribute features of the job data, two hypergraphs are constructed respectively, namely the resume feature hypergraph and the job feature hypergraph. Hypergraphs have a strong ability to depict and mine the non-linear associations between data samples. In a hypergraph, each node corresponds to the value of a specific attribute feature, while diverse hyperedges represent various resume data or job data. Hyperedges can connect any number of nodes to more accurately model multi-dimensional relationships.
[0092] Specifically, the nodes of the resume feature hypergraph are used to represent the first structured attribute features, the hyperedges of the resume feature hypergraph are used to represent the resume data, the nodes of the job feature hypergraph are used to represent the second structured attribute features, and the hyperedges of the job feature hypergraph are used to represent the job data. The total number of hyperedges in the resume feature hypergraph corresponds to the cumulative quantity of the job seeker's resume data, while the total number of hyperedges in the job feature hypergraph reflects the total quantity of the available job data.
[0093] Define the resume feature hypergraph as G R ={V R ,E R}, where V R represents the node set of the resume feature hypergraph, and E R represents the hyperedge set of the resume feature hypergraph. Define the job feature hypergraph as G J ={V J ,E J}, where V J represents the node set of the job feature hypergraph, and E J represents the hyperedge set of the job feature hypergraph.
[0094] Optionally, in some embodiments, generating the initial resume text embedding based on the first unstructured attribute feature and generating the initial job text embedding based on the second unstructured attribute feature includes:
[0095] Obtain the sub-components corresponding to the first unstructured attribute feature and obtain the sub-components corresponding to the second unstructured attribute feature. The sub-components include an attribute source, an attribute feature name, and an attribute feature value. The attribute source corresponding to the first unstructured attribute feature is resume data, and the attribute source corresponding to the second unstructured attribute feature is job data;
[0096] Input the sub-components corresponding to the first unstructured attribute feature and the sub-components corresponding to the second unstructured attribute feature into a pre-trained language model respectively to obtain the initial resume text embedding and the initial resume text embedding. The pre-trained language model is constructed based on BERT.
[0097] For unstructured attribute features, further data processing of the unstructured attribute features is required. First, the unstructured attribute features are processed into sub-components consisting of an attribute source, an attribute feature name, and an attribute feature value. Then, BERT is used as a text encoder to obtain an initial resume text embedding and an initial job text embedding. Since multivariate attribute data has a hierarchical semantic relationship between attribute values, such as "Real Estate / Construction / Building Materials / Engineering". In some embodiments, they are regarded as four attribute values associated with the same attribute feature to preserve the data hierarchy of the attributes. Therefore, the sub-component corresponding to each attribute feature consists of three semantically related parts: an attribute source (resume data or job data), an attribute feature name (age, city, etc.), and an attribute feature value (30, Beijing, etc.). In some embodiments, the pre-trained language model BERT-Base-Chinese is used as a text encoder to project the unstructured attribute features into the same semantic space.
[0098] As a specific embodiment, for the first unstructured attribute feature of the resume data, a first source embedding S corresponding to the attribute source can be obtained R , a first key embedding corresponding to the attribute feature name , and a first value embedding corresponding to the attribute feature value Similarly, for the second unstructured attribute feature of the job data, a second source embedding S corresponding to the attribute source can be obtained J , a second key embedding corresponding to the attribute feature name , and a second value embedding corresponding to the attribute feature value
[0099] The above process can be expressed by the following formula:
[0100]
[0101] where S R is used to represent the embedding representation of the source of the attribute feature as resume data, and S J is used to represent the embedding representation of the source of the attribute feature as job data. is used to represent the embedding representation of the attribute field k of the p-th feature in the resume data. is used to represent the embedding representation of the attribute feature value v of the p-th feature in the resume data. is used to represent the embedding representation of the attribute name k of the q-th feature in the job data. is used to represent the embedding representation of the attribute feature value v of the q-th feature in the job data. Bert(·) is used to represent the BERT model. represents the attribute field k corresponding to the p-th attribute feature in the resume data. Represents the attribute feature value v corresponding to the p-th feature in the resume data, Represents the attribute field k corresponding to the q-th attribute feature in the job data, Represents the attribute feature value v corresponding to the p-th feature in the job data, s R Used to indicate that the attribute source is the resume, s J Used to indicate that the attribute source is the job.
[0102] In the specific implementation, both the job information attributes of the recruitment information and the work experience attributes of the job seeker's resume contain long text data. In this embodiment, the multi-head attention layer included in the Text Encoder is used for information extraction. The BERT model consists of multiple stacked Transformer encoders, and each encoder layer contains two parts, namely multi-head self-attention and a feed-forward neural network. Since BERT is bidirectional, it is more accurate in understanding long texts, capturing complex language structures, and the associations between texts.
[0103] In this embodiment, the sub-components corresponding to the unstructured attribute features are obtained, namely the attribute source, the attribute feature name, and the attribute feature value, and then the BERT model is used as a text encoder for representation learning and long text information extraction, obtaining the initial resume text embedding and the initial job text embedding.
[0104] During the online recruitment process, there are various semantic associations between the resume data and the job data, such as internal semantic associations and cross-semantic associations. Internal semantic associations refer to the semantic associations existing between various attribute features from the same data source, that is, between various attribute features of the resume data and between various attribute features of the job data. For example, the expected working industry and the expected job category in the resume data respectively represent the job seeker's preferences for the expected industry and the expected job category, and there is a hierarchical semantic relationship between these two attribute features. Cross-semantic associations refer to the semantic associations existing between various attribute features from different data sources, that is, the semantic associations between the attribute features of the resume data and the attribute features of the job data. For example, the expected job category in the resume data and the job sub-category in the job data both represent the job category and have cross-semantic associations, except that the former represents the job seeker's preference and the latter represents the job requirements. To achieve the two-way matching of the resume and the job, the needs of both sides of the recruitment process need to be considered.
[0105] In step 103, a hierarchical interaction mechanism is used to capture the semantic associations of complex interactions between attributes. First, the structured data will be described below. Optionally, in some embodiments, the fusion of the resume feature embedding and the job feature embedding to obtain the resume structured embedding and the job structured embedding includes:
[0106] Learning the resume feature hypergraph based on the internal semantic associations between the attribute features of the resume data to obtain an initial resume hyperedge embedding;
[0107] Learning the job feature hypergraph based on the internal semantic associations between the attribute features of the job data to obtain an initial job hyperedge embedding;
[0108] Updating the initial resume hyperedge embedding and the initial job hyperedge embedding based on the cross-semantic associations between the attribute features of the resume data and the attribute features of the job data to obtain the resume structured embedding and the job structured embedding.
[0109] Please refer to Figure 2d , in this embodiment, a resume hypergraph learning module and a job hypergraph learning module are involved, which are respectively used to capture the internal semantic associations between the attribute features of resume data and the internal semantic associations between the attribute features of job data. The resume feature hypergraph is learned through the resume hypergraph learning module to capture the internal semantic associations between the attribute features of resume data. Similarly, the job feature hypergraph is learned through the job hypergraph learning module to capture the internal semantic associations between the attribute features of job data.
[0110] The following takes the resume feature hypergraph as an example for illustration. Suppose the resume feature hypergraph is denoted as G R ={V R , E R}, V R represents the node set of the resume feature hypergraph, and E R represents the hyperedge set of the resume feature hypergraph. is the hyperedge set of the resume feature hypergraph, N e is the number of hyperedges of the resume feature hypergraph, represents the j-th hyperedge in the l-th layer of the resume feature hypergraph neural network. is the set of hyperedge representations containing the node v.
[0111] Taking the resume feature hypergraph as an example, the internal semantic association can be regarded as a two-stage message passing process, that is, aggregating node information to hyperedges and aggregating hyperedge information to nodes, which is specifically expressed as follows:
[0112]
[0113] agg(S)=LayerNorm(Y + FFN(Y));
[0114] Y = LayerNorm(S + MultiHeadAttention(S));
[0115]
[0116] In the resume feature hypergraph, represents the embedding representation of the j-th hyperedge in the l-th layer obtained after feature fusion, represents the embedding representation of the i-th node in the (l - 1)-th layer obtained after feature fusion, represents the embedding representation of the i-th node in the l-th layer obtained after feature fusion, represents the embedding representation of the j-th hyperedge in the (l - 1)-th layer before feature fusion, is the input set embedding. When calculating the hyperedge , the corresponding input set is the node. When calculating the node , the corresponding input set is the hyperedge. FFN represents the feed-forward neural network, LayerNorm represents layer normalization, MultiHeadAttention represents the multi-head self-attention mechanism, and Softmax represents the softmax function. is the embedding of S after the multi-head self-attention mechanism, is a hyperparameter, h is the number of attention heads, d is the dimension, and the above equation is implemented using a 2-layer MLP.
[0117] To implement the two-stage message passing process, aggregating node information to hyperedges and hyperedge information to nodes, in this embodiment, the agg(·) function is constructed. Specifically, the agg(·) function is responsible for two-stage information passing: from nodes to edges and from edges to nodes. The input sets are different for different outputs. Specifically, in this embodiment, the self-attention function is adopted because it has strong expressive ability and can identify the most relevant elements in the set for information passing.
[0118] It should be understood that in this embodiment, positional encoding is not adopted. The embeddings of the hyperedge E (L) and the node V (L) in the last layer are obtained by stacking L Transformer layers together. Exemplarily, as an alternative implementation, the number of layers is set to L = 3.
[0119] Since the hypergraph is very large and dense, although the embeddings after message passing are very different, they are not easy to distinguish. Therefore, in some embodiments, the PairNorm technique is introduced to solve the over-smoothing problem. Specifically, additional normalization is added to the embeddings of each layer to keep the total pairwise embedding distance between layers unchanged.
[0120] Taking the hyperedge embedding as an example, the specific processing method of the hyperedge embedding is specifically expressed as follows:
[0121]
[0122] Among them, Denote the central representation of the hyperedge embedding, \(e\). j Denote the embedding representation of a certain hyperedge for pairnorm, \(e'\). j Denote the hyperedges in the hyperedge set (i.e., each \(e\)). j When performing pairnorm processing, all hyperedges in the hyperedge set are used for joint operation. \(|\varepsilon|\) represents the number of hyperedges. Denote the square of the Frobenius norm. Denote the square of the L2 norm. The superscript \(c\) indicates that the embedding is centralized. The node embeddings are processed in a similar manner, which is not specifically defined here.
[0123] Through the above-mentioned resume hypergraph learning module, the initial resume node embeddings and initial resume hyperedge embeddings are obtained. In the resume feature hypergraph, a hyperedge represents a resume, and a hyperedge contains the comprehensive representation of all node embeddings of the resume. The node features are composed of attributes and attribute values. Therefore, the hyperedge embedding of the resume feature hypergraph is regarded as the embedding of the resume data. Similarly, the structured attribute features of the job data are processed in the same way, and the initial job node embeddings and initial job hyperedge embeddings are respectively derived. The hyperedge embedding of the job feature hypergraph is regarded as the embedding of the job data.
[0124] As Figure 2d shown, in this embodiment, a connection layer and an inner attention layer (Inner Attention Layer) are used to capture the cross-semantic associations between the attribute features of the resume data and the attribute features of the job data. Finally, through a splitting operation and an average pooling layer, the resume structured embedding and the job structured embedding
[0125] For the unstructured attribute features, optionally, in some embodiments, the fusion of the initial resume text embedding and the initial job text embedding to obtain the resume text embedding and the job text embedding includes:
[0126] Input the initial resume text embedding into the resume text attention layer and update the initial resume text embedding based on the internal semantic associations between the attribute features of the resume data to obtain the resume interaction features.
[0127] Input the initial job text embedding into the job text attention layer and update the initial job text embedding based on the internal semantic associations between the attribute features of the job data to obtain the job interaction features.
[0128] After connecting the resume interaction features and the job interaction features, input them into the inner attention layer, and perform attribute interaction based on the cross-semantic association between the attribute features of the resume data and the attribute features of the job data to obtain interaction features;
[0129] After splitting the interaction features, input them into the average pooling layer to obtain the resume text embedding and the job text embedding.
[0130] In this embodiment, a Resume Text Attention Layer and a Job Text Attention Layer are used to capture the internal semantic association between the attribute features from the same attribute source. Input the initial resume text embedding into the resume text attention layer to obtain resume interaction features Input the initial job text embedding into the job text attention layer to obtain job interaction features The above process can be characterized as follows:
[0131]
[0132]
[0133] Among them, MultiHeadAttention(·) is used to represent the multi-head attention mechanism, and CONCAT(·) is used to represent the concatenation operation, is used to represent the embedding representation of the i-th resume, M is the number of resumes, is used to represent the embedding representation of the i-th job, N is the number of jobs, ⊕ is the concatenation symbol, S R is used to represent the embedding representation of the source of the attribute feature being the resume data, S J is used to represent the embedding representation of the source of the attribute feature being the job data, is used to represent the embedding representation of the attribute feature name k corresponding to the m-th attribute feature in the resume data, is used to represent the embedding representation of the attribute feature value v corresponding to the m-th attribute feature in the resume data, is used to represent the embedding representation of the attribute feature name k corresponding to the n-th attribute feature in the job data, is used to represent the embedding representation of the attribute feature value v corresponding to the n-th attribute feature in the job data.
[0134] Then, this embodiment provides an Inner Attention Layer, which uses the multi-head attention mechanism to capture the cross-semantic association between the attribute features from different attribute sources. Specifically, first, the resume interaction features and the job interaction features Connect them, and then use the multi-head attention layer to perform attribute interaction on the attribute features from different sources to obtain the interaction feature X after cross-semantic association atten Then, in order to interact with the subsequent structured features, the interaction feature X is further atten split into resume text interaction features and job text interaction features Finally, after passing through the average pooling layer, the embedding of unstructured data is obtained, that is, the resume text embedding and job text embedding The above process can be described as follows:
[0135]
[0136] X atten = MultiHeadAttention(X concant );
[0137]
[0138] Among them, Split(·) is used to represent the splitting operation, and AvgPool(·) is used to represent the average pooling operation.
[0139] In this embodiment, the resume text self-attention layer and the job text self-attention layer are used to capture the internal semantic associations of resume data and job data respectively. Then, the cross-semantic interaction between resume data and job data is captured through the internal attention layer. Finally, after the splitting operation and the average pooling layer, the resume text embedding and job text embedding
[0140] Most of the data in work experience and job description are semantically similar but expressed differently. For example, "office director" and "data management" in work experience have semantic overlap with the third point in the job description, "familiar with laws and policies related to employment, compensation, insurance, training, etc.". In addition, terms such as "reception", "meeting", "analysis ability" in work experience also have semantic overlap with the fourth point in the job description, "having good communication skills, analysis and judgment, decision-making, overall planning and organizational coordination abilities". In this embodiment, an internal attention layer is designed to learn the cross-semantic association information from different attribute sources, that is, the association information between resume and job, which can further improve the accuracy of the matching result.
[0141] Such as Figure 2a and Figure 2eAs shown in the figure, in order to obtain the comprehensive matching degree of resume data and job data, a training and prediction module is set in this embodiment. First, the structured data and unstructured data are concatenated through a fusion layer, and then the feature fusion of resume data and job data is achieved through an external attention layer. Finally, the classification module uses the MLP to obtain the matching prediction result of resume data and job data.
[0142] Optionally, in some embodiments, step 104 includes:
[0143] Input the total feature representation into the external attention layer for interactive fusion to obtain the resume-job feature representation;
[0144] Input the resume-job feature representation into the average pooling layer and the multi-layer perceptron for processing to obtain the output feature;
[0145] Calculate the matching probability between any resume data and any job data based on the output feature to obtain the prediction result;
[0146] Based on the prediction result, perform two-way recommendation on the resume data and the job data with the matching probability greater than or equal to the threshold to obtain the recommendation result.
[0147] Specifically, as Figure 2d shown, use the Concatenate Layer to concatenate the resume structured embedding, job structured embedding, resume text embedding, and job text embedding. Specifically, concatenate the resume structured embedding and the job structured embedding and the resume text embedding and the job text embedding obtained after text pre-training model learning and hierarchical interaction mechanism to obtain the total feature representation X RJ of resume data and job data, which can be specifically expressed as follows:
[0148]
[0149] Among them, CONCAT(·) is used to represent the concatenation operation.
[0150] Then, use the Outer Attention Layer to perform structured and unstructured feature interaction and fusion between resume data and job data to obtain the resume-job feature representation In some embodiments, to reduce the feature dimension and prevent overfitting, the resume job feature representation is first input into an average pooling layer for average pooling processing. Then, a multi-layer perceptron (MLP) is used for linear layer transformation to obtain the output feature H out . Finally, through sigmoid function prediction, the matching probability between the resume and the job is calculated to obtain the prediction result of the resume and job matching. The above process can be expressed as follows:
[0151]
[0152] where MLP(·) is used to represent the MLP, AvgPool(·) is used to represent the average pooling operation, and the predicted matching probability of the resume data and the job data is It should be understood that R is the set of resume data of job seekers, J is the set of job data for recruitment, and when predicting the matching probability, the matching probability of any resume r in R and any job j in J can be obtained.
[0153] In this embodiment, an outer attention layer is used to learn information from different data structures, fully considering the preferences of job seekers and the needs of recruiters. For example, there is a matching relationship between the structured data "age 33" in the resume and the unstructured data "age 30 - 40" in the job description. Through the method of this embodiment, the accuracy of the matching result can be further improved. The method provided by the embodiment of the present invention can learn the potential matching patterns between resumes and real-life jobs, effectively utilize structured and unstructured data and the semantic information between resumes and jobs, and provide mutual decision-making support for both sides in online recruitment.
[0154] As a specific embodiment, the loss function for model training is defined as cross-entropy loss, and the calculation formula is as follows:
[0155]
[0156] where y i and y i ′ represent the true matching label and the model prediction result respectively.
[0157] The beneficial effects of the method provided by the embodiments of the present invention are illustrated through specific experiments below. In the experiment, both the feature embedding and the hypergraph node embedding are set to 64 dimensions. The Adam optimizer is used to optimize the parameters during model training. An early stopping mechanism is adopted. If the accuracy on the validation set does not improve for five consecutive epochs, the training is terminated. To prevent overfitting, layer normalization and gradient clipping mechanisms are applied. The learning rate is set to 1e-06, the weight decay is set to 1e-09, and the mini-batch size is 16. The hidden size and the head size of the multi-head self-attention mechanism are set to 2048 and 8 respectively. To reduce the model complexity, the attribute features are padded or truncated to a fixed length: 256 for text-based attributes (such as job descriptions and work experience), and 16 for others. The baseline models are configured with similar parameter settings, and all models are fine-tuned to ensure a fair comparison. To evaluate the robustness of the model corresponding to the hypergraph-based online recruitment bilateral reciprocal recommendation method provided in this embodiment (subsequently referred to as the HRROR model), the dataset is divided into 70% for training, 15% for validation, and 15% for testing.
[0158] The baseline models are introduced first below. In this experiment, the following 4 models specifically designed for online recruitment are selected as baseline models: DPGNN, PJFNN, APJFNN, and BPJFNN. In addition, several popular generalized recommendation methods are selected for comparison: LightGCN, BERT, and BPR.
[0159] DPGNN: The dual-perspective graph neural network is a method based on graph representation learning, aiming to model the two-way selection preference by simulating the human-job matching process using dual perspectives.
[0160] PJFNN: The person-job matching neural network model is an end-to-end data-driven convolutional neural network (CNN) model. It designs a hierarchical representation structure and calculates the matching degree through cosine similarity to measure the distance between corresponding latent representations.
[0161] APJFNN: The ability-aware person-job matching neural network model is a model based on the end-to-end recurrent neural network (RNN). It designs four hierarchical ability-aware attention strategies and uses the bidirectional long short-term memory network (BiLSTM) to learn the representations of resume work experience and job ability requirements.
[0162] BPJFNN: The basic person-job matching neural network model is an RNN-based model and can be regarded as a simplified version of the APJFNN model. It uses BiLSTM to learn the representations of resume descriptions and job requirements.
[0163] LightGCN: LightGCN is a simplified graph convolutional neural network based on collaborative filtering for classification tasks. LightGCN only contains the most important components in GCN, namely neighborhood aggregation. LightGCN learns the embeddings of users and items by linearly propagating on the user-item interaction graph and uses the weighted sum of the embeddings learned at all layers as the final embedding.
[0164] BERT: BERT is a two-tower model with a text encoder that uses a language representation model called Bidirectional Encoder Representations from Transformers.
[0165] BPR: BPR is a method that can recommend items from implicit feedback and directly optimize ranking. BPR is a learning method based on bootstrapped sampling stochastic gradient descent and can also use the maximum a posteriori estimator derived from the Bayesian analysis of the problem to predict items.
[0166] In the experiment, the resume-job matching recommendation problem in online recruitment is conceptualized as a probability prediction task, and the prediction probability threshold is set to 0.5. If the predicted matching score of a sample exceeds 0.5, it is considered a positive match, indicating that both the resume and the job meet each other's needs and preferences. Otherwise, it is considered a mismatch. To comprehensively evaluate the performance of various models, the following five evaluation metrics are used in this experiment:.
[0167] Area Under Curve (AUC) of the ROC curve: Represents the area under the Receiver Operating Characteristic (ROC) curve, which is used to measure the ability of the model to distinguish positive and negative classes. The value range of AUC is from 0.5 to 1, where the value closer to 1 indicates better performance. An AUC of 0.5 means that the model's performance is no better than random guessing.
[0168] Accuracy: The ratio of the number of correctly predicted samples to the total number of samples. It is a common measure of overall classification performance.
[0169] Recall: Measures the proportion of actual positive samples that are correctly identified by the model, which measures the model's ability to detect positive samples.
[0170] Precision: The proportion of correctly predicted positive samples among the total number of predicted positive samples, which evaluates the accuracy of the model's prediction of positive samples.
[0171] The F1 score is the harmonic mean of precision and recall, which can balance the evaluation of the two metrics. In scenarios involving imbalanced datasets, it is particularly valuable because in such cases, it is crucial to consider both precision and recall simultaneously.
[0172] The experimental results are as Figure 3 shown in Table 3. According to the experimental results, it can be seen that the HRROR model corresponding to the hypergraph-based online recruitment bilateral reciprocal recommendation method provided in this embodiment is significantly better than all comparison models in all evaluation metrics, which proves that the hypergraph-based online recruitment bilateral reciprocal recommendation method provided in this embodiment can effectively provide information about job seeker resumes and job requirements and provide two-way recommendations for personnel matching.
[0173] Specifically, the HRROR model corresponding to the hypergraph-based online recruitment bilateral reciprocal recommendation method provided in this embodiment is significantly better than all baseline methods in different evaluation metrics. This clearly proves the effectiveness of the method provided in this embodiment in modeling structured data based on hypergraphs and unstructured text data based on a hierarchical interaction mechanism in the problem of resume-job matching reciprocal recommendation.
[0174] Table 3 Experimental Results
[0175]
[0176]
[0177] To further verify the effectiveness of each component of the HRROR model corresponding to the hypergraph-based online recruitment bilateral reciprocal recommendation method provided in this embodiment, several simplified model variants were designed in this experiment for ablation experiments:
[0178] HRROR w / o structured: This variant deletes the structured data in the resume data and job data.
[0179] HRROR w / o unstructured: This variant excludes the text data from the resume data and job data.
[0180] HRROR w / o inner: This variant omits the inner attention layer, which is used to capture cross-semantic associations from different attribute sources, especially the associations between resume structured data and job structured data, and between resume unstructured data and job unstructured data.
[0181] HRROR w / o outer: This variant excludes the outer attention layer, which is used to process resume data and job data of different data types while considering the needs of job seekers and recruiters.
[0182] The results of the ablation study are as Figure 4a and Figure 4b shown. As Figure 4a shown, the HRROR model corresponding to the hypergraph-based online recruitment bilateral reciprocal recommendation method provided in this embodiment is significantly better than the model variants with unstructured data and the model variants without unstructured data, which highlights that both structured data and unstructured text data provide important information for mutual resume-job recommendation modeling. As Figure 4b shown, the HRROR model corresponding to the hypergraph-based online recruitment bilateral reciprocal recommendation method provided in this embodiment is better than the model variants without the internal attention layer, which verifies the effectiveness of the internal attention layer in capturing cross-semantic associations from different attribute sources. In addition, the performance of the HRROR model corresponding to the hypergraph-based online recruitment bilateral reciprocal recommendation method provided in this embodiment is also better than the model variants without the external attention layer, which confirms the effectiveness of the external attention layer in information interaction and matching during the resume-job mutual recommendation process.
[0183] In this embodiment, separate feature hypergraphs are constructed for the structured attribute features corresponding to the resume data and the structured attribute features corresponding to the job data to obtain feature representations of high-order interactions within their respective domains. For the unstructured attribute features, a pre-trained language model is used in this embodiment to learn text representations. In addition, there are complex semantic relationships within the attribute features of the resume data or job data and between the attribute sources. To capture these semantic associations, a hierarchical interaction mechanism is designed for structured and unstructured data respectively in this embodiment. For the structured attribute features, a hypergraph learning module for capturing internal semantic associations and an internal attention layer for capturing cross-semantic associations are proposed. For the unstructured attribute features, a text self-attention layer is designed to capture internal semantic associations, and an internal attention layer is designed to capture cross-semantic associations. The complex semantic relationships are effectively captured through the above hierarchical interaction mechanism, generating powerful semantic embeddings for resumes and job postings. Subsequently, an external attention layer is also introduced in this embodiment to comprehensively consider the preferences and needs of job seekers and recruiters. Finally, a classification module is used for resume-job matching prediction. According to the above experimental results, the hypergraph-based online recruitment bilateral reciprocal recommendation method provided in this embodiment shows better performance compared with the recommendation models specifically designed for online recruitment and the general recommendation models.
[0184] Please refer to Figure 5 , the embodiment of the present invention also provides a hypergraph-based online recruitment bilateral reciprocal recommendation device 500, including:
[0185] A feature extraction module 501 is configured to perform feature extraction on multiple resume data and multiple job data respectively, so as to obtain the attribute features of each of the resume data and the attribute features of each of the job data. The attribute features of the resume data include a first structured attribute feature and a first unstructured attribute feature, and the attribute features of the job data include a second structured attribute feature and a second unstructured attribute feature;
[0186] An embedding generation module 502 is configured to generate a resume feature embedding based on the first structured attribute feature, generate a job feature embedding based on the second structured attribute feature, generate an initial resume text embedding based on the first unstructured attribute feature, and generate an initial job text embedding based on the second unstructured attribute feature;
[0187] A feature fusion module 503 is configured to fuse the resume feature embedding and the job feature embedding to obtain a resume structured embedding and a job structured embedding, and fuse the initial resume text embedding and the initial job text embedding to obtain a resume text embedding and a job text embedding;
[0188] A prediction and recommendation module 504 is configured to predict the matching probabilities of multiple resume data and multiple job data based on the total feature representation and determine a recommendation result based on the prediction result. The total feature representation is obtained by concatenating the resume structured embedding, the job structured embedding, the resume text embedding, and the job text embedding.
[0189] Optionally, the feature extraction module 501 includes:
[0190] A first extraction unit is configured to perform feature extraction on the structured data in multiple resume data to obtain the first structured attribute feature of each resume data;
[0191] A second extraction unit is configured to perform feature extraction on the unstructured data in multiple resume data to obtain the first unstructured attribute feature of each resume data;
[0192] A third extraction unit is configured to perform feature extraction on the structured data in multiple job data to obtain the second structured attribute feature of each job data;
[0193] A fourth extraction unit is configured to perform feature extraction on the unstructured data in multiple job data to obtain the second unstructured attribute feature of each job data.
[0194] Optionally, generating a resume feature embedding based on the first structured attribute feature and generating a job feature embedding based on the second structured attribute feature includes:
[0195] Construct a resume feature hypergraph based on the first structured attribute feature and generate a first node embedding and a first hyperedge embedding. The first hyperedge embedding is used to represent the resume feature embedding. The nodes of the resume feature hypergraph are used to represent the first structured attribute feature. The hyperedges of the resume feature hypergraph are used to represent the resume data. The number of hyperedges of the resume feature hypergraph is equal to the number of resume data.
[0196] Construct a job feature hypergraph based on the second structured attribute feature and generate a second node embedding and a second hyperedge embedding. The second hyperedge embedding is used to represent the job feature embedding. The nodes of the job feature hypergraph are used to represent the second structured attribute feature. The hyperedges of the job feature hypergraph are used to represent the job data. The number of hyperedges of the job feature hypergraph is equal to the number of job data.
[0197] Optionally, generating an initial resume text embedding based on the first unstructured attribute feature and generating an initial job text embedding based on the second unstructured attribute feature includes:
[0198] Obtain the sub-components corresponding to the first unstructured attribute feature and obtain the sub-components corresponding to the second unstructured attribute feature. The sub-components include an attribute source, an attribute feature name, and an attribute feature value. The attribute source corresponding to the first unstructured attribute feature is resume data, and the attribute source corresponding to the second unstructured attribute feature is job data.
[0199] Input the sub-components corresponding to the first unstructured attribute feature and the sub-components corresponding to the second unstructured attribute feature into a pre-trained language model respectively to obtain the initial resume text embedding and the initial resume text embedding. The pre-trained language model is constructed based on BERT.
[0200] Optionally, fusing the resume feature embedding and the job feature embedding to obtain a resume structured embedding and a job structured embedding includes:
[0201] Learn the resume feature hypergraph based on the internal semantic associations between the attribute features of the resume data to obtain an initial resume hyperedge embedding;
[0202] Learn the job feature hypergraph based on the internal semantic associations between the attribute features of the job data to obtain an initial job hyperedge embedding;
[0203] Update the initial resume hyperedge embedding and the initial job hyperedge embedding based on the cross-semantic associations between the attribute features of the resume data and the attribute features of the job data to obtain the resume structured embedding and the job structured embedding.
[0204] Optionally, fusing the initial resume text embedding and the initial job text embedding to obtain a resume text embedding and a job text embedding includes:
[0205] Input the initial resume text embedding into a resume text attention layer and update the initial resume text embedding based on the internal semantic associations between the attribute features of the resume data to obtain a resume interaction feature;
[0206] Input the initial job text embedding into a job text attention layer and update the initial job text embedding based on the internal semantic associations between the attribute features of the job data to obtain a job interaction feature;
[0207] Connect the resume interaction feature and the job interaction feature and input them into an internal attention layer to perform attribute interaction based on the cross-semantic associations between the attribute features of the resume data and the attribute features of the job data to obtain an interaction feature;
[0208] Split the interaction feature and input it into an average pooling layer to obtain the resume text embedding and the job text embedding.
[0209] Optionally, the prediction and recommendation module 504 includes:
[0210] An interaction fusion unit for inputting the total feature representation into an external attention layer for interaction fusion to obtain a resume-job feature representation;
[0211] A processing unit for inputting the resume-job feature representation into an average pooling layer and a multi-layer perceptron for processing to obtain an output feature;
[0212] A calculation unit for calculating the matching probability between any one of the resume data and any one of the job data based on the output feature to obtain a prediction result;
[0213] A recommendation unit for performing two-way recommendation on the resume data and the job data with the matching probability greater than or equal to a threshold based on the prediction result to obtain a recommendation result.
[0214] The hypergraph-based online recruitment bilateral reciprocal recommendation device 500 provided by the embodiments of the present application can execute the above method embodiments, and its implementation principle and technical effects are similar, which will not be elaborated here in this embodiment.
[0215] It should be noted that the division of units in the embodiments of this application is illustrative, merely a logical function division, and there may be other division methods in actual implementation. In addition, in each embodiment of this application, each functional unit may be integrated in a processing unit, may exist separately physically for each unit, or two or more units may be integrated in one unit. The above-mentioned integrated unit may be implemented in the form of hardware or in the form of a software functional unit.
[0216] If the above-mentioned integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it may be stored in a processor-readable storage medium. Based on such an understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of this technical solution, may be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to enable a computer device (which may be a personal computer, a server, or a network device, etc.) or a processor to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROM), random access memories (RAM), magnetic disks, or optical discs that can store program codes.
[0217] As Figure 6 shown, an embodiment of this application provides an electronic device 600, including: a memory 602, a processor 601, and a program stored on the memory 602 and executable on the processor 601; the processor 601 is used to read the program in the memory 602 to implement the steps in the online recruitment bilateral reciprocal recommendation method based on a hypergraph as described above.
[0218] The embodiments of the present application also provide a readable storage medium. A program is stored on the readable storage medium. When the program is executed by a processor, it implements each process of the above-mentioned embodiments of the online recruitment bilateral reciprocal recommendation method based on a hypergraph and can achieve the same technical effects. To avoid repetition, details are not described herein again. Among them, the readable storage medium may be any available medium or data storage device accessible by the processor, including but not limited to magnetic memories (such as floppy disks, hard disks, magnetic tapes, magneto-optical disks (MO), etc.), optical memories (such as compact disks (CD), digital versatile discs (DVD), Blu-ray discs (BD), high-definition versatile discs (HVD), etc.), and semiconductor memories (such as read-only memories (ROM), erasable programmable read-only memories (EPROM), electrically erasable programmable read-only memories (EEPROM), non-volatile memories (NAND FLASH), solid state disks (SSD) or solid state drives, etc.).
[0219] It should be noted that in this article, the term "including", "comprising" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not explicitly listed, or further includes elements inherent to such a process, method, article or device. Without further limitation, an element defined by the statement "including one..." does not exclude the existence of another identical element in the process, method, article or device including that element.
[0220] Through the description of the above embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus a necessary general hardware platform. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation. Based on such an understanding, the technical solution of the present application, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, disk, optical disc), and includes several instructions for causing a terminal (which can be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the methods described in various embodiments of the present application.
[0221] The embodiments of the present application have been described above in conjunction with the accompanying drawings. However, the present application is not limited to the above specific embodiments. The above specific embodiments are merely illustrative and not restrictive. Under the inspiration of the present application, those of ordinary skill in the art can also make many forms without departing from the purpose of the present application and the scope protected by the claims, and all of them belong to the protection scope of the present application.
Claims
1. A bilateral reciprocal recommendation method for online recruitment based on hypergraph, characterized by: include: Performing feature extraction on the plurality of resume data and the plurality of position data respectively to obtain attribute features of each of the resume data and attribute features of each of the position data, wherein the attribute features of the resume data include a first structured attribute feature and a first unstructured attribute feature, and the attribute features of the position data include a second structured attribute feature and a second unstructured attribute feature; Generate a resume feature embedding based on the first structured attribute feature, generate a position feature embedding based on the second structured attribute feature, generate an initial resume text embedding based on the first unstructured attribute feature, and generate an initial position text embedding based on the second unstructured attribute feature; The resume feature embedding and the position feature embedding are merged to obtain resume structured embedding and position structured embedding, and the initial resume text embedding and the initial position text embedding are merged to obtain resume text embedding and position text embedding; The matching probability of the plurality of resume data and the plurality of position data is predicted based on the total feature representation, and a recommendation result is determined based on the prediction result, wherein the total feature representation is obtained by concatenating the resume structured embedding, the position structured embedding, the resume text embedding and the position text embedding.
2. The method according to claim 1, characterized in that: The feature extraction is performed on the plurality of resume data and the plurality of position data respectively to obtain the attribute feature of each resume data and the attribute feature of each position data, including: Extracting features from the structured data in the plurality of resume data to obtain a first structured attribute feature of each resume data; Extracting features from the unstructured data in the plurality of resume data to obtain a first unstructured attribute feature of each of the resume data; Extracting features from the structured data in the plurality of position data to obtain a second structured attribute feature of each position data; Feature extraction is performed on the unstructured data in the plurality of position data to obtain a second unstructured attribute feature of each of the position data.
3. The method according to claim 1, characterized in that: The generating a resume feature embedding based on the first structured attribute feature, and generating a position feature embedding based on the second structured attribute feature, comprises: Based on the first structured attribute feature, a resume feature hypergraph is constructed and a first node embedding and a first hyperedge embedding are generated, wherein the first hyperedge embedding is used to represent the resume feature embedding, the nodes of the resume feature hypergraph are used to represent the first structured attribute feature, the hyperedges of the resume feature hypergraph are used to represent the resume data, and the number of hyperedges of the resume feature hypergraph is equal to the number of resume data; A position feature hypergraph is constructed based on the second structured attribute feature and a second node embedding and a second hyperedge embedding are generated. The second hyperedge embedding is used to characterize the position feature embedding. The nodes of the position feature hypergraph are used to characterize the second structured attribute feature. The hyperedges of the position feature hypergraph are used to characterize the position data. The number of hyperedges of the position feature hypergraph is equal to the number of the position data.
4. The method according to claim 1, characterized in that: The step of generating an initial resume text embedding based on the first unstructured attribute feature and generating an initial position text embedding based on the second unstructured attribute feature comprises: Obtaining a subcomponent corresponding to the first unstructured attribute feature, and obtaining a subcomponent corresponding to the second unstructured attribute feature, wherein the subcomponent includes an attribute source, an attribute feature name, and an attribute feature value, wherein the attribute source corresponding to the first unstructured attribute feature is resume data, and the attribute element corresponding to the second unstructured attribute feature is position data; The subcomponent corresponding to the first unstructured attribute feature and the subcomponent corresponding to the second unstructured attribute feature are respectively input into a pre-trained language model to obtain the initial resume text embedding and the initial resume text embedding, and the pre-trained language model is constructed based on BERT.
5. The method according to claim 3, characterized in that: The step of fusing the resume feature embedding and the position feature embedding to obtain resume structured embedding and position structured embedding includes: Learning the resume feature hypergraph based on the internal semantic associations between the attribute features of the resume data to obtain an initial resume hyperedge embedding; Learning the position feature hypergraph based on the internal semantic associations between the attribute features of the position data to obtain an initial position hyperedge embedding; The initial resume hyperedge embedding and the initial position hyperedge embedding are updated based on the cross-semantic association between the attribute features of the resume data and the attribute features of the position data to obtain the resume structured embedding and the position structured embedding.
6. The method according to claim 1, characterized in that: The step of fusing the initial resume text embedding and the initial position text embedding to obtain a resume text embedding and a position text embedding includes: Embedding the initial resume text into an input resume text attention layer and updating the initial resume text embedding based on internal semantic associations between attribute features of the resume data to obtain a resume interaction feature; Embedding the initial job text into the job text attention layer and updating the initial job text embedding based on the internal semantic association between the attribute features of the job data to obtain job interaction features; The resume interaction feature and the position interaction feature are connected and input into the internal attention layer, and attribute interaction is performed based on the cross-semantic association between the attribute features of the resume data and the attribute features of the position data to obtain the interaction feature; The interactive features are split and then input into an average pooling layer to obtain the resume text embedding and the position text embedding.
7. The method according to claim 1, characterized in that: The predicting the matching probability of the plurality of resume data and the plurality of position data based on the total feature representation and determining the recommendation result based on the prediction result includes: Input the total feature representation into the external attention layer for interactive fusion to obtain the resume position feature representation; Input the resume position feature representation into the average pooling layer and the multi-layer perceptron for processing to obtain output features; Calculate the matching probability between any of the resume data and any of the position data based on the output features to obtain a prediction result; Based on the prediction result, the resume data and the position data whose matching probability is greater than or equal to a threshold are bidirectionally recommended to obtain a recommendation result.
8. A hypergraph-based online recruitment bilateral reciprocal recommendation device, characterized in that: include: A feature extraction module, used to extract features from the plurality of resume data and the plurality of position data respectively, to obtain attribute features of each of the resume data and attribute features of each of the position data, wherein the attribute features of the resume data include a first structured attribute feature and a first unstructured attribute feature, and the attribute features of the position data include a second structured attribute feature and a second unstructured attribute feature; An embedding generation module, configured to generate a resume feature embedding based on the first structured attribute feature, generate a position feature embedding based on the second structured attribute feature, generate an initial resume text embedding based on the first unstructured attribute feature, and generate an initial position text embedding based on the second unstructured attribute feature; A feature fusion module, used to fuse the resume feature embedding and the position feature embedding to obtain a resume structured embedding and a position structured embedding, and to fuse the initial resume text embedding and the initial position text embedding to obtain a resume text embedding and a position text embedding; A prediction and recommendation module is used to predict the matching probability of multiple resume data and multiple position data based on the total feature representation and determine the recommendation result based on the prediction result. The total feature representation is obtained by splicing the resume structured embedding, the position structured embedding, the resume text embedding and the position text embedding.
9. An electronic device, comprising: A memory, a processor, and a program stored in the memory and executable on the processor; wherein the processor is used to read the program in the memory to implement the steps in the hypergraph-based online recruitment bilateral reciprocal recommendation method as described in any one of claims 1 to 7.
10. A readable storage medium for storing a program, characterized in that: When the program is executed by a processor, the steps in the hypergraph-based online recruitment bilateral reciprocal recommendation method as described in any one of claims 1 to 7 are implemented.