Resume screening method and device, electronic equipment and storage medium

By using capsule network to screen resumes in enterprise recruitment, key characteristics and their weights are determined, the problem of low efficiency and poor accuracy of resume screening in the existing technology is solved, and efficient and accurate resume screening is achieved.

CN120338740APending Publication Date: 2025-07-18SHANGHAI JUNXING TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510483337.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-17
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

In the current technology, it is difficult to efficiently and accurately select candidates who best match the recruitment position in enterprise recruitment. Traditional methods are difficult to fully explore the complex characteristic relationships in resumes, and the screening efficiency is inefficient and poor accuracy.

Method used

By obtaining job description information, determining the key features and weights of alternative resumes, using the capsule network for feature extraction and matching probability calculation, and filtering out the target resume that best matches the job description information.

Benefits of technology

It improves the accuracy and efficiency of resume screening, reduces labor costs, can capture the relationship between multiple complex features in the resume, and realizes the accurate representation of job description information and underlying features.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120338740A_ABST
    Figure CN120338740A_ABST
Patent Text Reader

Abstract

The invention discloses a resume screening method and device, electronic equipment and a storage medium. The method comprises the following steps: determining key features and weights of the key features according to data of alternative resumes; determining a key vector of the alternative resume according to each key feature and the weight of each key feature; determining an output vector of each primary capsule according to a key vector of the alternative resume, a primary coefficient in trained weight coefficients and at least one primary capsule in the primary capsule layer; according to the output vector of each primary capsule, a high-level coefficient in the weight coefficient and at least one high-level capsule in the high-level capsule layer, determining the output vector of each high-level capsule; determining the matching probability of the alternative resumes according to the output vector of each high-level capsule; and screening out a target resume corresponding to the position description information from the alternative resumes. According to the technical scheme, the resume screening accuracy can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence, and particularly to a resume screening method, device, electronic device, and storage medium. Background Art

[0002] In today's enterprise recruitment process, in the face of a vast amount of resume information, how to efficiently and accurately screen out the candidates most suitable for the recruitment position has become a key problem to be solved urgently.

[0003] Traditional resume screening methods often rely on manual screening or simple keyword matching, etc., and it is difficult to fully explore the complex feature relationships in the resume and effectively utilize the data, resulting in low screening efficiency and poor accuracy. Summary of the Invention

[0004] The present invention provides a resume screening method, device, electronic device, and storage medium, which can improve the accuracy of resume screening.

[0005] According to one aspect of the present invention, a resume screening method is provided, and the method includes:

[0006] Obtain job description information;

[0007] For at least one alternative resume, determine the key features corresponding to the job description information and the weights of each of the key features according to the data of the alternative resume;

[0008] Determine the key vector of the alternative resume according to each of the key features and the weights of each of the key features;

[0009] Determine the output vector of each of the primary capsules according to the key vector of the alternative resume, the primary coefficient in the trained weight coefficients, and at least one primary capsule in the primary capsule layer;

[0010] Determine the output vector of each of the senior capsules according to the output vector of each of the primary capsules, the senior coefficient in the weight coefficients, and at least one senior capsule in the senior capsule layer;

[0011] Determine the matching probability between the alternative resume and the job description information according to the output vector of each of the senior capsules;

[0012] Screen out the target resume corresponding to the job description information from each of the alternative resumes according to the matching probability between each of the alternative resumes and the job description information.

[0013] According to another aspect of the present invention, a resume screening device is provided, and the device includes:

[0014] A description information acquisition module for acquiring job description information;

[0015] A feature selection module for determining, for at least one alternative resume, key features corresponding to the job description information and weights of each of the key features according to data of the alternative resume;

[0016] A capsule network for determining key vectors of the alternative resume according to each of the key features and weights of each of the key features;

[0017] The capsule network for determining output vectors of each of the primary capsules according to the key vectors of the alternative resume, primary coefficients in trained weight coefficients, and at least one primary capsule in the primary capsule layer;

[0018] The capsule network for determining output vectors of each of the secondary capsules according to the output vectors of each of the primary capsules, secondary coefficients in the weight coefficients, and at least one secondary capsule in the secondary capsule layer;

[0019] The capsule network for determining a matching probability between the alternative resume and the job description information according to the output vectors of each of the secondary capsules;

[0020] A resume screening module for screening out target resumes corresponding to the job description information from each of the alternative resumes according to the matching probability between each of the alternative resumes and the job description information. According to another aspect of the present invention, there is provided an electronic device, the electronic device comprising:

[0021] At least one processor; and

[0022] A memory communicatively connected to the at least one processor; wherein,

[0023] The memory stores a computer program executable by the at least one processor, and the computer program is executed by the at least one processor so that the at least one processor can execute the resume screening method according to any embodiment of the present invention.

[0024] According to another aspect of the present invention, there is provided a computer-readable storage medium storing computer instructions for causing a processor to implement the resume screening method according to any embodiment of the present invention when executed.

[0025] The technical solution of the embodiment of the present invention screens key features from the data of alternative resumes, effectively improves the representativeness of features, reduces redundant information, and inputs the key features into the primary capsule layer with trained weight coefficients for processing to obtain the output vector of the primary capsule, and inputs the key features into the advanced capsule layer with trained weight coefficients for processing to obtain the output vector of the advanced capsule. Finally, based on the output vector of the advanced capsule, the matching probability between the alternative resume and the job description information is determined, and then the target resume corresponding to the job description information is screened out from multiple alternative resumes. It can capture the relationships between various complex features in the resume. At the same time, using the weight coefficients of the trained capsule layer to process the vector can accurately represent the correlation between the job description information and the underlying features, and then integrate the underlying features to form an accurate representation of multiple complex features in the alternative resume. Detecting the matching probability between the alternative resume and the job description information based on the accurate representation can improve the accuracy of the matching probability, solve the problem in the prior art that it is difficult to mine and utilize the complex feature relationships in the resume, resulting in low screening efficiency and poor accuracy, can effectively mine and accurately represent the complex relationships in the resume, improve the accuracy of resume screening, and reduce labor costs at the same time, and can effectively improve the resume screening efficiency.

[0026] It should be understood that the content described in this part is not intended to identify the key or important features of the embodiments of the present invention, nor is it used to limit the scope of the present invention. Other features of the present invention will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0027] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0028] Figure 1 is a flowchart of a resume screening method provided according to an embodiment of the present invention;

[0029] Figure 2 is a flowchart of a resume screening method provided according to an embodiment of the present invention;

[0030] Figure 3 is a schematic structural diagram of a resume screening device provided according to an embodiment of the present invention;

[0031] Figure 4 is a schematic structural diagram of an electronic device for implementing the resume screening method of the embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0032] To enable those skilled in the art to better understand the solution of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0033] It should be noted that the terms "first", "second", etc. in the specification and claims of the present invention and the above-mentioned drawings are used to distinguish similar objects, and do not necessarily need to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so that the embodiments of the present invention described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device comprising a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.

[0034] Figure 1 The figure is a flowchart of a resume screening method provided by an embodiment of the present invention. The embodiments of the present invention are applicable to the situation of screening out resumes corresponding to a position based on the position. This method can be executed by a resume screening device, which can be implemented in the form of hardware and / or software, and the resume screening device can be configured in an electronic device with resume screening functions, such as a server providing resume screening functions.

[0035] See Figure 1 The resume screening method shown includes:

[0036] S101. Obtain position description information.

[0037] Among them, the position description information may be the description text of the recruitment position. The user can directly input the position description information to screen out the resumes that match the position description information.

[0038] S102. For at least one alternative resume, determine the key features corresponding to the position description information and the weights of each key feature according to the data of the alternative resume.

[0039] Among them, the alternative resumes can be resumes that can be screened by the user with the authorization of the resume-providing user. The number of alternative resumes is huge. The data of the alternative resumes can be the content included in the alternative resumes. The key features can refer to the features related to the job description information. Among them, the key features can include a word. For example, a key feature is the Java language. Another example is that a key feature includes more than 3 years of back-end development experience. The weight of the key feature can indicate the degree of relevance between the key feature and the job description information. In fact, there is some noise such as template data and redundant data in the alternative resumes, and the degree of relevance of these data to the job description information is relatively small. Using these data for resume screening will reduce the accuracy of the screening results.

[0040] The data of the alternative resumes is divided into structured data and unstructured data. The structured data includes: education background (such as undergraduate or master, etc.), graduation institution (such as classification of domestic or foreign universities, etc.), working experience years, and skills mastered, etc.; the unstructured data contains content such as personal descriptions of project experience, personal advantages, and career planning. The structured data provides "hard thresholds" (such as education background and years), and the unstructured data supplements "soft capabilities" (such as project experience and technical depth). The combination of the two can more comprehensively evaluate the alternative resumes; by mining the explicit rules of the structured data and combining the implicit features of the unstructured data, the coverage rate and matching accuracy of the key features are improved. The efficient processing of the structured data reduces the model training cost, and after the features of the unstructured data are screened (such as retaining the key features with high information gain), the interpretability of the model can be enhanced (such as "recommend this candidate because he masters Java, Spring Boot and has high-concurrency optimization experience"). The structured data provides efficient and reliable basic features, while the unstructured data expands the feature dimension. The two jointly improve the comprehensiveness and accuracy of the screening model. The stability of the structured data and the flexibility of the unstructured data complement each other, and finally realize more efficient candidate screening that better meets the job requirements.

[0041] In the resume, especially in the unstructured data, there may be noise data. For example, the user's attribute information in the resume is irrelevant to the job description information. Specifically, the user's blood type and constellation are irrelevant to the job description information, and the user's blood type and constellation are redundant data. Another example is that the user's hobbies in the resume are irrelevant to the job description information. By extracting the key features, it is possible to focus on the key features in the resume for screening, improve the accuracy of the screening, thereby improving the anti-interference ability of the screening and having stronger robustness to interference information.

[0042] S103. Determine the key vector of the alternative resume according to each of the key features and the weight of each of the key features.

[0043] Feature extraction can be performed on the key feature and the weight of the key feature to obtain the key vector.

[0044] S104. Determine the output vectors of the primary capsules based on at least one of the primary capsules in the primary capsule layer, the primary coefficients in the trained weight coefficients, and the key vectors of the alternative resumes.

[0045] Among them, the weight coefficient can refer to the connection weight from the unit data of the input data to the input capsule. The capsule network includes a primary capsule layer and a secondary capsule layer. Each capsule layer corresponds to a weight coefficient. The weight coefficient of the primary capsule layer is the primary coefficient, and the weight coefficient of the secondary capsule layer is the secondary coefficient. The weight coefficient can include the primary coefficient of the primary capsule layer and the secondary coefficient of the secondary capsule layer. The key vector and the primary coefficient serve as the input of the primary capsule, and the output vector of the primary capsule serves as the output of the primary capsule.

[0046] Among them, the capsule is the basic unit in the capsule network. Each capsule can output a vector. The length of the vector represents the probability of the existence of the feature, and the direction represents the attribute of the feature. The primary capsule is used to detect local features, and the secondary capsule is used to detect complete objects. In one example, for the scenario of image detection, the primary capsule is used to detect local features such as edges and corners in the image, and the secondary capsule is used to identify complete objects such as cats or dogs. In the resume screening scenario, the primary capsule layer integrates local features (such as single skills), and the secondary capsule layer abstracts global patterns (such as the matching of skill combinations and job requirements). When the primary capsule transmits information to the secondary capsule, the weight coefficient between each primary capsule and the secondary capsule is calculated.

[0047] S105. Determine the output vectors of the secondary capsules based on at least one of the secondary capsules in the secondary capsule layer, the secondary coefficients in the weight coefficients, and the output vectors of the primary capsules.

[0048] The output vector of the primary capsule and the secondary coefficient serve as the input of the secondary capsule, and the output vector of the secondary capsule serves as the output of the secondary capsule.

[0049] In some embodiments, the capsule network includes a three-layer structure, specifically an input layer, a primary capsule layer, and a secondary capsule layer. The input layer receives the key vectors, and the number of its nodes corresponds to the number d of the key vectors. Among them, the number of key vectors is the same as the number of key features. The primary capsule layer consists of multiple primary capsules. Each primary capsule can capture the local feature relationship of the input features. The number of capsules in the primary capsule layer is k1, and the dimension of the output vector of each primary capsule is s1. The secondary capsule layer further integrates and abstracts the information of the primary capsule layer. The number of its capsules is k2, and the dimension of the output vector of each secondary capsule is s2. Each capsule represents a specific feature combination or pattern.

[0050] By using a capsule network, when learning features, more attention can be paid to the essential attributes of the features. When encountering new recruitment requirements or changes in resume formats, the capsule network can quickly adapt to the new situation based on the learned essential features of the features, without the need for a large amount of new data for retraining, having good generalization, and improving the adaptability and flexibility of the resume screening device.

[0051] The weight coefficient is used to dynamically determine how to transfer the output of the lower layer to the higher layer between two layers in the capsule network. The weight coefficient reflects the matching degree between the features represented by the two connected capsules. The higher the matching degree, the larger the weight coefficient; the lower the matching degree, the smaller the weight coefficient. The weight coefficient can well capture complex feature relationships and spatial information, enhance the modeling ability of the relationships between features, enable the capsules to adaptively establish effective connection relationships, better transfer and integrate feature information, and have the advantage of feature learning. For example, when analyzing the skills and project experience mentioned in the alternative resumes, the existing screening methods only simply match based on keywords. However, through the capsule network with weight coefficients, the deep connection between the skill combination and project experience can be understood. For example, understand the logical association between mastering Java and the Spring Boot framework and being responsible for backend development in a large e-commerce project, so as to more accurately evaluate the matching probability of the candidate corresponding to the alternative resume with the position of backend development engineer.

[0052] The weight coefficient can be optimized through multiple iterations to obtain the trained weight coefficient.

[0053] S106. Determine the matching probability between the alternative resume and the position description information according to the output vectors of the respective high-level capsules.

[0054] Among them, the output vector of the high-level capsule can be quantified to obtain the matching probability between the alternative resume and the position description information. The output vector of the high-level capsule indicates the matching probability between the feature corresponding to this high-level capsule and the position description information. The output vectors of all the high-level capsules finally obtained from all the key features extracted from an alternative resume indicate the matching probability between all the content in this alternative resume and the position description information. The higher the matching probability, the more the alternative resume meets the requirements of the position description information; the lower the matching probability, the less the alternative resume meets the requirements of the position description information.

[0055] In some embodiments, according to the output vectors of the respective high-level capsules, determine the output vector of the high-level capsule layer. For example, the output vector v of the high-level capsule layer ∈ R k2×s2 , and the output vector of this high-level capsule layer can be converted into a scalar matching probability through a scoring function f(v). In some embodiments, the matching probability can be converted by means of the modulus length of the vector or weighted summation of the vector. For example, the scoring function is: Among them, αi is the weight coefficient, vi is the output vector of the i-th high-level capsule in the high-level capsule layer, and ||vi|| represents the norm of vi.

[0056] S107. Screen out the target resume corresponding to the job description information from each of the alternative resumes according to the matching probability between each alternative resume and the job description information.

[0057] Among them, the alternative resumes can be sorted from high to low according to the matching probability, and at least one alternative resume with a high matching probability can be selected as the target resume. Or a probability threshold can be set, and the alternative resumes with a probability higher than the probability threshold can be selected and determined as the target resumes.

[0058] In an example, the three-layer structure included in the capsule network is: an input layer, a primary capsule layer, and a high-level capsule layer. The number of nodes in the input layer corresponding to the software engineer position is 25. The primary capsule layer is set with 10 capsule units, and the dimension of the output vector of each unit is 8; the high-level capsule layer is set with 5 capsule units, and the dimension of the output vector of each unit is 16. After the key features are vectorized, they are input into the capsule network. Inside the capsule network, the feature relationship is modeled through a dynamic routing mechanism. For example, for the key feature "master the Java language and have more than 3 years of back-end development experience" in the alternative resume, through the weight coefficient, the information of other relevant features is integrated and transmitted in the primary capsule layer and the high-level capsule layer, so that the capsule network can better capture the matching degree of this resume and the software engineer position at the overall feature level. The matching probability between the alternative resume and the job description information is calculated through a scoring function. If the probability threshold is 0.8 and the calculated matching probability is 0.85 > 0.8, this alternative resume is determined as the target resume.

[0059] The technical solution of the embodiment of the present invention effectively improves the representativeness of the features and reduces redundant information by screening key features in the data of the candidate resumes, and inputs the key features into the primary capsule layer of the trained weight coefficient for processing to obtain the output vector of the primary capsule, and inputs the key features into the advanced capsule layer of the trained weight coefficient for processing to obtain the output vector of the advanced capsule, and finally determines the matching probability between the candidate resume and the job description information based on the output vector of the advanced capsule, and then screens out the target resume corresponding to the job description information from multiple candidate resumes, and can capture the relationship between multiple complex features in the resume. At the same time, the weight coefficient of the trained capsule layer is used to process the vector, and the correlation between the job description information and the underlying features can be accurately characterized, and then the underlying features are integrated to form an accurate representation of multiple types of complex features in the candidate resume. The matching probability between the candidate resume and the job description information is detected based on the accurate representation, and the accuracy of the matching probability can be improved, which solves the problem that it is difficult to mine and utilize the complex feature relationship in the resume in the prior art, resulting in low screening efficiency and poor accuracy, and can effectively mine and accurately represent the complex relationship in the resume, improve the accuracy of resume screening, and reduce labor costs at the same time, and can effectively improve the efficiency of resume screening.

[0060] In an optional embodiment, the weight coefficient is trained in the following manner: determine the output vector of each primary capsule according to the key vector of the candidate resume, at least one primary coefficient in the weight coefficient and at least one primary capsule in the primary capsule layer; calculate the primary matching degree between each key vector and the output vector of each primary capsule; update each primary coefficient according to the primary matching degree; determine the output vector of each advanced capsule according to the output vector of each primary capsule, the advanced coefficient in the weight coefficient and at least one advanced capsule in the advanced capsule layer; calculate the advanced matching degree between the output vector of each primary capsule and the output vector of each advanced capsule; update the advanced coefficient according to the advanced matching degree; and determine the trained weight coefficient when it is determined that the weight coefficient training is satisfied.

[0061] The matching degree may refer to the matching degree between the features represented by the primary capsule and the features represented by the advanced capsule, which may be obtained by dot product calculation. The primary coefficient is positively correlated with the matching degree, and the advanced coefficient is positively correlated with the matching degree.

[0062] In some embodiments, the process of information transmission from the input layer to the primary capsule layer and then to the higher-level capsule layer: First, initialize the weight coefficients. The primary coefficients cij between the input layer nodes and the primary capsules in the primary capsule layer are all initialized to 0. Here, i represents the input layer node number (for example, 25 input layer neurons in the previous example, i = 1, 2, …, 25), and j represents the primary capsule number (for example, 10 primary capsules in the previous example, j = 1, 2, …, 10).

[0063] The first iteration calculation can specifically include: calculating the input sj of each primary capsule in the primary capsule layer, based on sj = ∑ i cijui, where ui is the output vector of the input layer node i. Since the primary coefficient cij = 0, the initial value of sj is small. Apply a non-linear function (such as the squash function) to sj to obtain the output vector vj of the primary capsule j in the primary capsule layer. Calculate the input sk of each higher-level capsule in the higher-level capsule layer, based on sk = ∑ i dijvj, where dij is the higher-level coefficient. Apply a non-linear function (such as the squash function) to sk to obtain the output vector vk of the higher-level capsule k in the higher-level capsule layer. When the primary capsule j in the primary capsule layer transmits information to the higher-level capsule k in the higher-level capsule layer, calculate the matching degree ajk between the primary capsule j and each higher-level capsule k in the higher-level capsule layer (for example, 5 higher-level capsules in the previous example, k = 1, 2, …, 5). The matching degree reflects the matching degree of the features represented by the two capsules and is usually obtained through the dot product operation. For example, if a higher-level capsule in the higher-level capsule layer represents a "senior software engineer", when calculating the matching degree between the primary capsules corresponding to "Java language" and "more than 3 years of back-end development experience" and the higher-level capsule corresponding to "senior software engineer", the primary capsule and the higher-level capsule are more relevant and the matching degree is larger. Update the weight coefficients cij and dij according to the matching degree, using the softmax function to constrain the sum of the primary coefficients between each input layer neuron and each primary capsule to be 1, and to constrain the sum of the higher-level coefficients between each primary capsule and each higher-level capsule to be 1. Such training can increase the higher-level coefficients between the primary capsules corresponding to "Java language" and "more than 3 years of back-end development experience" with higher relevance and the higher-level capsule corresponding to "senior software engineer".

[0064] Subsequent Iterations: Repeat the above steps multiple times, for example, 3 - 5 times. As the iteration progresses, the weight coefficients are continuously adjusted, and the information carried by the underlying capsules will be more accurately transmitted to the corresponding higher-level capsules. For example, in the second iteration, because the weight coefficients between the two junior capsules corresponding to "Java language" and "more than 3 years of back-end development experience" and the senior software engineer corresponding higher-level capsule increase, the information of these two junior capsules is more transmitted to this higher-level capsule, and the higher-level capsule can more accurately capture the relationship between these two features and "senior software engineer". When the number of iterations is greater than or equal to the iteration threshold (such as the iteration threshold is 3 - 5), it can be determined that the iteration ends. At this time, the weight coefficients are stable, and the trained weight coefficients are obtained. The capsule network completes the integration and transmission of feature information, enabling the information of the underlying capsules to be accurately transmitted to the corresponding higher-level capsules, thereby effectively capturing complex feature relationships and spatial information.

[0065] It can be seen that by dynamically adjusting the information transmission path according to the matching degree of the features represented by the capsules, an effective connection relationship can be adaptively established, and the feature information can be better transmitted and integrated to make the information transmission more flexible and accurate; by iteratively updating the coupling coefficient multiple times, the hierarchical relationship and complex associations between features can be better captured, and the deep connections between the features represented by the capsules at different layers can be discovered, so as to more accurately evaluate the matching degree between the resume and the position, and have good feature relationship capture ability; by dynamically adjusting the weight coefficients, the influence of irrelevant information is weakened, ensuring the accuracy of screening. Especially when dealing with unstructured data in the resume, it can focus on key features and has stronger robustness to interference, with high anti-interference ability.

[0066] Figure 2 The flowchart of a resume screening method provided by an embodiment of the present invention. On the basis of the above embodiment, the present invention embodiment specifies "determining the key features corresponding to the position description information and the weights of each key feature according to the data of the alternative resume" as: extracting at least one resume key information from the data of the alternative resume; combining each resume key information to obtain at least one feature combination; calculating the confidence and support of each feature combination; determining at least one key combination from each feature combination according to the confidence and support of each feature combination, and determining the resume key information in each key combination as the key feature; calculating the matching degree between each key feature in each key combination and the position description information to obtain the weights of each key feature in each key combination. It should be noted that for the parts not detailed in the embodiments of the present invention, reference can be made to the descriptions of other embodiments.

[0067] See Figure 2 The resume screening method shown includes:

[0068] S201. Obtain job description information.

[0069] S202. Extract at least one resume key information from the data of the alternative resumes.

[0070] A large amount of resume data can be collected. A large amount of resumes and corresponding recruitment job description information are collected from multiple public or authorized channels. The resumes cover structured data and unstructured data. Among them, the structured data may include educational background information and work experience information. Among them, the educational background information may include: educational level, graduated school, major studied, etc.; the work experience information may include work years, companies worked for, job titles and responsibilities, etc. The unstructured data may include: personal statement information and skill description information. Among them, the personal description information may include: written descriptions of one's own career development, advantages and characteristics, etc.; the skill description information may include: various professional skills mastered, etc. These data are integrated to form the original dataset D, which can be expressed as: D = {(ri, ji)|i = 1, 2,..., n}, where ri represents all data information corresponding to the i-th resume, ji represents the corresponding job description information, and n is the total number of groups of resumes and job descriptions collected. Among them, the job description information implies the core requirements of the recruitment position (such as skills, experience, and educational background, etc.), and the content included in the resumes that meet these requirements is used for matching judgment during subsequent model training. For example, if the job description requires "3 years of Java development experience", then the resumes that meet this requirement will be marked as "suitable", otherwise as "unsuitable".

[0071] Among them, the resume key information may refer to the information extracted from the alternative resumes. The data of the alternative resumes can be subjected to keyword extraction to obtain at least one keyword as the resume key information. Among them, the resume key information can be extracted from structured data or unstructured data. In some embodiments, the attribute value corresponding to the field in the structured data can be directly determined as the resume key information. In some embodiments, the sentences in the unstructured data can be segmented, and keywords can be extracted from the segmentation results and determined as the resume key information. In some embodiments, semantic understanding can be performed on the unstructured data to obtain the summary information of the unstructured data and determine it as the resume key information.

[0072] S203. Combine each of the resume key information to obtain at least one feature combination.

[0073] Among them, the feature combination may include at least one resume key information. The resume key information can be randomly combined to obtain at least one feature combination. Among them, the feature combination includes at least one resume key information and at least one content of the job description information.

[0074] In some embodiments, structured data and unstructured data may be distinguished, and data of different structural types may be processed accordingly to obtain a feature combination.

[0075] S204: Calculate the confidence and support of each feature combination.

[0076] Confidence is used to determine the reliability of the feature combination relative to the job description information. Confidence can refer to the probability that the consequent will also appear when the antecedent appears. In the embodiment of the present invention, confidence reflects the reliability or credibility from the antecedent to the consequent. Confidence is used to measure whether a feature combination is reliable. High confidence means that when the antecedent occurs, the consequent has a high probability of occurring; low confidence means that when the antecedent occurs, the consequent has a low probability of occurring. Support is used to determine the frequency of occurrence of the feature combination in resumes that match the job description information. In the embodiment of the present invention, support reflects the prevalence of the feature combination in resumes that match the job description information. Support is used to measure whether a feature combination is frequent. High support means that the item set or rule is very common in the data set and may have practical application value; low support means that the item set or rule is relatively rare and may lack practical significance.

[0077] S205. Determine at least one key combination in each of the feature combinations according to the confidence and support of each of the feature combinations, and determine the resume key information in each of the key combinations as the key feature.

[0078] Typically, a feature combination with high confidence and high support is selected as a key combination. In some embodiments, the confidence and support of each feature combination may be weighted and summed to obtain a weighted sum result, and based on the weighted sum result of each feature combination, the first n feature combinations that are higher than the corresponding threshold or ranked from high to low are selected as key combinations. In some embodiments, the confidence threshold and support threshold may be configured separately, and a feature combination with a confidence higher than the confidence threshold and a support higher than the support threshold may be determined as a key combination.

[0079] The at least one resume key information included in the key combination is determined as the same number of key features. The resume key information and the key features in the key combination correspond one to one.

[0080] In an example, the set of all resume key information is F = {f1, f2, ..., fm}, and the feature combination that meets the minimum support (set as minsup) and minimum confidence (set as minconf) requirements is found through frequent item set mining. The definition and mining process of frequent item sets are as follows:

[0081] The support calculation formula for frequent itemsets is:

[0082] Among them, X is a set of feature items (i.e., feature combination), and Support(X) represents its support degree in the original data set D. It represents the number of item sets X contained in the original data set, and |D| is the total number of the original data set. Only the item sets with a support degree greater than or equal to minsup are determined as frequent item sets.

[0083] For association rules (indicating that when the feature combination X appears, the feature combination Y is also likely to appear), its confidence calculation formula is:

[0084] Usually, the key contents in the job description information, such as programming languages, frameworks, working years, etc., are transformed into the consequent of the association rule. For example, the consequent "3 years of experience" of the association rule {Java, Spring Boot} {3 years of experience} directly comes from the job description information. In this way, when screening frequent item sets, the requirements in the job description information are used as the screening criteria. For example, if a certain job requires "Python and Django", by using the job description information as the consequent of the association rule, it is possible to give priority to focusing on the combinations containing these features, improving the job relevance.

[0085] By mining the association rules that meet the minimum confidence requirements, identify which feature combinations (such as specific professional skills and a certain number of years of relevant work experience, etc.) have a high influence on job matching, and form an associated feature set C. Each item in the associated feature set can be determined as a key combination.

[0086] In the scenario of screening job resumes for software development positions, through the associated screening of feature combinations, it is possible to discover the key combination of "Java language, Spring Boot framework and 3 years of work experience". The key combination can provide a more comprehensive reference for resume screening, considering resumes from multiple dimensions, no longer limited to a single condition, more accurately grasping the matching degree between the job requirements and the user's abilities in the resume, being able to discover potential feature relationships in job recruitment, so as to further improve the accuracy of resume screening; and screening resumes based on the key combination can quickly and accurately locate suitable resumes, avoiding the situation that the existing screening criteria are vague, which easily leads to missing suitable resumes or screening out unmatched resumes. The key combination mined by using association rules, such as "Python language, Django framework and 2 years of work experience", can be directly used for screening, reducing the workload of manual screening, improving the efficiency and accuracy of resume screening, and reducing the cost of resume screening.

[0087] In one example, for the structured data in the resume (such as programming languages, years of work experience, and educational background), there are clear classification tags and formats, which can be directly used as discrete features to participate in frequent itemset mining. For example, the key combination of "Java + Spring Boot + 3 years of experience" can be quickly identified without complex preprocessing, enabling efficient mining of frequent itemsets; the features of structured data (such as "3 years of work experience") are usually strongly related to the job requirements, and the confidence of their combination rules (such as "resumes that master Java and Spring Boot have a 70% probability of having 3 years of experience") is more reliable, which can improve the confidence of key combinations; and the feature dimension of structured data is low and the standardization degree is high, reducing the number of candidate item sets, lowering the computational complexity, and enhancing the efficiency of key feature extraction.

[0088] For the unstructured data in the resume, unstructured data (such as "high-concurrency system optimization" in the personal statement) can extract implicit ability features (such as "performance optimization experience") through natural language processing (NLP) technology. These features may form new association patterns with structured data (such as "high-concurrency optimization + master's degree"), which can achieve the supplementation of implicit features, enrich and accurately extract the resume content; the text information in the skill description (such as "familiar with microservice architecture") can supplement the technical details not covered by structured data, helping to discover more detailed rules (such as "microservice + Kubernetes cloud native development capabilities"), achieving the expansion of the feature coverage range.

[0089] S206. Calculate the matching degree between each key feature in each of the key combinations and the job description information to obtain the weight of each key feature in each of the key combinations.

[0090] The weight of a key feature is used to determine the matching degree between the key feature and the job information description. The similarity between the key feature and the job description information can be calculated to determine the matching degree, that is, the weight of the key feature. In some embodiments, algorithms such as term frequency–inverse document frequency (TF-IDF) and best match 25 (BM25) can be used to calculate the weight. The weight of the key feature can provide an initial priority for the weight coefficient of the subsequent capsule network. In one example, the content of the filtered key features and their weights is: {"Java": 0.8, "Spring Boot": 0.7, "e-commerce project": 0.9}.

[0091] S207. Determine the key vector of the alternative resume according to each of the key features and the weight of each of the key features.

[0092] A key feature and the weight corresponding to the key feature can be encoded to obtain a key vector, which represents the key feature. Feature screening is used to solve the problem of "which features to screen", and the capsule network is used to solve the problem of "how to understand the feature relationship". Feature screening first filters out noise, and the capsule network further weakens redundant information through dynamic weight coefficients to improve the anti-interference ability. Feature screening reduces the input dimension, lowers the computational complexity of the capsule network, and at the same time retains the deep association of key features, improving the resume screening efficiency.

[0093] S208. Determine the output vector of each primary capsule according to the key vector of the alternative resume, the primary coefficient in the trained weight coefficients, and at least one primary capsule in the primary capsule layer.

[0094] S209. Determine the output vector of each secondary capsule according to the output vector of each primary capsule, the secondary coefficient in the weight coefficients, and at least one secondary capsule in the secondary capsule layer.

[0095] S210. Determine the matching probability between the alternative resume and the job description information according to the output vector of each secondary capsule.

[0096] S211. Screen out the target resume corresponding to the job description information from each alternative resume according to the matching probability between each alternative resume and the job description information.

[0097] The technical solution of the embodiment of the present invention can accurately screen out the features related to the job description in the resume, improve the confidence of the key features, and at the same time, can reduce the amount of data processing required for resume screening, lower the computational complexity, and improve the screening efficiency by extracting the key resume information from the data of the alternative resume, combining the key resume information to obtain a feature combination, calculating the confidence and support of the feature combination, screening out the key combination in the feature combination, and then determining the resume key information in the key combination as the key feature.

[0098] In an optional embodiment, after determining at least one key combination in each feature combination according to the confidence and support of each feature combination, it further includes: calculating the information gain of each key combination; filtering each key combination according to the information gain of each key combination.

[0099] Among them, filtering the key combinations based on the information gain of the key combinations actually evaluates the importance of the key combination by calculating the degree of reduction in the uncertainty of all key combinations by the key combination. For all key combinations, calculate the information gain of each key combination, sort them from high to low according to the information gain, retain the top k key combinations, and delete the remaining key combinations. Or configure a gain threshold, retain multiple key combinations that are higher than or equal to the gain threshold, and delete the key combinations that are lower than the gain threshold.

[0100] The information gain actually screens out key combinations with high discrimination by comparing the distribution differences of key combinations in "matching job descriptions" and "non-matching job descriptions". For example, if the job description information emphasizes "optimization of high-concurrency systems", key features in the relevant personal statements (such as "response time reduced by 50%") will be given a higher information gain. The core requirements in the job description information (such as "Java + Spring Boot + 3 years of experience") perform better in the information gain calculation, so they are preferentially retained in the set S of the final key combinations to achieve feature priority sorting.

[0101] In an example, for the structured data in the resume, structured data (such as "Bachelor's degree from Tsinghua University") usually has a high information gain and can directly distinguish whether a candidate meets the hard requirements (such as educational threshold); the characteristics of structured data (such as "working years") are less affected by text noise, and the information gain calculation results are more stable, reducing the risk of model overfitting.

[0102] For the unstructured data in the resume, the text information in the unstructured data (such as "led the development of the core system") can be transformed into features such as "leadership" or "complex project experience". These soft capabilities can predict the job matching degree of the resume better than simple educational background or years of experience; by analyzing the keywords in the skill description (such as "machine learning", "blockchain"), emerging technology requirements can be quickly identified, making feature selection more flexible.

[0103] For example, for the software engineer position, the feature combination of "master the Java language and have more than 3 years of back-end development experience" has a high impact on job matching. For example, the support degree of this feature combination reaches 0.12 and the confidence level is 0.7. This feature combination is a key combination for the software engineer position; for the product manager position, the feature combination of "have experience in launching an Internet product from scratch and be familiar with user research methods" has a strong correlation. This feature combination is a key combination for the product manager position. Based on the key combinations, further screening is carried out through the information gain method to finally determine the key features corresponding to each position, reducing the number of input data of the capsule network from hundreds of features to about 20 - 30 key features for each position, achieving a significant reduction in the number of input data.

[0104] It can be seen that, based on the information gain of each key combination, further screening of the key combinations can reduce the data dimension, avoid data redundancy, and improve the operation efficiency of resume screening.

[0105] In an example, a large Internet company is recruiting software development engineers and screening resumes for the position of software development engineer.

[0106] Through the company's recruitment platform, major recruitment websites, and internal referral channels, 5,000 alternative resumes for software development positions and corresponding job description information were collected.

[0107] For example, the structured data included in the alternative resumes are: Educational background: such as graduating from XX Computer Science and Technology major at the undergraduate level and graduating from YY Software Engineering major at the master's level, etc. Work experience: For example, having 3 years of work experience, having worked in a certain company as a senior software engineer, responsible for the development and maintenance of the core business system, etc. The unstructured data included in the alternative resumes are: Personal statement: "In the past project, I led the optimization of the high-concurrency system, shortened the system response time by 50%, and improved the user experience." Skill description: "Proficient in programming languages such as Java, Python, and C++; familiar with development frameworks such as Spring Boot and Django."

[0108] After integration, the original dataset D is formed, D = {(ri, ji) | i = 1, 2,..., 5000}, where ri represents the i-th resume data and ji represents the corresponding job description information for software development.

[0109] Analyze the resume data characteristics, and set the minimum support minsup = 0.1 and the minimum confidence minconf = 0.6.

[0110] The set F of feature combinations includes the key resume information: programming languages (such as Java, Python, etc.), development frameworks (such as Spring Boot, Django, etc.), years of work experience, and education level, etc.

[0111] For example, through frequent itemset mining, it is found that the support degree of the feature combination {Java language, Spring Boot framework, 3 years of work experience} is: Support = 0.15 ≥ 0.1. So this feature combination is a frequent itemset.

[0112] For the association rule {Java language, Spring Boot framework} {3 years of work experience}, its confidence is: Confidence = 0.7 ≥ 0.6.

[0113] The set C of key combinations obtained by screening based on support and confidence includes: {Java language, Spring Boot framework, 3 years of work experience} and {Python language, Django framework, 2 years of work experience}, etc.

[0114] The set C of key combinations is screened using information gain.

[0115] For example, for the feature combinations {Java language, Spring Boot framework, 3 years of work experience} and {Python language, Django framework, 2 years of work experience}, calculate the information gain in distinguishing suitable and unsuitable. After calculation, it is found that the information gain of {Java language, Spring Boot framework, 3 years of work experience} is higher. The final set S of screened key combinations includes this key combination, etc. The finally retained key combinations will be used as the input features of the subsequent capsule network, improving the training and running efficiency of the model and more accurately screening out the target resumes suitable for the position of software development engineer.

[0116] In this example, the job description information provides an objective basis for the effectiveness of the key features in the resume, ensuring that the mined key combinations are directly related to the job requirements and providing clear screening criteria. By filtering out irrelevant features through the job description information and only focusing on the skill combinations related to the job requirements, redundant calculations can be reduced and the screening efficiency of key features can be improved. The finally screened key combinations are highly consistent with the job description, enhancing the result interpretability and facilitating user understanding and application.

[0117] In an optional embodiment, extracting at least one resume key information from the data of the alternative resumes includes: extracting at least one resume alternative information from the data of the alternative resumes; filtering each of the resume alternative information according to the job description information to obtain at least one resume key information.

[0118] Among them, the resume alternative information may refer to the valid information extracted from the resume data. The parameter values of the fields in the structured data can be directly determined as the resume alternative information; converting unstructured text (such as project descriptions) into structured text. For example, extracting skill information and project result quantification information from unstructured text to generate structured text, which is convenient for capsule network processing. Filtering the resume alternative information to obtain at least one resume key information may be screening out the resume key information strongly related to the job from the resume alternative information. In some embodiments, the features strongly related to the job requirements (such as skills, project experience, or educational background, etc.) extracted from the resume are filtered out of redundant or irrelevant information (such as irrelevant hobbies).

[0119] It can be seen that by filtering the resume alternative information in the data of alternative resumes, at least one key resume information can be obtained, which can reduce redundant information and improve the efficiency of resume screening.

[0120] In an optional embodiment, the resume screening method further includes: obtaining historical recruitment data of at least one sample job information, where the historical recruitment data of the sample job information includes: at least one sample resume and the true value screening result of each sample resume for the sample job information; for each sample resume, determining the key features corresponding to each sample job information and the weights of each key feature according to the data of the sample resume; according to each key feature and the weights of each key feature, determining the key vector of the sample resume; according to the key vectors of each sample job information, the primary coefficients in the trained weight coefficients and at least one primary capsule in the primary capsule layer, determining the output vector of each primary capsule; according to the output vectors of each primary capsule, the advanced coefficients in the weight coefficients and at least one advanced capsule in the advanced capsule layer, determining the output vector of each advanced capsule; according to the output vectors of each advanced capsule, determining the matching probability between the sample resume and each sample job information; according to the matching probability between each sample resume and each sample job information, determining the predicted screening result corresponding to each sample job information; according to the difference between the predicted screening result of each sample job information and the true value screening result of each sample job information, adjusting the structure of the capsule layer, where the capsule layer includes a primary capsule layer and an advanced capsule layer.

[0121] Among them, the structure of the capsule network can be determined through training. The historical recruitment data can be sample data for adjusting the capsule network. The sample job information can be the job description information of historical recruitment. The sample resume can be all alternative resumes obtained for the sample job information. The true value screening result can be the target resume screened from each sample resume for the sample job information. The historical recruitment data of one sample job information includes at least one sample resume and the true value screening result of this sample job information. The predicted screening result can refer to the result of the sample resume screened by using the resume screening method provided in the embodiments of the present invention from all sample resumes obtained for the sample job information. With the goal of narrowing the difference between the predicted screening result of the sample job information and the true value screening result of this sample job information, adjust the structure of the capsule layer.

[0122] In some embodiments, the structure of the capsule layer can include: the number of capsule layers and / or the parameters of capsule units, etc. Adjusting the structure of the capsule layer can include increasing or decreasing the number of capsule layers, and increasing or decreasing the dimension of the output vector in the parameters of capsule units.

[0123] Among them, the differences can be calculated using accuracy, recall, and F1-score.

[0124] The accuracy can be calculated based on the accuracy formula:

[0125]

[0126] Among them, TP (True Positive) represents the true positive example, that is, the number of resumes whose true value screening result is a match and the prediction is a match; TN (True Negative) represents the true negative example, that is, the number of resumes whose true value screening result is a non-match and the prediction is a non-match; FP (False Positive) represents the false positive example, that is, the number of resumes whose true value screening result is a non-match but the prediction is a match; FN (False Negative) represents the false negative example, that is, the number of resumes whose true value screening result is a match but the prediction is a non-match.

[0127] The recall formula:

[0128]

[0129] The F1-score formula (which is the harmonic mean of accuracy and recall):

[0130]

[0131] It can be seen that by optimizing and adjusting the structure of the capsule layer through historical data, the capsule layer can maintain a good screening effect in different scenarios, enhancing the practicability and adaptability of the resume screening method.

[0132] In an optional embodiment, determining at least one key combination among the feature combinations according to the confidence and support of each feature combination includes: determining at least one key combination among the feature combinations according to a pre-determined confidence threshold, support threshold, the confidence and support of each feature combination; the method further includes: updating the confidence threshold and the support threshold according to the difference between the predicted screening result of each sample job information and the true screening result of each sample job information.

[0133] Among them, a feature combination that is greater than or equal to the confidence threshold and greater than or equal to the support threshold can be determined as a key combination. A feature combination that is less than the confidence threshold or less than the support threshold is not determined as a key combination. With the goal of narrowing the difference between the predicted screening result of the sample job information and the true screening result of the sample job information, the confidence threshold and the support threshold are adjusted, for example, increasing or decreasing the values of the confidence threshold and the support threshold.

[0134] It can be seen that by optimizing and adjusting the confidence threshold and support threshold through historical data, the confidence threshold and support threshold can be ensured to maintain good screening effects in different scenarios, enhancing the practicability and adaptability of the resume screening method.

[0135] In one example, data preprocessing, feature screening (confidence support screening and information gain screening), and capsule networks can be integrated into an automated resume screening system. The resume screening system provides a user interface to facilitate users to input job description information, etc. The resume screening system automatically completes the full-process automated operation from feature screening and prediction of raw data to the output of the final screening results, improving the overall efficiency of the resume screening process. Through experiments, using historical recruitment data (such as 500 resumes successfully recruited in the past six months and corresponding job information as the test set) to evaluate the performance of the resume screening system, it is found that the accuracy rate is 80%, the recall rate is 75%, and the F1 score is 0.77. It is analyzed that the accuracy of resume screening for some emerging technology positions needs to be improved. Therefore, the capsule network is optimized by adding a layer of capsule layer, adjusting the output vector dimensions of some capsule units, and appropriately adjusting the minimum support parameter for association rule mining. After optimization and re-evaluation, the accuracy rate is increased to 85%, the recall rate is increased to 82%, and the F1 score reaches 0.83, significantly improving the resume screening effect, helping users to more efficiently and accurately screen out suitable resumes, and enhancing the efficiency and quality of resume screening.

[0136] The embodiments of the present invention are applicable to the resume screening work in the recruitment process of various enterprises and institutions, especially applicable to scenarios with numerous recruitment positions, a large number of received resumes, and high requirements for the matching degree between candidates and positions, such as large Internet enterprises recruiting different technical positions and financial institutions recruiting various professional talents, etc. It can help recruiters quickly and accurately screen out suitable candidates from a large number of resumes. The embodiments of the present invention, through innovative data feature processing methods, mine important feature combinations in resume data, and then perform feature selection based on this, fully excavating the hidden feature correlations in the data, effectively reducing the data dimension while retaining valuable information for position matching, providing high-quality input features for the subsequent model; using the characteristics of capsule networks that are good at capturing complex feature relationships and spatial information to deeply process the screened features, and better modeling the internal connections between features through its unique dynamic adjustment of weight coefficients, having advantages over traditional neural networks when processing data containing multiple complex features such as resumes, and being able to more accurately predict the matching degree between resumes and positions; integrating the entire resume screening process into an automated system to achieve full-process automated operation and improve recruitment efficiency; at the same time, through multi-index performance evaluation and targeted optimization strategies, ensuring that resume screening can maintain good screening effects in different scenarios, enhancing the practicability and adaptability of the method.

[0137] Figure 3 This is a schematic structural diagram of a resume screening device provided by an embodiment of the present invention. The embodiment of the present invention is applicable to the situation of screening out resumes corresponding to a position based on the position. The device can execute a resume screening method. The device can be implemented in the form of hardware and / or software, and the device can be configured in an electronic device carrying the resume screening function.

[0138] See Figure 3 The resume screening device shown in

[0139] A description information acquisition module 301, configured to acquire position description information;

[0140] A feature selection module 302, configured to determine, for at least one alternative resume, key features corresponding to the position description information and weights of each of the key features according to data of the alternative resume;

[0141] A capsule network 303, configured to determine a key vector of the alternative resume according to each of the key features and weights of each of the key features;

[0142] The capsule network 303 is configured to determine an output vector of each of the primary capsules according to the key vector of the alternative resume, a primary coefficient in the trained weight coefficients, and at least one primary capsule in the primary capsule layer;

[0143] The capsule network 303 is configured to determine an output vector of each of the senior capsules according to the output vectors of each of the primary capsules, a senior coefficient in the weight coefficients, and at least one senior capsule in the senior capsule layer;

[0144] The capsule network 303 is configured to determine a matching probability between the alternative resume and the position description information according to the output vectors of each of the senior capsules;

[0145] A resume screening module 304, configured to screen out a target resume corresponding to the position description information from each of the alternative resumes according to the matching probability between each of the alternative resumes and the position description information.

[0146] The technical solution of the embodiment of the present invention effectively improves the representativeness of the features and reduces redundant information by screening key features in the data of the candidate resumes, and inputs the key features into the primary capsule layer of the trained weight coefficient for processing to obtain the output vector of the primary capsule, and inputs the key features into the advanced capsule layer of the trained weight coefficient for processing to obtain the output vector of the advanced capsule, and finally determines the matching probability between the candidate resume and the job description information based on the output vector of the advanced capsule, and then screens out the target resume corresponding to the job description information from multiple candidate resumes, and can capture the relationship between multiple complex features in the resume. At the same time, the weight coefficient of the trained capsule layer is used to process the vector, and the correlation between the job description information and the underlying features can be accurately characterized, and then the underlying features are integrated to form an accurate representation of multiple types of complex features in the candidate resume. The matching probability between the candidate resume and the job description information is detected based on the accurate representation, and the accuracy of the matching probability can be improved, which solves the problem that it is difficult to mine and utilize the complex feature relationship in the resume in the prior art, resulting in low screening efficiency and poor accuracy, and can effectively mine and accurately represent the complex relationship in the resume, improve the accuracy of resume screening, and reduce labor costs at the same time, and can effectively improve the efficiency of resume screening.

[0147] Optionally, the resume screening device further includes: a weight coefficient training module, which is used to:

[0148] Determine an output vector of each primary capsule according to a key vector of the candidate resume, at least one primary coefficient in the weight coefficients, and at least one primary capsule in the primary capsule layer;

[0149] Determine the output vector of each of the high-level capsules according to the output vector of each of the primary capsules, the high-level coefficients in the weight coefficients, and at least one high-level capsule in the high-level capsule layer;

[0150] Calculating the matching degree between the output vector of each primary capsule and the output vector of each high-level capsule;

[0151] updating the advanced coefficient and each of the primary coefficients according to the matching degree;

[0152] When it is determined that the weight coefficient training is completed, the trained weight coefficient is determined.

[0153] Optionally, the feature selection module 302 is specifically used for:

[0154] Extracting at least one resume key information from the data of the candidate resume;

[0155] Combining each of the resume key information to obtain at least one feature combination;

[0156] Calculate the confidence and support of each of the feature combinations;

[0157] Based on the confidence and support of each of the feature combinations, determine at least one key combination among the feature combinations, and determine the resume key information in each of the key combinations as key features;

[0158] Calculate the matching degree between each key feature in each of the key combinations and the job description information to obtain the weights of each key feature in each of the key combinations.

[0159] Optionally, the resume screening device further includes: an information gain screening module, configured to:

[0160] After determining at least one key combination among the feature combinations based on the confidence and support of each of the feature combinations, calculate the information gain of each key combination;

[0161] Filter each of the key combinations according to the information gain of each key combination.

[0162] Optionally, the feature selection module 302 is specifically configured to:

[0163] Extract at least one resume alternative information from the data of the alternative resumes;

[0164] Filter each of the resume alternative information according to the job description information to obtain at least one resume key information.

[0165] Optionally, the resume screening device further includes: a capsule structure adjustment module, configured to:

[0166] Obtain the historical recruitment data of at least one sample job information, where the historical recruitment data of the sample job information includes: at least one sample resume and the true value screening results of each of the sample resumes for the sample job information;

[0167] For each of the sample resumes, determine the key features corresponding to each of the sample job information and the weights of each of the key features according to the data of the sample resume;

[0168] Determine the key vector of the sample resume according to each of the key features and the weights of each of the key features;

[0169] Determine the output vector of each of the primary capsules according to the key vectors of each of the sample job information, the primary coefficients in the trained weight coefficients, and at least one primary capsule in the primary capsule layer;

[0170] Determine the output vector of each of the secondary capsules according to the output vectors of each of the primary capsules, the secondary coefficients in the weight coefficients, and at least one secondary capsule in the secondary capsule layer;

[0171] Determine the matching probability between the sample resume and each of the sample job information according to the output vectors of the respective advanced capsules;

[0172] Determine the predicted screening results corresponding to each of the sample job information according to the matching probabilities between the respective sample resumes and the respective sample job information;

[0173] Adjust the structure of the capsule layer according to the differences between the predicted screening results of each of the sample job information and the true screening results of each of the sample job information, where the capsule layer includes a primary capsule layer and an advanced capsule layer.

[0174] Optionally, the feature selection module 302 is specifically configured to:

[0175] Determine at least one key combination among the respective feature combinations according to a pre-determined confidence threshold, support threshold, the confidence and support of each of the feature combinations;

[0176] The resume screening device further includes: a threshold adjustment module, configured to:

[0177] Update the confidence threshold and the support threshold according to the differences between the predicted screening results of each of the sample job information and the true screening results of each of the sample job information.

[0178] The resume screening device provided by the embodiments of the present invention can execute the resume screening method provided by any embodiment of the present invention, and has the corresponding functional modules and beneficial effects for executing the method.

[0179] In the technical solution of the embodiments of the present invention, the acquisition, storage, and application of the resume data involved all comply with the provisions of relevant laws and regulations and do not violate public order and good customs.

[0180] Figure 4 FIG. shows a schematic structural diagram of an electronic device 400 that can be used to implement the embodiments of the present invention. The electronic device is intended to represent various forms of digital computers, such as, a laptop computer, a desktop computer, a workbench, a personal digital assistant, a server, a blade server, a mainframe computer, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as, a personal digital processor, a cellular phone, a smart phone, a wearable device (such as a helmet, glasses, a watch, etc.) and other similar computing devices. The components shown herein, their connections and relationships, and their functions are only examples and are not intended to limit the implementation of the present invention described herein and / or claimed.

[0181] As Figure 4As shown, the electronic device 400 includes at least one processor 401 and a memory communicatively connected to the at least one processor 401, such as read-only memory (ROM) 402, random access memory (RAM) 403, etc. The memory stores a computer program executable by the at least one processor. The processor 401 can perform various appropriate actions and processes according to the computer program stored in the read-only memory (ROM) 402 or the computer program loaded from the storage unit 408 into the random access memory (RAM) 403. In the RAM 403, various programs and data required for the operation of the electronic device 400 can also be stored. The processor 401, the ROM 402, and the RAM 403 are connected to each other via a bus 404. The input / output (I / O) interface 405 is also connected to the bus 404.

[0182] Multiple components in the electronic device 400 are connected to the I / O interface 405, including: an input unit 406, such as a keyboard, a mouse, etc.; an output unit 407, such as various types of displays, speakers, etc.; a storage unit 408, such as a magnetic disk, an optical disc, etc.; and a communication unit 409, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 409 allows the electronic device 400 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.

[0183] The processor 401 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the processor 401 include but are not limited to a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, a digital signal processor (DSP), and any appropriate processor, controller, microcontroller, etc. The processor 401 executes the various methods and processes described above, such as the resume screening method.

[0184] In some embodiments, the resume screening method can be implemented as a computer program tangibly embodied in a computer-readable storage medium, such as the storage unit 408. In some embodiments, part or all of the computer program can be loaded and / or installed onto the electronic device 400 via the ROM 402 and / or the communication unit 409. When the computer program is loaded into the RAM 403 and executed by the processor 401, one or more steps of the resume screening method described above can be executed. Alternatively, in other embodiments, the processor 401 can be configured to execute the resume screening method by any other appropriate means (e.g., by means of firmware).

[0185] The various embodiments of the systems and techniques described above in this specification can be implemented in digital electronic circuitry, integrated circuit systems, field programmable gate arrays (FPGA), application specific integrated circuits (ASIC), application specific standard products (ASSP), systems-on-chip (SOC), complex programmable logic devices (CPLD), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include: being implemented in one or more computer programs that are executable and / or interpretable on a programmable system including at least one programmable processor, which may be a special-purpose or general-purpose programmable processor that receives data and instructions from, and transmits data and instructions to, a storage system, at least one input device, and at least one output device.

[0186] The computer programs for implementing the methods of the present invention can be written in any combination of one or more programming languages. These computer programs can be provided to a processor of a general purpose computer, special purpose computer, or other programmable data processing apparatus, such that the computer programs, when executed by the processor, cause the functions / operations specified in the flowchart and / or block diagram to be implemented. The computer programs can be executed entirely on the machine, partly on the machine, as a stand-alone software package partly on the machine and partly on a remote machine or entirely on the remote machine or server.

[0187] In the context of the present invention, a computer-readable storage medium can be a tangible medium that can contain or store a computer program for use by or in connection with an instruction execution system, apparatus, or device. The computer-readable storage medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. Alternatively, the computer-readable storage medium can be a machine-readable signal medium. More specific examples of the machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0188] To provide interaction with a user, the systems and techniques described herein can be implemented on an electronic device having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and a pointing device (e.g., a mouse or a trackball) through which the user can provide input to the electronic device. Other kinds of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).

[0189] The systems and techniques described herein can be implemented in a computing system including backend components (e.g., as a data server), or a computing system including middleware components (e.g., an application server), or a computing system including frontend components (e.g., a user computer having a graphical user interface or a web browser through which the user can interact with an implementation of the systems and techniques described herein), or a computing system including any combination of such backend components, middleware components, or frontend components. The components of the system can be interconnected to each other by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: local area network (LAN), wide area network (WAN), blockchain network, and the Internet.

[0190] The computing system can include a client and a server. The client and the server are generally remote from each other and typically interact through a communication network. The client-server relationship is created by computer programs running on the respective computers and having a client-server relationship with each other. The server can be a cloud server, also known as a cloud computing server or a cloud host, which is a host product in the cloud computing service system, solving the defects of difficult management and weak business scalability existing in traditional physical hosts and VPS (Virtual Private Server) services.

[0191] It should be understood that various forms of the processes shown above can be used, with steps reordered, added, or deleted. For example, the steps recited in the present invention can be executed in parallel, sequentially, or in a different order, as long as the desired results of the technical solution of the present invention can be achieved, and no limitation is imposed herein.

[0192] The above specific embodiments do not constitute a limitation to the protection scope of the present invention. Those skilled in the art should understand that various modifications, combinations, sub-combinations and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions and improvements made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.

Claims

1. A resume screening method, characterized in that, The method includes: Obtaining job description information; For at least one alternative resume, determining key features corresponding to the job description information and weights of each of the key features according to data of the alternative resume; Determining a key vector of the alternative resume according to each of the key features and weights of each of the key features; Determining output vectors of each of the primary capsules according to the key vector of the alternative resume, a primary coefficient in the trained weight coefficients, and at least one primary capsule in the primary capsule layer; Determining output vectors of each of the secondary capsules according to the output vectors of each of the primary capsules, a secondary coefficient in the weight coefficients, and at least one secondary capsule in the secondary capsule layer; Determining a matching probability between the alternative resume and the job description information according to the output vectors of each of the secondary capsules; Screening out a target resume corresponding to the job description information from each of the alternative resumes according to the matching probability between each of the alternative resumes and the job description information.

2. The method according to claim 1, wherein The weight coefficients are obtained through training in the following manner: Determining output vectors of each of the primary capsules according to the key vector of the alternative resume, at least one primary coefficient in the weight coefficients, and at least one primary capsule in the primary capsule layer; Determining output vectors of each of the secondary capsules according to the output vectors of each of the primary capsules, a secondary coefficient in the weight coefficients, and at least one secondary capsule in the secondary capsule layer; Calculating a matching degree between the output vectors of each of the primary capsules and the output vectors of each of the secondary capsules; Updating the secondary coefficient and each of the primary coefficients according to the matching degree; Determining the trained weight coefficients when it is determined that the training of the weight coefficients is completed.

3. The method according to claim 1, wherein The determining key features corresponding to the job description information and weights of each of the key features according to data of the alternative resume includes: Extracting at least one resume key information from the data of the alternative resume; Combining each of the resume key information to obtain at least one feature combination; Calculating a confidence level and a support degree of each of the feature combinations; Determining at least one key combination from each of the feature combinations according to the confidence level and the support degree of each of the feature combinations, and determining resume key information in each of the key combinations as key features; Calculating a matching degree between each key feature in each of the key combinations and the job description information to obtain weights of each key feature in each of the key combinations.

4. The method according to claim 3, wherein After determining at least one key combination from each of the feature combinations according to the confidence level and the support degree of each of the feature combinations, it further includes: Calculating an information gain of each key combination; Filtering each of the key combinations according to the information gain of each key combination.

5. The method according to claim 3, characterized in that, The extracting at least one resume key information from the data of the alternative resume includes: Extracting at least one resume alternative information from the data of the alternative resume; Filtering each of the resume alternative information according to the job description information to obtain at least one resume key information.

6. The method according to claim 1, wherein It further includes: Obtain historical recruitment data of at least one sample job information, where the historical recruitment data of the sample job information includes: at least one sample resume and the true value screening results of each sample resume for the sample job information; For each sample resume, determine the key features corresponding to each sample job information and the weights of each key feature according to the data of the sample resume; Determine the key vector of the sample resume according to each key feature and the weights of each key feature; According to the key vectors of each sample job information, the primary coefficients in the trained weight coefficients, and at least one primary capsule in the primary capsule layer, determine the output vectors of each primary capsule; According to the output vectors of each primary capsule, the advanced coefficients in the weight coefficients, and at least one advanced capsule in the advanced capsule layer, determine the output vectors of each advanced capsule; Determine the matching probability between the sample resume and each sample job information according to the output vectors of each advanced capsule; Determine the predicted screening results corresponding to each sample job information according to the matching probabilities between each sample resume and each sample job information; Adjust the structure of the capsule layer according to the difference between the predicted screening results of each sample job information and the true value screening results of each sample job information, where the capsule layer includes a primary capsule layer and an advanced capsule layer.

7. The method according to claim 3, wherein The determining at least one key combination among each feature combination according to the confidence level and support degree of each feature combination includes: Determine at least one key combination among each feature combination according to a pre-determined confidence level threshold, support degree threshold, the confidence level and support degree of each feature combination; The method further includes: Update the confidence level threshold and the support degree threshold according to the difference between the predicted screening results of each sample job information and the true value screening results of each sample job information.

8. A resume screening device, characterized in that, The device includes: A description information acquisition module, configured to acquire job description information; A feature selection module, configured to determine, for at least one alternative resume, the key features corresponding to the job description information and the weights of each key feature according to the data of the alternative resume; A capsule network, configured to determine the key vector of the alternative resume according to each key feature and the weights of each key feature; The capsule network is configured to determine the output vectors of each primary capsule according to the key vector of the alternative resume, the primary coefficients in the trained weight coefficients, and at least one primary capsule in the primary capsule layer; The capsule network is configured to determine the output vectors of each advanced capsule according to the output vectors of each primary capsule, the advanced coefficients in the weight coefficients, and at least one advanced capsule in the advanced capsule layer; The capsule network is configured to determine the matching probability between the alternative resume and the job description information according to the output vectors of each advanced capsule; A resume screening module, configured to screen out the target resume corresponding to the job description information from each alternative resume according to the matching probability between each alternative resume and the job description information.

9. An electronic device, characterized in that, The electronic device includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein, the memory stores a computer program executable by the at least one processor, and the computer program is executed by the at least one processor to enable the at least one processor to execute the resume screening method according to any one of claims 1-7.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions for implementing the resume screening method according to any one of claims 1-7 when executed by a processor.