Personal resume correlation degree analysis method and device, equipment and medium

By using multiple matching methods and multidimensional feature analysis, a resume association graph is constructed, which solves the problem of low accuracy in the association analysis of resumes. It enables accurate identification and visualization of social relationships between resumes, supporting in-depth talent operation decision-making.

CN120996009APending Publication Date: 2025-11-21BEIJING HYDROPHIS NETWORK TECH CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202511128125.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-12
Publication Date
2025-11-21

AI Technical Summary

Technical Problem

Existing resume analysis methods have low accuracy in correlation analysis and cannot fully explore the multidimensional features and deep social relationship networks in resumes, resulting in a lack of panoramic relationship insights in team building and industry talent network mapping.

Method used

Structured and unstructured text in resumes are extracted using a multiple matching method. Social relationship factors are identified based on multidimensional features to construct a correlation graph. The correlation between resumes is calculated. Techniques such as cosine similarity, dynamic time warping, and Jaccard coefficient are used, combined with user annotations to adjust weights, to construct an accurate resume correlation graph.

Benefits of technology

It improves the accuracy of correlation analysis between resumes, realizes the visualization and structured display of social relationships between resumes, and supports accurate decision-making in team building and industry talent networks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120996009A_ABST
    Figure CN120996009A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of feature analysis, and discloses a personal resume correlation degree analysis method and device, equipment and a medium, and the method comprises the steps: extracting a structured text and an unstructured text in each personal resume in a preset personal resume set through a multi-matching method, extracting multi-dimensional features of each personal resume in the personal resume set based on the structured text and the unstructured text, and identifying social relation factors among the resumes from the personal resume set according to the multi-dimensional features and a preset multi-dimensional weight, and according to the social relation factors, constructing an association graph between the resumes in the personal resume set, and according to the association graph, calculating the association degree between all the resumes in the personal resume set. The accuracy of correlation degree analysis between personal resumes can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of feature analysis technology, and in particular to a method, apparatus, equipment and medium for personal resume correlation analysis. Background Technology

[0002] In the fields of human resource management and talent recruitment, the demand for intelligent analysis of personal resumes is growing. Current technologies for analyzing personal resumes mostly remain at a superficial level, such as keyword matching and structured data comparison, making it difficult to fully uncover the multidimensional characteristics and deep social networks hidden within the resume.

[0003] On the one hand, existing technologies have significant shortcomings in the presentation and deep correlation mining of resume analysis results. At the result presentation level, matching information is often presented in a simple list format, failing to integrate resume data with the underlying social relationship network. At the correlation mining level, it cannot construct correlation graphs between different resumes based on multi-dimensional factors such as education level (e.g., alumni of the same institution), career trajectory (e.g., work experience), and social networks (e.g., shared contacts). It also struggles to reveal the hidden network structure within group resumes (e.g., alumni circles in specific industries, cross-enterprise professional communities) through techniques such as block segmentation and correlation path tracing. This lack of ability to mine and visualize implicit relationship networks directly results in a failure to provide decision-makers with a panoramic view of relationships in scenarios such as team building (e.g., quickly identifying talent combinations with a foundation for collaboration) and industry talent network mapping (e.g., locating core talent and their surrounding circles), severely limiting the application value of resume analysis technology in in-depth talent management scenarios.

[0004] On the other hand, existing solutions typically divide resume text into structured fields (such as name and university) and unstructured fields (such as project experience and skill descriptions), but lack the ability to deeply analyze the semantics of unstructured text. For example, analyzing job titles and project experience solely through keyword matching fails to capture the temporal logic (such as the impact of tenure on career trajectory), semantic hierarchy (such as the importance differences between different job levels), and domain specialization (such as the relevance of industry-specific skills) within the text. This results in low accuracy in existing resume analysis methods when analyzing the correlation between resumes. Summary of the Invention

[0005] This invention provides a method, apparatus, device, and medium for analyzing the correlation between personal resumes, in order to solve the problem of low accuracy in analyzing the correlation between resumes in existing resume analysis methods.

[0006] Firstly, a method for analyzing the correlation of personal resumes is provided, including: The structured and unstructured text of each resume in the preset set of resumes is extracted using a multi-matching method. Based on the structured text and the unstructured text, multidimensional features of each personal resume in the personal resume collection are extracted; Based on the multidimensional features and preset multidimensional weights, social relationship factors between each resume are identified from the personal resume set. Based on the social relationship factors, a correlation graph is constructed between each resume in the personal resume set, and the correlation degree between all resumes in the personal resume set is calculated based on the correlation graph.

[0007] Secondly, a personal resume correlation analysis device is provided, including: The multidimensional feature extraction module is used to extract the structured and unstructured text of each personal resume in the preset set of personal resumes through multiple matching methods, and extract multidimensional features of each personal resume in the set of personal resumes based on the structured and unstructured text. The social relationship identification module is used to identify social relationship factors between each resume in the set of personal resumes based on the multidimensional features and preset multidimensional weights. The association graph construction module is used to construct an association graph between each resume in the personal resume set based on the social relationship factors, and to calculate the degree of association between all resumes in the personal resume set based on the association graph.

[0008] Thirdly, a computer device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the steps of the above-described personal resume correlation analysis method.

[0009] Fourthly, a computer-readable storage medium is provided, which stores a computer program that, when executed by a processor, implements the steps of the aforementioned personal resume correlation analysis method.

[0010] The aforementioned method, apparatus, equipment, and medium for personal resume correlation analysis can extract structured and unstructured text from each resume in a preset set of personal resumes using a multiple matching method. Based on the structured and unstructured text, multidimensional features of each resume in the set are extracted. Social relationship factors between resumes are identified based on these multidimensional features and preset multidimensional weights. A correlation graph is constructed based on these social relationship factors, and the correlation degree between all resumes in the set is calculated using the correlation graph. This improves the accuracy of personal resume correlation analysis. Attached Figure Description

[0011] To more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings used in the description of the embodiments of the present invention will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0012] Figure 1 This is a schematic diagram of an application environment for the personal resume correlation analysis method in one embodiment of the present invention; Figure 2 This is a flowchart illustrating a personal resume correlation analysis method according to an embodiment of the present invention; Figure 3 This is a schematic diagram of a personal resume correlation analysis device according to an embodiment of the present invention; Figure 4 This is a schematic diagram of the structure of a computer device according to an embodiment of the present invention; Figure 5 This is another structural schematic diagram of a computer device according to one embodiment of the present invention. Detailed Implementation

[0013] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0014] The personal resume correlation analysis method provided in this invention can be applied to, for example... Figure 1 In this application environment, the client communicates with the server via a network. The server can extract multi-dimensional features of each resume in the resume collection based on the structured and unstructured text from the client. Based on these multi-dimensional features and preset multi-dimensional weights, it identifies social relationship factors between each resume in the resume collection. Based on these social relationship factors, it constructs a correlation graph between each resume in the resume collection and calculates the correlation degree between all resumes in the resume collection based on the correlation graph. This improves the accuracy of correlation analysis between resumes. The client can be, but is not limited to, various personal computers, laptops, smartphones, tablets, and portable wearable devices. The server can be implemented using a standalone server or a server cluster consisting of multiple servers. The invention will be described in detail below through specific embodiments.

[0015] Please see Figure 2 As shown, Figure 2 A flowchart illustrating the personal resume correlation analysis method provided in this embodiment of the invention includes the following steps: S1. Extract the structured and unstructured text from each resume in the preset set of resumes using a multi-matching method.

[0016] In this embodiment of the invention, the step of extracting the structured and unstructured text from each resume in a preset set of resumes using a multiple matching method includes: Extract the visual features of the individual's resume; Based on the visual features, the personal resume is matched with similar resume templates in a preset resume template library to obtain matching results; If the matching result is successful, the structured text and unstructured text in the personal resume are confirmed according to the resume template corresponding to the matching result; If the matching result is a failure, then all field areas in the personal resume are identified based on the visual features; Calculate the structural rigidity score and semantic complexity score for each field region in the personal resume; The structured and unstructured regions in the personal resume are determined based on the structural solidity score and the semantic complexity score. Perform structured text keyword matching on the structured region to obtain keyword matching results, and confirm the structured text of the structured region based on the keyword matching results; The title field of the unstructured region is identified, and the title field is matched with a preset unstructured text title corpus to obtain the corpus matching result. The unstructured text of the unstructured region is confirmed based on the corpus matching result.

[0017] In detail, the extraction of visual features of the personal resume refers to the extraction of layout features of the personal resume (margins, number of columns, title position, etc.).

[0018] In detail, matching the personal resume with similar resume templates in a preset resume template library based on the visual features involves comparing the visual features of the resume to be parsed with standard templates in the preset template library (such as common job application resumes, academic resumes, industry-customized templates, etc.).

[0019] In detail, confirming the structured and unstructured text in the personal resume based on the corresponding resume template in the matching result means extracting the structured and unstructured text at fixed positions based on the resume template.

[0020] In detail, the structure solidity score is used to measure the degree of standardization of field area format, with a value range of 0-1. It can be calculated by analyzing the uniformity of text format (such as whether tables or fixed date formats are used) and the fixedness of area location (such as contact information in headers / footers). The higher the score, the more likely it is to be structured text (such as name and education level).

[0021] In detail, the semantic complexity score is used to measure the semantic certainty of the content of the field area, and the value ranges from 0 to 1 (the higher the score, the more ambiguous the semantics, and the more likely it is to be unstructured text).

[0022] In detail, the semantic complexity score can be obtained by analyzing whether the region contains a large number of adjectives and long sentences (such as "self-evaluation" which is mostly descriptive language and has high complexity), and by analyzing the proportion of entity words (such as names of people, organizations, and times) in the region (such as company names and positions in "work experience" which are low-complexity entities).

[0023] S2. Extract multidimensional features of each personal resume in the personal resume set based on the structured text and the unstructured text.

[0024] In this embodiment of the invention, the extraction of multidimensional features from each resume in the personal resume set based on the structured and unstructured text is achieved by extracting structured (e.g., name) and unstructured (e.g., project experience) text from the resumes. Text regions are segmented using similar resume template matching or visual features. For unstructured text, cosine similarity and dynamic time warping are used to calculate semantic and career trajectory similarity. For structured text, Jaccard coefficients are used to match social relationship factors, and these are combined to form multidimensional features.

[0025] In this embodiment of the invention, the step of extracting multidimensional features of each personal resume in the personal resume collection based on the structured text and the unstructured text includes: Based on cosine similarity, the semantic similarity between each personal resume in the personal resume set and the unstructured text in the preset reference resume is calculated. Based on dynamic time warping, the similarity of career trajectory between each personal resume in the personal resume set and the reference resume is calculated according to the unstructured text. Based on the Jaccard coefficient, the social relationship factor of each personal resume in the personal resume set and the reference resume is calculated according to the structured text. The semantic similarity, career trajectory similarity, and social relationship factors are combined to obtain the multidimensional features.

[0026] In this embodiment of the invention, the structured text may refer to information such as name, age, address, and university of graduation in the personal resume.

[0027] In this embodiment of the invention, the unstructured text may refer to fields such as job title, length of service, project experience, and personal skills in the personal resume.

[0028] In detail, the dynamic time warping method calculates the career trajectory similarity between each personal resume in the personal resume set and the reference resume based on the unstructured text, converts the employment time sequence in the resume (such as "2020.01-2022.12 Engineer of a certain company") into a timeline sequence, aligns the time sequences of different resumes through the dynamic time warping algorithm, calculates trajectory differences such as employment duration and job change frequency, and outputs a similarity value of 0-1.

[0029] In this embodiment of the invention, the multidimensional features include semantic similarity between the personal resume and a preset reference resume, career trajectory similarity, and social relationship factors.

[0030] Specifically, the preset reference resume can be a personal resume designed according to the job requirements of the target position; or, the preset reference resume can be the personal resume of an employee in the target position.

[0031] In this embodiment of the invention, the step of calculating the semantic similarity between each personal resume in the personal resume set and the unstructured text in a preset reference resume based on cosine similarity includes: Extract job title keywords, professional skill keywords, and project experience from the unstructured text; The project experience is segmented into sentences to obtain the segmentation results. Useless sentences are filtered from the segmentation results to obtain the project experience filtering results. Deep semantic encoding is performed on the job title keywords, the professional skill keywords, and the project experience filtering results to obtain job title codes, professional skill codes, and project experience codes. Semantic vectors for the job title code, the professional skill code, and the project experience code are generated based on a hierarchical attention mechanism, resulting in a job title vector, a professional skill vector, and a project experience vector. Dynamic job weights are set for the job name vector based on a preset job level pyramid diagram and the job name keywords. Based on time-series awareness and the project experience, dynamic project experience weights are set for the project experience vector; Calculate the cosine similarity between each individual resume and the reference resume for the vectors of job title, professional skills, and project experience, to obtain the job similarity, skill similarity, and project experience similarity. The semantic similarity is obtained by weighting and summing the job similarity, skill similarity, and project experience similarity according to the dynamic job weight, the preset professional skill weight, and the dynamic project experience weight.

[0032] In detail, the filtering of useless sentences in the sentence processing results can refer to filtering out some non-substantive content, such as "efficient collaboration" and "positive communication".

[0033] In detail, the deep semantic encoding of the job title keywords, the professional skill keywords, and the project experience filtering results can be performed using a pre-trained professional domain model.

[0034] In detail, the pre-trained occupational domain model is a language model fine-tuned on an occupational text corpus (such as job postings and industry reports).

[0035] In detail, the job level pyramid diagram is a diagram of predefined job hierarchy relationships (such as "junior engineer → senior engineer → technical manager → technical director") that quantifies the differences in importance of different jobs and sets different weights.

[0036] In detail, the dynamic project experience weighting of the project experience vector based on time sequence awareness and project experience is set by considering the impact of the project occurrence time on the current ability. The more recent the project experience time, the higher the weight of the project experience.

[0037] In detail, the hierarchical attention mechanism is a deep learning technique used to assign dynamic attention weights to different levels of text in a resume, such as job titles, professional skills, and project experience, and to prioritize capturing key semantics (such as senior job titles and core skill terms).

[0038] In this embodiment of the invention, the step of calculating the social relationship factor of each personal resume in the personal resume set and the reference resume based on the Jaccard coefficient according to the structured text includes: The structured text of each personal resume in the personal resume collection and the reference resume is attribute-encoded to obtain an attribute vector; Based on the attribute vectors of an individual's resume, similar contacts are matched against a pre-defined vector database to obtain a set of similar contacts; Based on the attribute vectors of the reference resume, similar contacts are matched in a preset vector database to obtain a set of reference contacts; Calculate the Jaccard coefficient between the set of similar contacts and the set of reference contacts for each personal resume to obtain the social relationship factor for each personal resume.

[0039] In detail, the Jaccard coefficient is a statistical measure used to measure the similarity between two sets. Its core idea is to quantify their similarity by comparing the ratio of the size of the intersection to the size of the union of the two sets. In detail, the step of encoding the structured text of each personal resume and the reference resume in the personal resume collection to obtain an attribute vector means converting the structured text into a numerical vector. For example, categorical attributes (such as "industry": Internet / Finance / Healthcare) are converted into vectors through one-hot encoding; continuous attributes (such as "years of work experience") are directly normalized and used as vector dimensions; and textual attributes (such as "company name") are encoded into semantic vectors through a pre-trained model (such as BERT).

[0040] In detail, the vector database is a pre-built large-scale contact vector library that stores a large number of attribute vectors of known individuals (such as industry experts, job seekers, and recruiters).

[0041] In this embodiment of the invention, by extracting multidimensional features of each personal resume in the personal resume set based on the structured text and the unstructured text, the accuracy of subsequent identification of social relationship factors between each resume can be improved.

[0042] S3. Identify the social relationship factors between each resume from the set of personal resumes based on the multidimensional features and preset multidimensional weights.

[0043] In this embodiment of the invention, the step of identifying social relationship factors between each resume within the personal resume set based on the multidimensional features and preset multidimensional weights is achieved by calculating a similarity score by combining multidimensional features and preset weights, filtering and sorting the resumes, and then adjusting the weights based on user annotations. Social relationship matching is obtained by analyzing structured text such as educational background similarity and address association, and the scores are then fused to obtain the social relationship factors between resumes.

[0044] In this embodiment of the invention, the step of identifying social relationship factors between each resume within the personal resume set based on the multidimensional features and preset multidimensional weights includes: The similarity score of each personal resume in the personal resume set is calculated based on the multidimensional features and preset multidimensional weights. Based on the similarity score, a preset number of personal resumes are selected from the personal resume set to obtain a sorted personal resume set. Obtain the user's acceptance result annotation for the sorted personal resume set, and adjust the sorted personal resume set according to the acceptance result annotation to obtain the adjusted resume set; The social relationship matching degree between each individual resume in the adjusted resume set is identified based on the structured text. The social relationship factor between each personal resume in the ranked set is calculated based on the social relationship matching degree and the similarity score.

[0045] In this embodiment of the invention, the similarity score of each personal resume in the personal resume set is calculated by weighting and fusing semantic similarity, career trajectory similarity and social relationship factors in the multidimensional features.

[0046] In this embodiment of the invention, the step of calculating the similarity score of each personal resume in the personal resume set based on the multidimensional features and preset multidimensional weights can be achieved by multiplying the semantic similarity, career trajectory similarity, and social relationship factor in the multidimensional features by preset first weights, second weights, and third weights, respectively, and then summing them to obtain the similarity score.

[0047] Specifically, the first weight can be 0.5, the second weight can be 0.3, and the third weight can be 0.2.

[0048] In this embodiment of the invention, obtaining user acceptance result annotations for the sorted personal resume set, and adjusting the sorted personal resume set according to the acceptance result annotations to obtain an adjusted resume set, includes: Based on the acceptance result annotation, determine whether the user accepts the sorting order of the personal resume sorting set; If accepted, the sorted set of personal resumes is confirmed as the adjusted set of resumes; If not accepted, the annotation content of each personal resume in the personal resume sorting set is identified based on the annotation data; Based on the annotations in each individual resume, determine the overall weighting adjustment target and direction; The multidimensional features are dynamically adjusted according to the weight adjustment target and the weight adjustment direction to obtain the adjusted multidimensional weights. Based on the adjusted multidimensional weights and the multidimensional features, the similarity score of each personal resume in the personal resume ranking set is recalculated. The personal resume ranking set is then re-ranked based on the recalculated similarity to obtain the adjusted resume set.

[0049] Specifically, the labeled data includes the user's labeling results regarding whether each personal resume in the sorted set of personal resumes is acceptable.

[0050] In this embodiment of the invention, the annotation content of the personal resume refers to the user's annotation of the similarity between the unstructured text in the personal resume and the unstructured text in the reference resume.

[0051] In detail, the annotation content of the personal resume may refer to the annotation content obtained after the user annotates the content that does not match the reference resume the most.

[0052] In detail, the step of determining the overall weight adjustment target and direction based on the annotation content of each individual resume refers to identifying the features (semantic similarity or career trajectory similarity, excluding social relationship factors) corresponding to the annotation content in the individual resume. For example, if the annotation content is project experience, then the feature corresponding to semantic similarity in the individual resume is the same; if the annotation content is a time series of employment, then the feature corresponding to career trajectory similarity in the individual resume is the same. After identifying the corresponding features, the number of features corresponding to all annotation content is counted to determine the features that need to be adjusted. For example, if there are more individual resumes with semantic similarity corresponding to the annotation content in the sorted set of individual resumes, then the adjustment direction is to reduce the first weight corresponding to semantic similarity and increase the second weight corresponding to career trajectory similarity.

[0053] In this embodiment of the invention, the step of sorting and selecting a preset number of personal resumes from the personal resume set based on the similarity score means sorting each personal resume in descending order of similarity score, and then selecting the top N resumes from the sorted results as the final output, where N is the preset number. The preset number can be set according to the sample size of the personal resume set. For example, if the personal resume set contains 20 personal resumes, the preset number can be half of the sample size of the personal resume set (10).

[0054] In this embodiment of the invention, a preset number of personal resumes are selected from the personal resume set based on the similarity score to obtain a sorted set of personal resumes. This process can filter out irrelevant samples, reduce the number of samples, and improve processing efficiency.

[0055] In this embodiment of the invention, the step of identifying the social relationship matching degree between each individual resume in the adjusted resume set based on the structured text includes: Extract basic identity information and educational background information from the structured text; Calculate the similarity of educational information among the individual resumes in the adjusted resume set to obtain the educational similarity. Determine whether the educational similarity is greater than or equal to a preset educational similarity threshold; If the educational similarity is greater than or equal to the educational similarity threshold, then a first weighting coefficient is determined based on the educational information between the individual resumes; Multiply the first weight coefficient by the preset first matching degree to obtain the social relationship matching degree; If the educational similarity is less than the educational similarity threshold, then a second weighting coefficient is determined based on the basic identity information between the personal resumes; Multiply the second weight coefficient by the preset second matching degree to obtain the social relationship matching degree.

[0056] In detail, the calculation of the similarity of educational information between personal resumes can be calculated by calculating the similarity of the graduating institutions contained in the educational information. When the educational similarity is greater than the educational similarity threshold, it indicates that the applicants corresponding to the two personal resumes may have graduated from the same university.

[0057] In this embodiment of the invention, the preset first matching degree can be 0.8, and the preset second matching degree can be 1.

[0058] Specifically, the determination of the first weighting coefficient based on educational information between individual resumes includes: Extract the major and graduation date from the academic information; Calculate the semantic similarity of the professional names between individuals' resumes to obtain the professional similarity score; The professional similarity is normalized to obtain the normalized similarity. Determine if the graduation dates are the same across individuals' resumes; If they are the same, then the normalized similarity is confirmed to be the first weight coefficient; If they are different, calculate the difference in graduation time between the individual resumes to obtain the graduation time difference; The first weight coefficient is obtained by multiplying the reciprocal of the graduation time difference by the normalized similarity.

[0059] Specifically, the confirmation of the second weighting coefficient based on basic identity information between personal resumes includes: Calculate the address similarity between personal resumes based on the aforementioned basic identity information; Determine whether the address similarity is less than a preset address similarity threshold; If the address similarity is greater than or equal to the address similarity threshold, then the kinship relationship between the personal resumes is determined based on the basic identity information to obtain the kinship probability, and the kinship probability is confirmed as the second weighting coefficient. If the address similarity is less than the address similarity threshold, then the second weighting coefficient is confirmed to be 0.2.

[0060] In detail, the step of calculating the address similarity between personal resumes based on the basic identity information may involve extracting the address information from the basic identity information, converting the address information into latitude and longitude, calculating the geographical distance based on the latitude and longitude, and normalizing the geographical distance into an address similarity score, with the closer the distance, the higher the similarity score.

[0061] In detail, the step of determining the kinship between personal resumes based on the basic identity information can be achieved by extracting the name and age from the basic identity information, determining whether the surnames are the same, and if they are the same, calculating the age difference. If the age difference is in the range of 0-5 and 25-30, it is determined that the submitters of the two personal resumes may be related as father and son / father and daughter or brothers and sisters.

[0062] Specifically, when it is determined from the basic identity information that the surnames are different, the probability of the kinship relationship can be set to 0.1.

[0063] Specifically, when it is determined from the basic identity information that the surnames are the same, but the age difference is not in the ranges of 0-5 and 25-30, the probability of the kinship relationship can be set to 0.3.

[0064] Specifically, when it is determined from the basic identity information that the surnames are the same and the age difference is in the ranges of 0-5 and 25-30, the probability of the kinship relationship can be set to 0.8.

[0065] In this embodiment of the invention, the step of calculating the social relationship factor between each personal resume in the personal resume ranking set based on the social relationship matching degree and the similarity score can be achieved by weighted fusion of the social relationship matching degree and the similarity score.

[0066] In detail, the weighted fusion of the social relationship matching degree and the similarity score can be performed by assigning each a weight of 0.5.

[0067] In this embodiment of the invention, by identifying the social relationship factors between each resume in the personal resume set based on the multidimensional features and preset multidimensional weights, the accuracy of subsequently constructing the association graph between each resume in the personal resume set can be improved.

[0068] S4. Construct a correlation graph between each resume in the personal resume set based on the social relationship factors, and calculate the correlation degree between all resumes in the personal resume set based on the correlation graph.

[0069] In this embodiment of the invention, the construction of a relationship graph between each resume within the personal resume set based on the social relationship factor is achieved by using the social relationship factor as the connection weight to construct the resume relationship graph. The graph is then divided into blocks, and the strength of the relationships between blocks is calculated. Strong relationship paths (such as education level and social connections) are analyzed to visualize and structure the implicit relationships between resumes.

[0070] In this embodiment of the invention, the social relationship factor refers to the social relationship factor among all personal resumes in the personal resume ranking set.

[0071] In this embodiment of the invention, constructing a correlation graph between each resume in the personal resume set based on the social relationship factor means that any personal resume in the personal resume ranking set has a social relationship factor calculated with other personal resumes in the personal resume ranking set, and a correlation is established between it and the personal resume with the highest social relationship factor, ultimately obtaining the overall correlation graph.

[0072] In this embodiment of the invention, in order to make the association map more accurate, the embodiment of the invention also includes regional association marking / display within the association map, that is, treating the resume within a certain association range within the association map as a local block and analyzing its association with other blocks.

[0073] In detail, after constructing the relationship graph between each resume in the personal resume set based on the social relationship factors, the method further includes: Identify identical blocks based on a preset social relationship factor threshold; The correlation strength between all blocks is calculated using the fully connected average method. Strongly correlated blocks are identified based on the correlation strength. Perform association path analysis on strongly correlated blocks to obtain the path analysis results; Block association marking is performed based on the path analysis results.

[0074] In detail, the identification of identical blocks based on a preset social relationship factor threshold is to divide the personal resumes in the set of resumes with a correlation strength higher than the threshold into the same local blocks according to the preset threshold (such as social relationship factor ≥ 0.6).

[0075] In detail, the method of using the fully connected average to calculate the association strength between all blocks is described. For all records in different blocks, the social relationship factor is calculated pairwise, and the average value is taken as the overall association strength between blocks.

[0076] In detail, the step of identifying strongly correlated blocks based on the correlation strength involves setting a strong correlation threshold (e.g., correlation strength > 0.5) and filtering out block pairs that are higher than the threshold.

[0077] In detail, the process of performing association path analysis on strongly related blocks to obtain path analysis results, and marking the blocks according to the path analysis results, refers to analyzing the dominant factors that generate a strong association between two blocks. For example, if two blocks are associated through "alumni of the same institution" (educational similarity ≥ threshold), the association path is marked as "educational association"; if they are associated through "common social contacts" (high Jaccard coefficient), they are marked as "social association".

[0078] In this embodiment of the invention, the step of calculating the correlation degree between all resumes in the personal resume set based on the correlation graph is as follows: for a single resume, the social relationship factor between each resume in the correlation graph is directly used as the correlation degree between the two; for blocks in the correlation graph divided according to a preset social relationship factor threshold, the fully connected average method is used to calculate the social relationship factor for each pair of resumes in different blocks and take the average value to obtain the overall correlation degree between blocks.

[0079] As can be seen, in the above scheme, based on a pre-set set of personal resumes, structured text (such as name, age, etc.) and unstructured text (such as project experience, skill descriptions, etc.) are extracted from each resume. Visual features (layout, margins, etc.) of the resumes are extracted for similar resume template matching. If a match is successful, the text is located according to the template; otherwise, structured and unstructured regions are divided using structural fixity scores (measuring the degree of format standardization) and semantic complexity scores (measuring content certainty). Then, for unstructured text, cosine similarity is used in conjunction with a hierarchical attention mechanism, a pre-trained career domain model, and a job level pyramid to calculate semantic similarity. Career trajectory similarity is calculated using a dynamic time warping algorithm. For structured text, social similarity is calculated based on the Jaccard coefficient through attribute encoding and vector database matching. The process involves: 1) summarizing relationship factors to form multidimensional features; 2) calculating similarity scores based on multidimensional features and preset weights (e.g., semantic similarity 0.5, career trajectory similarity 0.3, social factors 0.2); 3) selecting a preset number of resumes and dynamically adjusting weights based on user annotations; 4) calculating social relationship matching degree using educational information (school, major, time similarity) and basic identity information (address, surname, age difference); and 5) fusing similarity scores to obtain social relationship factors; and 6) establishing connections between each resume and the other resumes with the strongest associations based on social relationship factors to construct a relationship graph. 7) dividing the resume into blocks based on thresholds and calculating the relationship strength between blocks using the fully connected average method. 8) analyzing the dominant factors of strongly associated blocks (e.g., educational association, social association) and marking paths to achieve visualization and structured analysis of social relationships between resume sets.

[0080] It should be understood that the sequence number of each step in the above embodiments does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present invention.

[0081] In one embodiment, a personal resume correlation analysis device is provided, which corresponds one-to-one with the personal resume correlation analysis method described in the above embodiments. For example... Figure 3 As shown, the personal resume correlation analysis device includes a multi-dimensional feature extraction module 101, a social relationship identification module 102, and a correlation graph construction module 103. Detailed descriptions of each functional module are as follows: The multidimensional feature extraction module 101 is used to extract the structured text and unstructured text in each personal resume in the preset personal resume set through multiple matching method, and extract multidimensional features of each personal resume in the personal resume set based on the structured text and the unstructured text. The social relationship identification module 102 is used to identify the social relationship factors between each resume in the personal resume set based on the multidimensional features and preset multidimensional weights. The association graph construction module 103 is used to construct an association graph between each resume in the personal resume set based on the social relationship factors, and to calculate the degree of association between all resumes in the personal resume set based on the association graph.

[0082] In one embodiment, the multidimensional feature extraction module 101, when performing the extraction of structured and unstructured text from each personal resume in the pre-acquired set of personal resumes, specifically performs the following: Extract the visual features of the individual's resume; Based on the visual features, the personal resume is matched with similar resume templates in a preset resume template library to obtain matching results; If the matching result is successful, the structured text and unstructured text in the personal resume are confirmed according to the resume template corresponding to the matching result; If the matching result is a failure, then all field areas in the personal resume are identified based on the visual features; Calculate the structural rigidity score and semantic complexity score for each field region in the personal resume; The structured and unstructured regions in the personal resume are determined based on the structural solidity score and the semantic complexity score. Perform structured text keyword matching on the structured region to obtain keyword matching results, and confirm the structured text of the structured region based on the keyword matching results; The title field of the unstructured region is identified, and the title field is matched with a preset unstructured text title corpus to obtain the corpus matching result. The unstructured text of the unstructured region is confirmed based on the corpus matching result.

[0083] In one embodiment, the multidimensional feature extraction module 101, when performing the extraction of multidimensional features of each personal resume in the personal resume set based on the structured text and the unstructured text, is specifically used for: Based on cosine similarity, the semantic similarity between each personal resume in the personal resume set and the unstructured text in the preset reference resume is calculated. Based on dynamic time warping, the similarity of career trajectory between each personal resume in the personal resume set and the reference resume is calculated according to the unstructured text. Based on the Jaccard coefficient, the social relationship factor of each personal resume in the personal resume set and the reference resume is calculated according to the structured text. The semantic similarity, career trajectory similarity, and social relationship factors are combined to obtain the multidimensional features.

[0084] In one embodiment, the multidimensional feature extraction module 101, when performing the calculation of the semantic similarity between each personal resume in the personal resume set and the unstructured text in the preset reference resume based on cosine similarity, is specifically used for: Extract job title keywords, professional skill keywords, and project experience from the unstructured text; The project experience is segmented into sentences to obtain the segmentation results. Useless sentences are filtered from the segmentation results to obtain the project experience filtering results. Deep semantic encoding is performed on the job title keywords, the professional skill keywords, and the project experience filtering results to obtain job title codes, professional skill codes, and project experience codes. Semantic vectors for the job title code, the professional skill code, and the project experience code are generated based on a hierarchical attention mechanism, resulting in a job title vector, a professional skill vector, and a project experience vector. Dynamic job weights are set for the job name vector based on a preset job level pyramid diagram and the job name keywords. Based on time-series awareness and the project experience, dynamic project experience weights are set for the project experience vector; Calculate the cosine similarity between each individual resume and the reference resume for the vectors of job title, professional skills, and project experience, to obtain the job similarity, skill similarity, and project experience similarity. The semantic similarity is obtained by weighting and summing the job similarity, skill similarity, and project experience similarity according to the dynamic job weight, the preset professional skill weight, and the dynamic project experience weight.

[0085] In one embodiment, the social relationship identification module 102, when performing the step of identifying the social relationship factors between each resume from the personal resume set based on the multidimensional features and preset multidimensional weights, is specifically used for: The similarity score of each personal resume in the personal resume set is calculated based on the multidimensional features and preset multidimensional weights. Based on the similarity score, a preset number of personal resumes are selected from the personal resume set to obtain a sorted personal resume set. Obtain the user's acceptance result annotation for the sorted personal resume set, and adjust the sorted personal resume set according to the acceptance result annotation to obtain the adjusted resume set; The social relationship matching degree between each individual resume in the adjusted resume set is identified based on the structured text. The social relationship factor between each personal resume in the ranked set is calculated based on the social relationship matching degree and the similarity score.

[0086] In one embodiment, the social relationship identification module 102, when performing the step of identifying the social relationship matching degree between each individual resume in the adjusted resume set based on the structured text, is specifically used for: Extract basic identity information and educational background information from the structured text; Calculate the similarity of educational information among the individual resumes in the adjusted resume set to obtain the educational similarity. Determine whether the educational similarity is greater than or equal to a preset educational similarity threshold; If the educational similarity is greater than or equal to the educational similarity threshold, then a first weighting coefficient is determined based on the educational information between the individual resumes; Multiply the first weight coefficient by the preset first matching degree to obtain the social relationship matching degree; If the educational similarity is less than the educational similarity threshold, then a second weighting coefficient is determined based on the basic identity information between the personal resumes; Multiply the second weight coefficient by the preset second matching degree to obtain the social relationship matching degree.

[0087] In one embodiment, the association map construction module 103 is further configured to: Identify identical blocks based on a preset social relationship factor threshold; The correlation strength between all blocks is calculated using the fully connected average method. Strongly correlated blocks are identified based on the correlation strength. Perform association path analysis on strongly correlated blocks to obtain the path analysis results; Block association marking is performed based on the path analysis results.

[0088] This invention provides a personal resume correlation analysis device. Based on a preset set of personal resumes, it extracts structured text (such as name, age, etc.) and unstructured text (such as project experience, skill descriptions, etc.) from each resume. It then performs similar resume template matching by extracting visual features (layout, margins, etc.). If a match is successful, the text is located according to the template; otherwise, it divides the structured and unstructured regions using structural fixity scores (measuring the degree of format standardization) and semantic complexity scores (measuring content certainty). For unstructured text, it calculates semantic similarity using cosine similarity combined with a hierarchical attention mechanism, a pre-trained career domain model, and a job level pyramid graph. It also calculates career trajectory similarity using a dynamic time warping algorithm. For structured text, it matches attribute encoding and a vector database based on the Jaccard coefficient. The process involves calculating social relationship factors and summarizing them to form multidimensional features. Then, based on these multidimensional features and preset weights (e.g., semantic similarity 0.5, career trajectory similarity 0.3, social factors 0.2), a similarity score is calculated. A preset number of resumes are selected, and the weights are dynamically adjusted based on user annotations. Social relationship matching is then calculated using educational information (school, major, time similarity) and basic identity information (address, surname, age difference). The similarity scores are then integrated to obtain the social relationship factor. Finally, based on the social relationship factor, connections are established between each resume and the other resumes with the strongest associations, constructing a relationship graph. Blocks are then divided based on thresholds, and the strength of associations between blocks is calculated using the fully connected average method. The dominant factors of strongly associated blocks (e.g., educational association, social association) are analyzed, and path marking is performed to achieve visualization and structured analysis of social relationships between resume sets. Specific limitations of the personal resume association analysis device can be found in the limitations of the personal resume association analysis method described above, and will not be repeated here. Each module in the above-mentioned personal resume association analysis device can be implemented entirely or partially through software, hardware, or a combination thereof. The above modules can be embedded in the processor of the computer device in hardware form or independent of it, or they can be stored in the memory of the computer device in software form, so that the processor can call and execute the corresponding operations of the above modules.

[0089] In one embodiment, a computer device is provided, which may be a server, and its internal structure diagram may be as follows: Figure 4 As shown, the computer device includes a processor, memory, network interface, and database connected via a system bus. The processor provides computing and control capabilities. The memory includes non-volatile and / or volatile storage media and internal memory. The non-volatile storage media stores the operating system, computer programs, and database. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage media. The network interface is used to communicate with external clients via a network connection. When the computer program is executed by the processor, it implements the functions or steps of a personal resume correlation analysis method on the server side.

[0090] In one embodiment, a computer device is provided, which may be a client, and its internal structure diagram may be as follows: Figure 5 As shown, the computer device includes a processor, memory, network interface, display screen, and input system connected via a system bus. The processor provides computing and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage media. The network interface is used to communicate with an external server via a network connection. When the computer program is executed by the processor, it implements the functions or steps of a client-side method for personal resume correlation analysis. In one embodiment, a computer device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to perform the following steps: The structured and unstructured text of each resume in the preset set of resumes is extracted using a multi-matching method. Based on the structured text and the unstructured text, multidimensional features of each personal resume in the personal resume collection are extracted; Based on the multidimensional features and preset multidimensional weights, social relationship factors between each resume are identified from the personal resume set. Based on the social relationship factors, a correlation graph is constructed between each resume in the personal resume set, and the correlation degree between all resumes in the personal resume set is calculated based on the correlation graph.

[0091] In one embodiment, a computer-readable storage medium is provided having a computer program stored thereon, the computer program performing the following steps when executed by a processor: The structured and unstructured text of each resume in the preset set of resumes is extracted using a multi-matching method. Based on the structured text and the unstructured text, multidimensional features of each personal resume in the personal resume collection are extracted; Based on the multidimensional features and preset multidimensional weights, social relationship factors between each resume are identified from the personal resume set. Based on the social relationship factors, a correlation graph is constructed between each resume in the personal resume set, and the correlation degree between all resumes in the personal resume set is calculated based on the correlation graph.

[0092] It should be noted that the functions or steps that can be implemented by the computer-readable storage medium or computer device described above can be referred to the relevant descriptions on the server side and client side in the foregoing method embodiments. To avoid repetition, they will not be described one by one here.

[0093] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium, and when executed, it can include the processes of the embodiments of the above methods. Any references to memory, storage, databases, or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), dual data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.

[0094] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the above-described division of functional units and modules is used as an example. In practical applications, the above functions can be assigned to different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above.

[0095] Finally, it should be noted that if any software tools or components not belonging to this company appear in the embodiments of the application, they are merely illustrative examples and do not represent actual use. The embodiments described above are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be included within the protection scope of the present invention.

Claims

1. A method for analyzing the correlation of personal resumes, characterized in that, include: The structured and unstructured text of each resume in the preset set of resumes is extracted using a multi-matching method. Based on the structured text and the unstructured text, multidimensional features of each personal resume in the personal resume collection are extracted; Based on the multidimensional features and preset multidimensional weights, social relationship factors between each resume are identified from the personal resume set. Based on the social relationship factors, a correlation graph is constructed between each resume in the personal resume set, and the correlation degree between all resumes in the personal resume set is calculated based on the correlation graph.

2. The personal resume correlation analysis method as described in claim 1, characterized in that, The step of extracting structured and unstructured text from each resume in a pre-defined set of resumes using a multi-matching method includes: Extract the visual features of the individual's resume; Based on the visual features, the personal resume is matched with similar resume templates in a preset resume template library to obtain matching results; If the matching result is successful, the structured text and unstructured text in the personal resume are confirmed according to the resume template corresponding to the matching result; If the matching result is a failure, then all field areas in the personal resume are identified based on the visual features; Calculate the structural rigidity score and semantic complexity score for each field region in the personal resume; The structured and unstructured regions in the personal resume are determined based on the structural solidity score and the semantic complexity score. Perform structured text keyword matching on the structured region to obtain keyword matching results, and confirm the structured text of the structured region based on the keyword matching results; The title field of the unstructured region is identified, and the title field is matched with a preset unstructured text title corpus to obtain the corpus matching result. The unstructured text of the unstructured region is confirmed based on the corpus matching result.

3. The personal resume correlation analysis method as described in claim 1, characterized in that, The extraction of multidimensional features for each personal resume in the personal resume collection based on the structured text and the unstructured text includes: Based on cosine similarity, the semantic similarity between each personal resume in the personal resume set and the unstructured text in the preset reference resume is calculated. Based on dynamic time warping, the similarity of career trajectory between each personal resume in the personal resume set and the reference resume is calculated according to the unstructured text. Based on the Jaccard coefficient, the social relationship factor of each personal resume in the personal resume set and the reference resume is calculated according to the structured text. The semantic similarity, career trajectory similarity, and social relationship factors are combined to obtain the multidimensional features.

4. The personal resume correlation analysis method as described in claim 3, characterized in that, The step of calculating the semantic similarity between each personal resume in the personal resume set and the unstructured text in the preset reference resumes based on cosine similarity includes: Extract job title keywords, professional skill keywords, and project experience from the unstructured text; The project experience is segmented into sentences to obtain the segmentation results. Useless sentences are filtered from the segmentation results to obtain the project experience filtering results. Deep semantic encoding is performed on the job title keywords, the professional skill keywords, and the project experience filtering results to obtain job title codes, professional skill codes, and project experience codes. Semantic vectors for the job title code, the professional skill code, and the project experience code are generated based on a hierarchical attention mechanism, resulting in a job title vector, a professional skill vector, and a project experience vector. Dynamic job weights are set for the job name vector based on a preset job level pyramid diagram and the job name keywords. Based on time-series awareness and the project experience, dynamic project experience weights are set for the project experience vector; Calculate the cosine similarity between each individual resume and the reference resume for the vectors of job title, professional skills, and project experience, to obtain the job similarity, skill similarity, and project experience similarity. The semantic similarity is obtained by weighting and summing the job similarity, skill similarity, and project experience similarity according to the dynamic job weight, the preset professional skill weight, and the dynamic project experience weight.

5. The personal resume correlation analysis method as described in claim 1, characterized in that, The step of identifying social relationship factors between each resume within the personal resume set based on the multidimensional features and preset multidimensional weights includes: The similarity score of each personal resume in the personal resume set is calculated based on the multidimensional features and preset multidimensional weights. Based on the similarity score, a preset number of personal resumes are selected from the personal resume set to obtain a sorted personal resume set. Obtain the user's acceptance result annotation for the sorted personal resume set, and adjust the sorted personal resume set according to the acceptance result annotation to obtain the adjusted resume set; The social relationship matching degree between each individual resume in the adjusted resume set is identified based on the structured text. The social relationship factor between each personal resume in the ranked set is calculated based on the social relationship matching degree and the similarity score.

6. The personal resume correlation analysis method as described in claim 5, characterized in that, The step of identifying the social relationship matching degree between each individual resume in the adjusted resume set based on the structured text includes: Extract basic identity information and educational background information from the structured text; Calculate the similarity of educational information among the individual resumes in the adjusted resume set to obtain the educational similarity. Determine whether the educational similarity is greater than or equal to a preset educational similarity threshold; If the educational similarity is greater than or equal to the educational similarity threshold, then a first weighting coefficient is determined based on the educational information between the individual resumes; Multiply the first weight coefficient by the preset first matching degree to obtain the social relationship matching degree; If the educational similarity is less than the educational similarity threshold, then a second weighting coefficient is determined based on the basic identity information between the personal resumes; Multiply the second weight coefficient by the preset second matching degree to obtain the social relationship matching degree.

7. The personal resume correlation analysis method as described in claim 1, characterized in that, The method further includes: Identify identical blocks based on a preset social relationship factor threshold; The correlation strength between all blocks is calculated using the fully connected average method. Strongly correlated blocks are identified based on the correlation strength. Perform association path analysis on strongly correlated blocks to obtain the path analysis results; Block association marking is performed based on the path analysis results.

8. A personal resume correlation analysis device, characterized in that, include: The multidimensional feature extraction module is used to extract the structured and unstructured text of each personal resume in the preset set of personal resumes through multiple matching methods, and extract multidimensional features of each personal resume in the set of personal resumes based on the structured and unstructured text. The social relationship identification module is used to identify social relationship factors between each resume in the set of personal resumes based on the multidimensional features and preset multidimensional weights. The association graph construction module is used to construct an association graph between each resume in the personal resume set based on the social relationship factors, and to calculate the degree of association between all resumes in the personal resume set based on the association graph.

9. A computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the personal resume correlation analysis method as described in any one of claims 1 to 7.

10. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, it implements the personal resume correlation analysis method as described in any one of claims 1 to 7.

Citation Information

Cited By

  • Human resource file storage management system based on knowledge graph

    CN121581827A