Resume processing method and system based on multi-modal semantic matching

By using multimodal semantic matching technology to parse and match resumes, constructing candidate and job profiles, and generating recommendation reports using deep learning and large language models, this technology solves the problems of low efficiency and low accuracy in resume parsing and matching in existing technologies, and achieves efficient and accurate person-job matching.

CN122045514APending Publication Date: 2026-05-15BEIJING HESI HUIZHI INFORMATION TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
BEIJING HESI HUIZHI INFORMATION TECHNOLOGY CO LTD
Filing Date
2026-02-02
Publication Date
2026-05-15

Smart Images

  • Figure CN122045514A_ABST
    Figure CN122045514A_ABST
Patent Text Reader

Abstract

The invention provides a resume processing method and system based on multi-modal semantic matching, and the method comprises the steps: obtaining a resume document of a candidate, and processing the resume document to obtain a candidate portrait; semantic analysis is carried out on the position description of the target position, and a position portrait is constructed; and performing semantic matching on the candidate portrait and the post portrait based on a double-tower deep neural network model to obtain a person-post matching degree score of the resume document and the target post, and generating a candidate recommendation report based on a large language model. According to the method, the resume screening efficiency is improved, and meanwhile, the resume analysis accuracy is improved, so that the man-post matching is more accurate.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of data processing technology, and in particular to a resume processing method and system based on multimodal semantic matching. Background Technology

[0002] In large corporations or headhunting firms, HR professionals face thousands of resumes daily. These resumes vary widely in format (e.g., PDF / image / Word) and layout. Traditional resume screening methods rely on manual reading or simple keyword searches (such as Ctrl+F for "Java"), which leads to two problems: extremely low efficiency and significant missed opportunities, as excellent candidates may be filtered out simply because their resumes don't include specific keywords. Meanwhile, with increasing competition for talent, companies are focusing more on uncovering candidates' deeper potential and job fit, rather than simply matching qualifications / years of experience. Summary of the Invention

[0003] In view of this, the purpose of this invention is to provide a resume processing method and system based on multimodal semantic matching, so as to improve the efficiency of resume screening and the accuracy of resume parsing, thereby making the matching of people and jobs more accurate.

[0004] To achieve the above objectives, the technical solutions adopted in the embodiments of the present invention are as follows: In a first aspect, embodiments of the present invention provide a resume processing method based on multimodal semantic matching, comprising: obtaining a candidate's resume document and processing the resume document to obtain a candidate profile; performing semantic parsing on the job description of the target position to construct a job profile; performing semantic matching on the candidate profile and the job profile based on a dual-tower deep neural network model to obtain a person-job matching score between the resume document and the target position, and generating a candidate recommendation report based on a large language model.

[0005] Optionally, the resume document is processed to obtain a candidate profile, including: using a deep learning-based document layout analysis model to identify layout elements in the resume document and obtain the position coordinates and logical relationships of the layout elements to obtain the resume document structure; wherein, the layout elements include at least: text block areas, table areas, image areas, and heading levels; based on the resume document structure, multimodal information is extracted from the resume document and the multimodal information is fused to obtain the candidate profile.

[0006] Optionally, multimodal information is extracted from the resume document based on its structure, and the multimodal information is fused to obtain a candidate profile, including: for text blocks, named entity recognition is performed based on a fine-tuned large language model to obtain entity information; for image regions, visual feature information of the candidate is extracted based on a fine-tuned visual model; and the entity information and visual feature information are fused to obtain a candidate profile.

[0007] Optionally, semantic parsing is performed on the job description of the target position to construct a job profile, including: semantic parsing of the job description of the target position, extracting the job requirements of the target position, and constructing an initial job profile based on the job requirements; and expanding the initial job profile based on a preset industry knowledge graph to obtain a job profile.

[0008] Optionally, semantic matching of candidate profiles and job profiles is performed based on a dual-tower deep neural network model to obtain a person-job matching score between the resume document and the target job. This includes: inputting the candidate profile and job profile into the dual-tower deep neural network model for vectorization processing to obtain candidate profile vectors and job profile vectors; calculating the cosine similarity between the candidate profile vector and the job profile vector, and determining a first preset number of candidate resumes based on the cosine similarity; and interacting with the job description and candidate resumes based on a preset ranking model to obtain the matching result between the candidate resume and the target job and the person-job matching score.

[0009] Optionally, a candidate recommendation report is generated based on a large language model, including: generating a candidate recommendation report based on a large oracle model and matching results; generating a recommendation list based on the job-person matching score and the recommendation report; displaying the recommendation list to recruiters and obtaining their feedback.

[0010] Optional features include: acquiring candidate tracking data at each stage of the recruitment process and optimizing the dual-tower deep neural network model based on the tracking data.

[0011] Secondly, embodiments of the present invention provide a resume processing system based on multimodal semantic matching, comprising: a candidate profile construction module, used to obtain candidate resume documents and process the resume documents to obtain candidate profiles; a job profile construction module, used to perform semantic parsing on the job description of the target job and construct a job profile; and a person-job matching module, used to perform semantic matching on the candidate profile and job profile based on a dual-tower deep neural network model to obtain a person-job matching score between the resume document and the target job, and generate a candidate recommendation report based on a large language model.

[0012] Thirdly, embodiments of the present invention provide an electronic device including a processor and a memory, the memory storing computer-executable instructions executable by the processor, the processor executing the computer-executable instructions to implement the steps of any of the methods provided in the first aspect above.

[0013] Fourthly, embodiments of the present invention provide a computer-readable storage medium storing a computer program, which, when executed by a processor, performs the steps of the method provided in any of the first aspects above.

[0014] The embodiments of the present invention bring the following beneficial effects: The resume processing method and system based on multimodal semantic matching provided by this invention first obtains the candidate's resume document and processes it to obtain a candidate profile; then, it performs semantic parsing on the job description of the target position to construct a job profile; next, it performs semantic matching on the candidate profile and job profile based on a dual-tower deep neural network model to obtain a person-job matching score between the resume document and the target position, and generates a candidate recommendation report based on a large language model. In this method, the candidate's resume document is parsed to construct a candidate profile, and the job description is semantically parsed to construct a job profile. Then, a dual-tower deep neural network model is used to perform semantic matching on the candidate profile and job profile to obtain a person-job matching score. Based on the person-job matching score, the candidate's resume is recommended to HR, and a candidate recommendation report is generated using a large language model, thereby improving the accuracy of resume parsing and enabling more precise person-job matching. Simultaneously, this method can automatically screen resumes, allowing HR to focus only on the screened resumes, thus improving the efficiency of resume screening.

[0015] Other features and advantages of the invention will be set forth in the description which follows, and will be apparent in part from the description, or may be learned by practicing the invention. The objects and other advantages of the invention are realized and obtained in accordance with the structures particularly pointed out in the description, claims and drawings.

[0016] To make the above-mentioned objects, features and advantages of the present invention more apparent and understandable, preferred embodiments are described below in detail with reference to the accompanying drawings. Attached Figure Description

[0017] To more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the drawings used in the description of the specific embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present invention. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.

[0018] Figure 1 A flowchart illustrating a resume processing method based on multimodal semantic matching provided in an embodiment of the present invention; Figure 2 This is a flowchart illustrating a multimodal resume parsing method provided in an embodiment of the present invention. Figure 3 A schematic diagram of a matching process based on a dual-tower model provided in an embodiment of the present invention; Figure 4A flowchart of model back optimization provided in an embodiment of the present invention; Figure 5 This is a schematic diagram of the structure of a resume processing system based on multimodal semantic matching provided in an embodiment of the present invention; Figure 6 This is a schematic diagram of the structure of an electronic device provided in an embodiment of the present invention. Detailed Implementation

[0019] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0020] Currently, existing resume parsing and matching technologies mainly suffer from the following problems: (1) Weak parsing ability: Traditional OCR cannot handle resume documents with complex layouts (such as two columns), which will cause text to be out of order (reading the left column and then the right column becomes reading one line), and it cannot parse the content in the image.

[0021] (2) Poor semantic understanding: The keyword matching method cannot understand synonyms and hyponyms, which will lead to misunderstandings for resumes with piled-up keywords.

[0022] (3) Lack of multimodal analysis: For positions such as designers and models, the existing system directly ignores the portfolio images in the resume, thus missing the most core evaluation dimensions.

[0023] (4) Poor interpretability: AI scoring systems usually only give a score, and HR cannot know the specific reasons for the score, thus distrusting the algorithm results.

[0024] (5) Static model: The matching rules are fixed and cannot be dynamically adjusted according to the actual recruitment feedback of the enterprise (such as the preference of a certain department).

[0025] Based on this, the present invention provides a resume processing method and system based on multimodal semantic matching, which can improve the efficiency of resume screening and the accuracy of resume parsing, thereby making the matching of people and jobs more accurate.

[0026] To facilitate understanding of this embodiment, a resume processing method based on multimodal semantic matching disclosed in this invention will first be described in detail. This method can be executed by electronic devices, such as computers, smartphones, and tablets. See also... Figure 1The flowchart shown illustrates a resume processing method based on multimodal semantic matching, indicating that the method mainly includes the following steps S101 to S103: Step S101: Obtain the candidate's resume document and process the resume document to obtain the candidate profile.

[0027] In one implementation, resume documents in various formats such as PDF, Word, and images can be collected from multiple channels, including recruitment websites, corporate referral systems, and headhunter emails. Then, a document layout analysis model based on deep learning (such as LayoutLMv3) can be used to identify layout elements such as text blocks, tables, images (avatars / portfolios), and heading levels in the document, transforming the unstructured document into a structured intermediate representation that includes position coordinates and logical relationships.

[0028] Furthermore, the resume documents undergo multimodal information extraction and structural reconstruction, extracting fields such as candidate name, contact information, education, work experience, project experience, and skill tags. The professionalism of the candidate's avatar is analyzed, or the stylistic features of the designer's portfolio are extracted. Finally, the multimodal information is integrated to generate a standardized 360-degree candidate profile.

[0029] Step S102: Perform semantic parsing on the job description of the target position to construct a job profile.

[0030] In one implementation, deep semantic parsing is performed on the job description (JD) of the target position to extract hard requirements (education, experience) and soft skills (communication skills, stress resistance, etc.) to construct a job profile. This profile is then combined with an industry knowledge graph and extended using a graph neural network (GNN). For example, when the JD requires familiarity with PyTorch, the industry knowledge graph can infer that candidates with TensorFlow experience are also highly relevant, thus solving the problem of missed screening due to keyword mismatch.

[0031] Step S103: Based on the dual-tower deep neural network model, perform semantic matching on the candidate profile and the job profile to obtain the person-job matching score between the resume document and the target job, and generate a recommendation report for the candidate based on the large language model.

[0032] In one implementation, a Two-Tower Deep Neural Network (Two-TowerDSSM) model is constructed for both the resume and job description sides. Candidate profiles and job descriptions are mapped to high-dimensional dense vectors, and coarse-ranking recall is performed by calculating the cosine similarity between the two vectors. Subsequently, a Cross-Encoder model is used to refine the recall results, considering the cross-influence between features (e.g., a master's degree has a higher weight in the algorithm's job descriptions), and a candidate-job fit score is calculated.

[0033] In this embodiment of the invention, a large language model (LLM) can also be used to generate natural language recommendation reasons, i.e., a recommendation report. For example: Recommendation reason: This candidate has 5 years of experience at a competitor company, and their skill set is highly complementary to the team. Simultaneously, an interactive screening interface is provided to HR, who can like / dislike the recommendation results or input feedback, such as: preferring candidates with management experience. The system can adjust the sorting logic in real time based on HR feedback.

[0034] The resume processing method based on multimodal semantic matching provided in this invention first parses the candidate's resume document to construct a candidate profile and performs semantic parsing on the job description to construct a job profile. Then, it uses a dual-tower deep neural network model to perform semantic matching between the candidate profile and the job profile to obtain a person-job matching score. Based on the person-job matching score, it recommends the candidate's resume to HR and uses a large language model to generate a candidate recommendation report, thereby improving the accuracy of resume parsing and making person-job matching more precise. At the same time, the above method can automatically screen resumes, and HR only needs to focus on the screened resumes, thereby improving the efficiency of resume screening.

[0035] In one implementation, for the preceding step S101, i.e., when processing the resume document to obtain the candidate profile, the following methods may be used, including but not limited to: First, a document layout analysis model based on deep learning is used to identify the layout elements in the resume document and obtain the position coordinates and logical relationships of the layout elements to obtain the resume document structure; among them, the layout elements include at least: text block areas, table areas, image areas and heading levels.

[0036] In practice, after obtaining the resume document, the built-in file conversion model converts resumes in Word (doc / docx), PDF, image (jpg / png), HTML and other formats into a unified PDF intermediate format and generates a high-resolution rendering image. Then, the document layout analysis model is used to identify text block areas, table areas, image areas and heading levels in the resume document, and records the position coordinates and logical relationships of the layout elements to obtain the resume document structure.

[0037] Specifically, it identifies different logical blocks (i.e., text blocks) in the resume, such as: personal information area, education experience area, project experience area, sidebar, etc. For resume documents with column layout, the document layout analysis model can accurately determine the reading order (i.e., left column-right column, top-bottom); it identifies skill list tables or project timelines in the resume document, reconstructs the row and column relationships of the tables, and avoids content misalignment; it detects visual objects in the document such as facial photos, portfolio thumbnails, and company logos, and records their coordinates for subsequent processing; it identifies information such as title, subtitle, and body text, and constructs a document tree (DOM Tree) to ensure that the logical structure of the parsed content is consistent with the original document.

[0038] Then, based on the resume document structure, multimodal information is extracted from the resume document and fused to obtain the candidate profile.

[0039] In practice, after obtaining the document structure, the resume document undergoes in-depth information extraction and standardization, specifically including the following processes: (1) For text blocks, named entity recognition is performed based on the fine-tuned large language model to obtain entity information.

[0040] In one implementation, named entity recognition is performed using a finely tuned 7B / 13B parameter LLM (such as LayoutLMv3), including: accurately extracting basic information (name, gender, etc.), school (distinguishing between bachelor's, master's, and doctoral degrees), major, company, position, and time period (automatically calculating years of service) of candidates from unstructured paragraphs; inferring the candidate's actual job level based on project descriptions (e.g., inferring the actual job level as P7+ based on the management of the responsible team); inferring age based on graduation time (if not specified); and pre-establishing a skill thesaurus, for example, unifying JS, JavaScript, and EcmaScript into standard skill tags, using this skill thesaurus to identify skill entities in the candidate's resume document, and identifying the level of skill mastery (proficient / familiar / understanding).

[0041] (2) For the image region, extract the visual feature information of the candidate based on the fine-tuned visual model.

[0042] In one implementation, for a candidate's headshot, the Face API is used to detect the face and assess the professionalism of the headshot (whether the candidate is dressed formally, clarity, etc.) for screening service-related positions; for a candidate's portfolio, a pre-trained visual aesthetics scoring model is used to analyze the images of works attached to the resume and extract feature vectors such as color matching and composition style, which serve as auxiliary screening dimensions for UI / UX designer positions.

[0043] (3) The entity information and visual feature information are fused to obtain the candidate profile.

[0044] In one implementation, all the extracted information is assembled into a standardized JSON Profile (User Profile), which includes the candidate's basic information, educational background, work experience (timeline ordered), skills graph, project portfolio, visual features, etc., to obtain a candidate profile, which is then saved to Elasticsearch and a graph database.

[0045] For ease of understanding, this embodiment of the invention also provides a flowchart of multimodal resume parsing, see [link to flowchart]. Figure 2 As shown, the original resume file is first obtained, and the file type of the original resume file is determined. If the original resume file is in PDF or Word format, the text and corresponding coordinate information in the original resume file are extracted. If the original resume file is an image, the original resume file is subjected to OCR file detection. Then, the extracted information is rendered as an image and input into the fine-tuned LayoutLMv3 model to perform in-depth information extraction and standardization processing on the resume file.

[0046] Specifically, the LayoutLMv3 model's encoding layer encodes the textual, layout, and visual information of the resume file: text is converted into vectors through word embeddings, layout information is added to the text token representation through position embeddings, and the image is segmented into fixed-size patches (e.g., 16×16 pixels). Each patch is converted into a vector through linear projection or a small CNN, and then patch position embeddings are added. Next, LayoutLMv3 performs downstream tasks, including: semantic entity recognition, identifying entities such as name, school, and company in the resume file; relation extraction, such as identifying the company-position-time relationships between candidates in the resume file; and document image classification, such as determining the resume type and language.

[0047] Finally, the extracted information is structured and assembled, and standardized based on a knowledge base to obtain the final candidate profile. For example, JS is unified into JavaScript.

[0048] In this embodiment of the invention, LayoutLMv3 can be pre-trained using large-scale unlabeled document data through the following self-supervised tasks: (1) Masked language modeling: Randomly mask part of the text token and let the model predict the masked words based on the context and layout / image information.

[0049] (2) Masked image modeling: Randomly mask part of the image patch and reconstruct the content of the masked patch using the remaining image, text and layout information.

[0050] (3) Text-image alignment: Learn the semantic consistency between text tokens and their corresponding image regions.

[0051] After pre-training, LayoutLMv3 can be fine-tuned for specific tasks, such as: document classification: labeling entire pages of documents; sequence labeling (e.g., named entity recognition): identifying whether each token belongs to a certain field (e.g., date); table structure recognition: understanding the row and column relationships in a table. During fine-tuning, task-specific classification headers (e.g., fully connected layers) can be added above the Transformer output layer, and end-to-end training can be performed using labeled data.

[0052] Considering that job descriptions are often vague and direct matching is ineffective, it is necessary to process the job descriptions first and construct a job profile. Specifically, for step S102 mentioned above, that is, when performing semantic parsing on the job description of the target position to construct a job profile, the following methods can be used, including but not limited to: First, semantic parsing is performed on the job description of the target position to extract the job requirements, and an initial job profile is constructed based on the job requirements.

[0053] In practice, LLM is used to semantically analyze job descriptions and extract job requirements, including hard requirements (education > bachelor's degree, experience > 5 years, location = Beijing) and soft requirements (stress resistance, teamwork, etc.). Simultaneously, bonus and mandatory items in the job description are identified, assigned different weights, and an initial job profile is constructed based on the extracted job requirements.

[0054] Then, based on the preset industry knowledge graph, the initial job profile is expanded to obtain the job profile.

[0055] In practical implementation, a heterogeneous industry knowledge graph encompassing job positions, skills, companies, schools, and industries is constructed. This industry knowledge graph, combined with the initial job profile, is used for reasoning and expansion to obtain the final job profile. Furthermore, the industry knowledge graph can also be used to reason about and expand candidate profiles. Specifically, industry knowledge graph reasoning includes: (1) Skill association reasoning: Define the association between skills in the industry knowledge graph. For example, PyTorch similar to TensorFlow, Spring Boot part of Java Ecosystem. When the job description requires PyTorch, but the resume only contains TensorFlow, combine the industry knowledge graph, use graph algorithms to calculate the semantic distance between the two, and determine that the two are highly related based on the semantic distance, thereby avoiding omissions.

[0056] (2) Company background reasoning: Industry knowledge graphs contain information such as competitive / cooperative relationships between companies and industry attributes. Using industry knowledge graphs, it is possible to identify whether a candidate comes from a competitor company or from an upstream or downstream company, and to give bonus points accordingly.

[0057] (3) School-level reasoning: Based on knowledge bases such as QS rankings and Double First-Class University lists, the system automatically scores candidates’ academic qualifications in different levels, rather than simply matching school names, thereby improving the accuracy of matching.

[0058] (4) Alternative reasoning: When hard conditions are not met (such as academic qualifications), look for candidates with high-value alternative skills (such as winning an ACM gold medal).

[0059] In this embodiment of the invention, through graph reasoning, the job description can be transformed into a job requirement subgraph containing extended semantics, and the resume can be transformed into a talent capability subgraph.

[0060] In one implementation, for the aforementioned step S103, i.e., when performing semantic matching of candidate profiles and job profiles based on a dual-tower deep neural network model to obtain the person-job matching score between the resume document and the target job, the following methods may be adopted, including but not limited to: First, the candidate profile and job profile are input into a dual-tower deep neural network model for vectorization processing to obtain candidate profile vectors and job profile vectors.

[0061] Then, the cosine similarity between the candidate profile vector and the job profile vector is calculated, and based on the cosine similarity, a first preset number of candidate resumes is determined.

[0062] In practical implementation, the dual-tower deep neural network model includes a job tower and a resume tower. The job tower uses a pre-trained language model (such as BERT) to encode job profiles, converting them into fixed-length vector representations. The resume tower uses another similar pre-trained language model to encode candidate profiles, also converting them into fixed-length vector representations. In this embodiment, historical resume submission records (submitted = positive sample, unsubmitted = negative sample) can be used for comparative learning training. The job tower and resume tower are trained in parallel, each focusing on feature learning for either the job or the candidate.

[0063] In this embodiment of the invention, candidate profiles and job profiles are input into a resume database and a job database, respectively, for vectorization processing to obtain candidate profile vectors and job profile vectors. The candidate profile vectors are then stored in a vector database. During real-time matching, the job profile vectors are used to search the vector database, and the cosine similarity between the candidate profile vectors and job profile vectors is calculated. Based on the cosine similarity, the Top-N (e.g., 1000) resumes are selected as candidate resumes. The dual-tower deep neural network model provides fast resume coarse-sorting and recall speed and can improve the problem of word mismatch (semantic matching).

[0064] Finally, based on the pre-set fine-tuning model, the job description and candidate resumes are interacted to obtain the matching results between the candidate resume and the target position and the person-job matching score.

[0065] In practice, deep interactive computation is performed on the top-N resumes retrieved, and the job description and candidate resume are input into the fine-ranking model for interaction. In this embodiment of the invention, the fine-ranking model can adopt a Cross-Encoder model (such as BERT for sentence pair classification). This model enables deep attention interaction between JD features and resume features at the bottom layer of the model, while adding manually defined features (e.g., whether the education level matches, whether the years of experience meet the requirements, job-hopping frequency, etc.), and finally outputs a person-job matching score (0-100 points).

[0066] The fine-ranking model in this embodiment of the invention can capture subtle semantic differences, such as distinguishing between an algorithm engineer who has worked in sales and a salesperson who has worked in algorithms, thereby improving the accuracy of person-job matching.

[0067] For ease of understanding, this embodiment of the invention also provides a schematic diagram of a matching process based on a dual-tower model, see [link / reference]. Figure 3 As shown, firstly, candidate profiles and job profiles are input into the resume pyramid and job pyramid respectively, resulting in candidate profile vectors and resume profile vectors. Then, based on the candidate profile vectors and resume profile vectors, an ANN (Approximate Nearest Neighbor) retrieval is performed to recall the top-1000 candidate resumes. Next, a Cross-Encoder model is used to perform Full Self-Attention interaction on the job description and candidate resumes to obtain the matching results between the candidate resumes and the target job and the person-job matching score. Finally, based on preset business rules, the recalled candidates are re-ranked to obtain the top-50 candidate resumes to generate a recommendation list, and a large language model is used to generate recommendation reasons.

[0068] In one implementation, for the aforementioned step S103, i.e., when generating a candidate recommendation report based on a large language model, the following methods may be used, including but not limited to: First, based on the Big Prophecy model and matching results, a recommendation report for the candidate is generated. In practice, the text generation capabilities of LLM are utilized to generate a recommendation summary (i.e., the recommendation report) based on the matching results, including: highlighting key strengths, such as: this candidate is not only proficient in the Java skills required by the JD, but also possesses cloud-native experience that is not mentioned in the JD but is urgently needed by the team (a plus); risk warnings, such as: note that this candidate has changed jobs 3 times in the past 3 years, and their stability may be questionable; and interview suggestions, such as: it is recommended to focus on assessing their system design capabilities in high-concurrency scenarios during the interview.

[0069] Then, a recommendation list is generated based on the job-person matching score and the recommendation report. In practice, the job-person matching score and the recommendation report are displayed as a recommendation list to recruiters through an interactive interface.

[0070] Finally, the recommended list is displayed to recruiters, and their feedback is obtained. In practice, HR can intervene in the recommendation list. For example, HR can click "not interested" and select a reason (e.g., low education level), and the system will update the ranking in real time, reducing the weight of candidates with similar characteristics (low education level). Alternatively, HR can click "add to favorites," and the system will recommend more similar candidates. Furthermore, HR can input the command "put those with e-commerce backgrounds first," and the system will parse the command and dynamically adjust the weight parameters of the ranking model.

[0071] In one implementation, person-job matching is not a one-time action, but a long-term optimization process. Based on this, the method further includes: acquiring tracking data of candidates at each stage of the recruitment process, and optimizing the dual-tower deep neural network model based on the tracking data.

[0072] In practice, data on candidates' performance in subsequent recruitment stages, such as interviews, offers, onboarding, and probation, can be continuously tracked. Information such as interview pass rates and onboarding performance can be used as delayed feedback signals. The parameters of the dual-tower deep neural network model can be updated in reverse through reinforcement learning algorithms, so that the model gradually evolves from matching resumes to matching high-performing talents.

[0073] In this embodiment of the invention, the ATS (Admissions Management System) and EHR (Employment Human Resources System) can be integrated to record candidate's entire lifecycle data: resume approval, interview approval, offer acceptance, onboarding, probation period approval, performance rating, etc. This embodiment sets the optimization objective of the dual-tower deep neural network model to ensure successful onboarding and high performance. Samples that are ultimately hired and perform well are assigned extremely high positive weights; samples that fail the interview are assigned negative weights. The model is regularly (e.g., monthly) optimized using the latest end-to-end data. The model gradually learns the implicit preferences of companies (e.g., a department prefers graduates from certain types of schools or a position may not require high academic qualifications but values ​​practical experience). Simultaneously, fairness constraints are introduced to monitor and eliminate discriminatory biases in the model based on gender, age, and region, ensuring recruitment compliance.

[0074] For ease of understanding, this embodiment of the invention also provides a model back-optimization flowchart, see [link / reference]. Figure 4 As shown, after generating the recommendation list, the model can be optimized based on HR behavior feedback. Specifically, HR feedback includes immediate and delayed feedback. Immediate feedback includes clicks / favorites, where HR can click "not interested" or "add to favorites," and sample pool A is generated based on the feedback results. Delayed feedback includes interviews / onboarding, where data on candidates' performance in subsequent recruitment stages such as interviews, offers, onboarding, and probation is continuously tracked, and information such as interview pass rates and onboarding performance is used as delayed feedback signals, generating sample pool B based on the delayed feedback information. Furthermore, the samples are weighted according to interview results and performance, for example: passing the interview increases the weight by 1; excellent performance increases the weight by 5; and failing the interview decreases the weight by 1. Then, using resume samples from sample pools A and B, the model is trained, the model's loss value is calculated, and the parameters of the job pyramid and Cross-Encoder are updated.

[0075] The resume processing method based on multimodal semantic matching provided in this invention employs the LayoutLMv3 model, which integrates text, layout, and image information for pre-training. In the resume parsing task, the model not only reads text but also views images and layouts, thereby accurately identifying the reading order of columns in the resume, distinguishing heading levels by font size, and recognizing table boundaries by alignment, thus improving the parsing accuracy for complex design resumes (such as two-column or mixed text and image layouts).

[0076] An industry knowledge graph was constructed, encompassing a standard skills tree, enterprise graph, university database, and job system. A graph reasoning mechanism was introduced during the matching process: hierarchical reasoning: if a job description requires backend development, the graph automatically matches lower-level terms such as Java development and Go development; associative reasoning: if a job description requires familiarity with distributed systems, the graph can infer that candidates with Dubbo and Spring Cloud skills meet the requirements; alternative reasoning: when hard requirements are not met, the graph can find candidates with high-value alternative skills. This gives the system reasoning capabilities and improves the problem of missed screening caused by hard keyword matching.

[0077] The dual-tower DSSM model is used to encode resumes and job descriptions into vectors independently, and a vector retrieval engine (ANN) is used to achieve millisecond-level retrieval of hundreds of millions of resumes. The Cross-Encoder model is used to perform deep full-attention interaction on the small number of resumes recalled to capture fine-grained semantic matching features, thereby improving the retrieval speed and accuracy of large-scale resume databases.

[0078] By leveraging the generative capabilities of large language models, white-box processing is achieved. This not only outputs job-person matching scores but also automatically generates recommendation summaries (summarizing candidates' strengths using natural language), gap analysis (clearly pointing out candidates' weaknesses, such as the lack of mention of English proficiency), and interview question generation (automatically generating a customized interview question bank based on doubts in the candidate's resume (such as project gaps) and the core requirements of the JD to assist interviewers in asking questions). This improves the professionalism of the interview, provides candidates with more accurate job recommendations, and enhances the job application experience.

[0079] By using an importance sampling algorithm, subsequent interview evaluations and onboarding performance are used as real feedback signals to backtrack and correct the parameters of the early matching model, enabling the model to overcome time delays and improve its accuracy and interpretability.

[0080] In addition to the resume processing method based on multimodal semantic matching provided in the foregoing embodiments, this invention also provides a resume processing system based on multimodal semantic matching, see [link to relevant documentation]. Figure 5 The diagram shows a structural schematic of a resume processing system based on multimodal semantic matching. The device may include the following parts: The candidate profile building module 501 is used to obtain the candidate's resume document and process the resume document to obtain the candidate profile.

[0081] The job profile building module 502 is used to perform semantic parsing on the job description of the target job and build a job profile.

[0082] The person-job matching module 503 is used to perform semantic matching of candidate profiles and job profiles based on a dual-tower deep neural network model, obtain the person-job matching score between the resume document and the target job, and generate a recommendation report for candidates based on a large language model.

[0083] The resume processing system based on multimodal semantic matching provided in this invention first parses the candidate's resume document to construct a candidate profile and performs semantic parsing on the job description to construct a job profile. Then, it uses a dual-tower deep neural network model to perform semantic matching between the candidate profile and the job profile to obtain a person-job matching score. Based on the person-job matching score, it recommends the candidate's resume to HR and uses a large language model to generate a candidate recommendation report, thereby improving the accuracy of resume parsing and enabling more precise person-job matching. At the same time, the above method can automatically screen resumes, and HR only needs to focus on the screened resumes, thereby improving the efficiency of resume screening.

[0084] The device provided in this embodiment of the invention has the same implementation principle and technical effect as the aforementioned method embodiment. For the sake of brevity, any parts not mentioned in the device embodiment can be referred to the corresponding content in the aforementioned method embodiment.

[0085] This invention also provides an electronic device, specifically, the electronic device includes a processor and a storage device; the storage device stores a computer program, and the computer program, when run by the processor, executes the method described in any of the above embodiments.

[0086] Figure 6 This is a schematic diagram of the structure of an electronic device provided in an embodiment of the present invention. The electronic device 100 includes: a processor 60, a memory 61, a bus 62, and a communication interface 63. The processor 60, the communication interface 63, and the memory 61 are connected through the bus 62. The processor 60 is used to execute executable modules, such as computer programs, stored in the memory 61.

[0087] The memory 61 may include high-speed random access memory (RAM) or non-volatile memory, such as at least one disk storage device. Communication between this system network element and at least one other network element is achieved through at least one communication interface 63 (which can be wired or wireless), such as the Internet, wide area network, local area network, metropolitan area network, etc.

[0088] Bus 62 can be an ISA bus, PCI bus, or EISA bus, etc. The bus can be divided into address bus, data bus, control bus, etc. For ease of representation, Figure 6The symbol is represented by a single double-headed arrow, but this does not mean that there is only one bus or one type of bus.

[0089] The memory 61 is used to store programs. After receiving an execution instruction, the processor 60 executes the program. The method executed by the device for defining the flow process disclosed in any of the foregoing embodiments of the present invention can be applied to the processor 60 or implemented by the processor 60.

[0090] Processor 60 may be an integrated circuit chip with signal processing capabilities. In implementation, each step of the above method can be completed by the integrated logic circuitry in the hardware of processor 60 or by instructions in software form. Processor 60 can be a general-purpose processor, including a Central Processing Unit (CPU), a Network Processor (NP), etc.; it can also be a Digital Signal Processor (DSP), an Application Specific Integrated Circuit (ASIC), a Field-Programmable Gate Array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. It can implement or execute the methods, steps, and logic block diagrams disclosed in the embodiments of this invention. The general-purpose processor can be a microprocessor or any conventional processor. The steps of the methods disclosed in the embodiments of this invention can be directly embodied in the execution of a hardware decoding processor, or executed by a combination of hardware and software modules in the decoding processor. The software modules can reside in random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, registers, or other mature storage media in the art. The storage medium is located in memory 61. Processor 60 reads the information in memory 61 and, in conjunction with its hardware, completes the steps of the above method.

[0091] The computer program product of the readable storage medium provided in the embodiments of the present invention includes a computer-readable storage medium storing program code. The instructions included in the program code can be used to execute the methods described in the foregoing method embodiments. For specific implementation, please refer to the foregoing method embodiments, which will not be repeated here.

[0092] If the aforementioned functions are implemented as software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this invention, essentially, or the part that contributes to the prior art, or a portion of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0093] Finally, it should be noted that the above-described embodiments are merely specific implementations of the present invention, used to illustrate the technical solutions of the present invention, and not to limit it. The scope of protection of the present invention is not limited thereto. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that any person skilled in the art can still modify or easily conceive of changes to the technical solutions described in the foregoing embodiments within the technical scope disclosed in the present invention, or make equivalent substitutions for some of the technical features; and these modifications, changes, or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be covered within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the scope of the claims.

Claims

1. A resume processing method based on multimodal semantic matching, characterized in that, include: Obtain the candidate's resume document and process the resume document to obtain the candidate profile; Semantic parsing is performed on the job description of the target position to construct a job profile; Based on a dual-tower deep neural network model, semantic matching is performed on the candidate profile and the job profile to obtain the person-job matching score between the resume document and the target job, and a recommendation report for the candidate is generated based on a large language model.

2. The method according to claim 1, characterized in that, The candidate profile is obtained by processing the resume document, including: A document layout analysis model based on deep learning is used to identify the layout elements in the resume document and obtain the position coordinates and logical relationships of the layout elements to obtain the resume document structure; wherein, the layout elements include at least: text block areas, table areas, image areas and heading levels; Based on the resume document structure, multimodal information is extracted from the resume document, and the multimodal information is fused to obtain a candidate profile.

3. The method according to claim 2, characterized in that, Based on the resume document structure, multimodal information is extracted from the resume document, and the multimodal information is fused to obtain a candidate profile, including: For the text block, named entity recognition is performed based on the fine-tuned large language model to obtain entity information; For the image region, visual feature information of the candidate is extracted based on the fine-tuned visual model; The entity information and the visual feature information are fused to obtain a candidate profile.

4. The method according to claim 1, characterized in that, Semantic parsing is performed on the job description of the target position to construct a job profile, including: Semantic parsing is performed on the job description of the target position to extract the job requirements, and an initial job profile is constructed based on the job requirements; Based on a pre-defined industry knowledge graph, the initial job profile is expanded to obtain a new job profile.

5. The method according to claim 1, characterized in that, Based on a dual-tower deep neural network model, semantic matching is performed on the candidate profile and the job profile to obtain a person-job fit score between the resume document and the target job, including: The candidate profile and the job profile are input into a dual-tower deep neural network model for vectorization processing to obtain the candidate profile vector and the job profile vector. Calculate the cosine similarity between the candidate profile vector and the job profile vector, and determine a first preset number of candidate resumes based on the cosine similarity. Based on a pre-defined ranking model, the job description and the candidate's resume are interacted to obtain the matching result between the candidate's resume and the target position, as well as the person-job matching score.

6. The method according to claim 5, characterized in that, A recommendation report for the candidates is generated based on a large language model, including: Based on the big oracle model and the matching results, a recommendation report for the candidate is generated; A recommendation list is generated based on the job-person matching score and the recommendation report; The recommended list is displayed to recruiters, and feedback is obtained from them.

7. The method according to claim 1, characterized in that, Also includes: The candidate's tracking data at each stage of the recruitment process is obtained, and the dual-tower deep neural network model is optimized based on the tracking data.

8. A resume processing system based on multimodal semantic matching, characterized in that, include: The candidate profile building module is used to obtain the candidate's resume document and process the resume document to obtain the candidate profile. The job profile building module is used to perform semantic parsing on the job description of the target job and build a job profile. The person-job matching module is used to perform semantic matching between the candidate profile and the job profile based on a dual-tower deep neural network model, obtain the person-job matching score between the resume document and the target job, and generate a recommendation report for the candidate based on a large language model.

9. An electronic device, characterized in that, The method includes a processor and a memory, the memory storing computer-executable instructions executable by the processor, the processor executing the computer-executable instructions to implement the steps of the method according to any one of claims 1 to 7.

10. A computer-readable storage medium storing a computer program thereon, characterized in that, The computer program is executed by the processor to perform the steps of the method described in any one of claims 1 to 7.