Talent feature extraction and matching method based on vertical field large model

Through the talent feature extraction and matching method based on the vertical field big model, the problems of low efficiency and insufficient accuracy of talent matching in the vertical field are solved, and efficient and accurate talent acquisition and decision-making transparency are achieved, and are suitable for specialized fields such as finance and medical care.

CN120338738APending Publication Date: 2025-07-18SUZHOU RENLIAN TIMES TECHNOLOGY CO LTD
View PDF 0 Cites 5 Cited by

Patent Information

Application Number
CN202510423298.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-07
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

The existing talent matching technology has problems such as low matching efficiency, high cost, difficulty in identifying implicit capabilities, lack of semantic understanding, and evaluation deviation in vertical fields, resulting in inaccurate matching and unable to meet the enterprise's needs for efficient and accurate talent acquisition.

Method used

A talent feature extraction and matching method based on a vertical field big model is adopted. Through structured text segmentation, explicit and implicit ability feature extraction, entity linking, mixed search strategies and reinforced learning optimization, high-dimensional job requirements and professional ability portraits are generated, and model parameters are optimized in combination with user feedback.

Benefits of technology

It has achieved in-depth analysis and accurate matching of talent characteristics, improved the accuracy of job matching and decision-making transparency, had a higher level of intelligence and industry adaptability, and continuously improved the quality of talent recommendations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120338738A_ABST
    Figure CN120338738A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of evaluation management, in particular to a vertical domain large model-based talent feature extraction and matching method, which comprises the following steps of: performing standardized analysis and structured processing on enterprise recruitment demand information and candidate occupational backgrounds, extracting explicit qualification features by using a vertical domain large model and generating implicit ability features; performing background enhancement on the feature data by associating an industry knowledge graph, and constructing a post demand portrait and a vocational ability portrait; a mixed retrieval strategy is adopted to fuse hard standard screening and soft capacity matching, and a preliminary screening candidate set is generated; carrying out man-post adaptation degree actuarial calculation through an SFT-LLM model, and outputting a visual analysis report containing key matching element comparison and risk early warning; and finally, based on the actual recording feedback data of the recruiters, continuously improving the talent recommendation accuracy by dynamically adjusting model parameters, and providing efficient and reliable talent selection decision support for the specialized fields of finance, medical treatment and the like.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of evaluation management, and specifically to a method for extracting and matching talent characteristics based on a large model in a vertical field. Background Art

[0002] In enterprise management and market competition, efficient and accurate talent matching is crucial for the success of an enterprise organization. However, existing talent or expert matching technologies still have certain limitations.

[0003] Traditional manual matching relies on the experience and network of recruitment specialists, headhunting consultants, or business connectors, resulting in low matching efficiency, high costs, and difficulty in scaling up. When faced with a large number of candidates or cross-domain requirements, the matching speed is slow, the coverage is limited, and it is easily affected by subjective biases. Especially in the scenario of expert matching that requires "private domain knowledge" in a specific field, the difficulty and time cost of manual evaluation are even higher. In addition, early enterprise human resource management systems mostly adopted keyword-based retrieval methods to search for resumes or talent pools through job requirements or skill keywords. However, this method overly relies on literal matching and lacks an understanding of semantics, context, and skill relevance, resulting in difficulty in accurately identifying synonyms, abbreviations, and different expression methods, easily missing suitable candidates or generating a large number of irrelevant matches, with relatively low matching accuracy and affecting the enterprise's talent acquisition efficiency.

[0004] In recent years, intelligent recruitment decision-making systems have been gradually applied to tasks such as resume parsing and job description generation, but directly applying them to accurate talent matching still faces challenges. Their basic data sources are extensive, but they have insufficient understanding of professional terms, sub-skills, and implicit experience in vertical industries such as finance, biomedicine, and semiconductors, resulting in limited entity relationship mining capabilities of the knowledge graph and difficulty in accurately quantifying the mapping relationship between skill relevance and job competency. In addition, these existing intelligent recruitment systems are not optimized for the "demand - talent" dual-end matching, prone to inaccurate matching problems, and may have evaluation bias phenomena, affecting the reliability of decision-making. At the same time, the effective utilization of private data such as enterprise historical recruitment data and successful matching cases is still a difficult point, restricting the improvement of the matching capabilities of these intelligent recruitment systems in specific scenarios.

[0005] Therefore, there are certain deficiencies in the existing technology in terms of the depth of talent characteristic extraction, the domain adaptability of the knowledge graph, the professionalism in dealing with vertical field requirements, and the automation efficiency, and it cannot meet the current market demand for efficient and accurate talent matching. For this reason, a method for extracting and matching talent characteristics based on a large model in a vertical field is proposed. Summary of the Invention

[0006] The purpose of the present invention is to provide a method for talent feature extraction and matching based on a vertical domain large model, aiming to improve the accuracy of talent matching and the transparency of decision-making. By performing structured segmentation and semantic analysis on enterprise demand texts and professional background texts, using the vertical domain large model to extract explicit qualification features and infer and generate implicit ability features with confidence scores; combining with an industry knowledge graph for entity association and context enhancement to construct multi-dimensional job demand portraits and professional ability portraits; adopting a hybrid retrieval strategy to fuse structured filtering and semantic similarity matching to generate a set of pre-screened candidates; performing deep fine ranking through the SFT-LLM model optimized by tasks, outputting a quantitative adaptation score and an interpretable matching report; finally, constructing a reinforcement learning reward mechanism by capturing user interaction feedback to realize dynamic optimization of model parameters, improving the accuracy of talent matching and the transparency of decision-making.

[0007] To achieve the above object, the present invention provides the following technical solutions:

[0008] A method for talent feature extraction and matching based on a vertical domain large model, comprising:

[0009] Perform text preprocessing and block segmentation on the input enterprise demands and professional backgrounds to obtain structured texts; use the vertical domain large model to extract explicit qualification features, and perform context reasoning on the structured texts to generate implicit ability features with confidence scores; connect the explicit qualification features and the implicit ability features to a talent knowledge graph for entity linking and querying to generate a job demand portrait and a professional ability portrait;

[0010] Based on the job demand portrait, adopt a hybrid retrieval strategy to generate a set of candidate portraits from the professional ability portrait; input the job demand portrait and the set of candidate portraits into the SFT-LLM model, output the job adaptability, and output a human-job matching analysis report based on the attention mechanism to obtain a candidate list;

[0011] Calculate the preference feedback of the user on the candidate list through an interaction interface, and map the preference feedback to a reward signal based on direct rules; optimize the internal parameters of the SFT-LLM model according to the reward signal using a reinforcement learning algorithm.

[0012] Further, the extraction process of the explicit qualification features includes:

[0013] Define a target feature list for clarifying the explicit feature types and explicit feature formats;

[0014] Design dedicated prompt templates for different explicit feature types to instruct the vertical domain large model to extract the explicit qualification features; among them, the skill features are extracted using an enumeration prompt, and the experience features are extracted using a relation extraction prompt;

[0015] Input the dedicated prompt template and the structured text into the vertical domain large model to obtain the explicit qualification features;

[0016] Align the terms of the explicit qualification features with the vertical domain ontology library, and trigger an artificial term review queue for terms not included.

[0017] Further, the extraction process of the implicit ability features includes:

[0018] Use the vertical domain large model to perform semantic reasoning on the structured text to generate local implicit ability features and preliminary confidence scores;

[0019] According to the preliminary confidence scores, use the attention mechanism to adjust the weights of the local implicit ability features, combine the time series information between the structured texts, and use the time decay function to calculate the final aggregation weights to generate the implicit ability features with the confidence scores.

[0020] Further, the process of entity linking and querying includes:

[0021] Screen out the entity list from the explicit qualification features; for each entity string in the entity list, call the entity linking API provided by the talent knowledge graph to return the matching graph nodes and graph confidence;

[0022] Select the best graph nodes according to the matching graph nodes and the graph confidence, perform graph queries on the best graph nodes to obtain graph association information; supplement the graph association information to the corresponding explicit qualification features, and combine the implicit ability features to generate the job requirement portrait and the professional ability portrait.

[0023] Further, the implementation process of the hybrid retrieval strategy includes:

[0024] Analyze the hard constraint conditions for the job requirement portrait, and extract the semantic core to generate a demand vector representation;

[0025] According to the hard constraint conditions, query the professional ability portrait to obtain a preliminary subset of candidate portraits;

[0026] Obtain the professional vector representation of the preliminary subset of candidate portraits, perform KNN search, calculate the similarity score between the demand vector representation and the professional vector representation, and screen out the candidate ID list;

[0027] Output the candidate portrait set according to the candidate ID list.

[0028] Further, the generation process of the candidate list includes:

[0029] Encode the job requirement portrait and the candidate portrait set into a multi-modal input structure adapted to the SFT-LLM model, including Pointwise input, Pairwise input, and Listwise input; for the Pointwise input, directly output the job suitability; for the Pairwise input and the Listwise input, output the job suitability of the relative ranking.

[0030] Analyze the internal attention weight distribution of the SFT-LLM model, extract positive matching evidence and mismatching points, and generate the human-job matching analysis report; rank the candidates according to the job suitability and store them in association with the human-job matching analysis report to obtain the candidate list.

[0031] Furthermore, the calculation process of the reward signal includes:

[0032] Collect the user's sorting adjustment operations, explicit score inputs, and employment status marks for candidates from the interaction interface using the event-driven mechanism, and quantify them into preference interaction data.

[0033] Perform single-feedback mapping on the preference interaction data one by one to generate sorting adjustment rewards, explicit score rewards, and employment status rewards.

[0034] Aggregate the sorting adjustment rewards, the explicit score rewards, and the employment status rewards through multi-feedback, and adjust them according to the demand priority and candidate scarcity to obtain the reward signal.

[0035] Compared with the prior art, the beneficial effects of the present invention are:

[0036] 1. The present invention integrates the large model in the vertical field, the knowledge graph, and the multi-modal feature extraction technology to deeply analyze the job requirements and the candidate's professional background. Through multi-task learning, explicit qualification features and implicit ability features are extracted, and entity linking is performed in combination with the knowledge graph to construct high-dimensional job requirement portraits and professional ability portraits. Compared with traditional keyword matching, the present invention realizes more accurate and comprehensive talent feature modeling, improving the accuracy of human-job matching.

[0037] 2. The present invention adopts a hybrid retrieval strategy of hard screening, vector semantic search, and keyword weighted matching, combines KNN nearest neighbor search to calculate the matching degree between the job requirements and the candidate portrait, and extracts matching evidence based on the attention mechanism to generate a human-job matching analysis report. The candidate ranking adopts three matching strategies of Pointwise, Pairwise, and Listwise, and dynamically calculates the suitability, which can more accurately identify high-quality candidates and improve the job suitability compared with fixed rule screening.

[0038] 3. Collect data such as the screening behavior and employment decisions of enterprise HR through an event-driven feedback mechanism, quantify it as a reward signal, and dynamically optimize the matching model parameters in combination with reinforcement learning. At the same time, use the optimized matching model parameters to identify the key factors affecting the matching quality, continuously optimize the matching strategy, and achieve the adaptive optimization of the talent matching result. Compared with the static matching algorithm, the present invention has a higher level of intelligence and industry adaptability, and can continuously improve the quality of talent recommendation. BRIEF DESCRIPTION OF THE DRAWINGS

[0039] Figure 1 It is a schematic flowchart of the method for extracting and matching talent characteristics based on a large model in a vertical field provided by the present invention;

[0040] Figure 2 It is a schematic flowchart of the process for generating a job requirement portrait and a professional ability portrait of the present invention;

[0041] Figure 3 It is a schematic flowchart of the calculation process of the reward signal of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0042] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative work shall fall within the protection scope of the present invention.

[0043] Please refer to Figures 1 to 3 , the present invention provides a method for extracting and matching talent characteristics based on a large model in a vertical field, and the technical solution is as follows:

[0044] Embodiment 1:

[0045] Talent characteristic extraction and matching is the core link of modern human resource management and has important applications in enterprise recruitment, career recommendation, and team optimization. However, traditional talent matching methods mainly rely on keyword extraction in resumes and manual experience judgment, making it difficult to comprehensively capture the implicit ability characteristics of talents, such as leadership and adaptability, and it is also difficult to meet diverse job requirements. In addition, existing feature extraction technologies do not have a deep enough understanding of the semantics of text and lack analysis of industry background and context, resulting in insufficient accuracy of matching results and affecting the recruitment efficiency and scientific nature of talent allocation. Therefore, it is necessary to develop a method for extracting and matching talent characteristics based on a large model in a vertical field to optimize feature extraction, improve matching accuracy, and ensure the efficiency and consistency of person-job matching.

[0046] To improve the depth of talent characteristic extraction and the accuracy of matching results, asFigure 1 As shown, a talent feature extraction and matching method based on a large model in a vertical domain is used, including:

[0047] Step 1: Perform text preprocessing and block segmentation on the input job background and career background to obtain structured text.

[0048] Among them, the job background refers to the job requirements released by the enterprise, usually the job description in the recruitment information. The career background refers to the career information of the job seeker, usually from the resume or personal file.

[0049] The text preprocessing process is as follows: First, perform format conversion by calling libraries such as Apache POI or pdfminer.six to convert the file content into plain text format. Then perform text cleaning, remove or replace non-standard characters and special formatting characters, and uniformly replace redundant spaces, tab characters, and blank lines with standard formats. According to specific requirements, unify or retain the case of specific words. In addition, based on a predefined dictionary of common technical terms, correct obvious spelling mistakes, such as correcting "MySql" to "MySQL".

[0050] The implementation process of block segmentation is as follows: First, use a rule-based splitter to scan the cleaned text to identify common resume section headings (such as "Personal Information", "Educational Background", "Work Experience", and "Professional Skills", etc.) and their format features (such as occupying a single line, followed by a colon, and bold font, etc.). According to the identified headings, split the text into different logical blocks and associate corresponding type tags with each block. Finally, output a structured data representation for subsequent processing and analysis.

[0051] Through the structured text, the enterprise requirements are decomposed into clear modules of job requirements and candidate expectations, and the career background is split into independent units of candidate experiences and capabilities. This enables the subsequent feature extraction model to focus on specific types of blocks, such as "Work Experience" and "Skills" blocks, and mainly extract technical skills. This targeted processing avoids ineffective analysis in irrelevant text (such as "Personal Information"), thereby improving the speed and accuracy of feature extraction.

[0052] Step 2: Use a large model in the vertical domain to extract explicit qualification features and perform context reasoning on the structured text to generate implicit ability features with confidence scores.

[0053] Among them, explicit qualification features refer to the abilities, qualifications, or background information directly mentioned in the text, such as "proficient in Java", while implicit ability features need to be inferred, for example, inferring "leadership" from "led the team to complete the project".

[0054] Furthermore, the process of extracting the explicit qualification feature includes:

[0055] Define the target feature list to specify explicit feature types and explicit feature formats;

[0056] Designing dedicated prompt templates for different explicit feature types to instruct the vertical domain big model to extract the explicit qualification features; wherein, skill features use enumeration extraction prompts, and experience features use relational extraction prompts;

[0057] Inputting the dedicated prompt template and the structured text into the vertical field big model to obtain the explicit qualification feature;

[0058] The explicit qualification features are aligned with the vertical domain ontology library, and a manual terminology review queue is triggered for unlisted terms.

[0059] The target feature list includes at least: skill name (such as Java), years of experience and position (such as 5 years, back-end development engineer), education background (including degree, major and school; such as undergraduate, computer science, Tsinghua University), and job description (such as developing microservice systems). These are used to comprehensively extract and analyze job and candidate information to ensure the accuracy of job matching.

[0060] Specialized prompt templates are reusable instruction templates designed for extracting specific types of information, aiming to improve the consistency and efficiency of extraction tasks. Among them, enumeration extraction prompts are a special prompt template, the purpose of which is to allow the vertical domain big model to find and list all items belonging to a specific category from a given text, such as skill lists, tool lists, and language lists. Relation extraction prompts are another special prompt template, whose purpose is not only to identify entities in the text (such as companies and positions), but also to extract specific relationships between these entities, such as a work experience (including multiple related information such as company, position, time, responsibilities, etc.).

[0061] Taking professional background as an example, the structured text is shown in Table 1. The prompt template and structured text block are input into the vertical field large model, such as using BERT variants fine-tuned with industry corpus (such as BioBERT for the medical field and FinBERT for the financial field). The explicit qualification features are shown in Table 2.

[0062] Table 1 Structured text examples

[0063] Block Number Label Content Block 1 Experience 2018 - 2023; ABC Technology Company; Position held: Back - end Development Engineer Block 2 Responsibilities Responsible for developing and maintaining a Java - based microservices system using Spring Boot Block 3 Projects Participated in multiple large - scale projects Block 4 Education Graduated from Tsinghua University, majoring in Computer Science, with a solid programming foundation

[0064] Table 2 Examples of explicit qualification characteristics

[0065] Block Number Label Content Block 1 Experience 5 years as a back - end development engineer at ABC Technology Company Block 2 Responsibilities Java; Spring Boot Block 3 Projects Developed microservices system; Maintained microservices system Block 4 Education Bachelor's degree in Computer Science, Tsinghua University

[0066] In addition, the term alignment process includes matching the extracted terms with the ontology library, using exact matching or fuzzy matching based on word form and aliases. Suppose terms such as "Java" and "Python" are successfully matched to the standard terms in the ontology library, while "Hibernate" and "Spring Boot" are not found. The unmatched terms "Hibernate", "Spring Boot" and the company name "ABC Technology Company" are automatically added to the "Manual Term Review Queue". The data administrator will process this queue regularly, review the validity of these terms, and decide whether to add them to the ontology library or use them as aliases of existing terms.

[0067] Using a dedicated prompt template to guide the large model in the vertical domain makes the extraction task more explicit, thus improving the efficiency and accuracy of talent feature extraction. Term alignment is carried out through the vertical domain ontology library to ensure the unity and standardization of feature expression. The manual term review queue mechanism is used to capture new knowledge, which is incorporated into the ontology library after review, forming a knowledge accumulation cycle and enhancing the system's adaptability. In addition, ontology library alignment also plays a verification role, discovering and correcting extraction errors, and improving the quality and reliability of talent features.

[0068] Furthermore, the extraction process of the implicit ability features includes:

[0069] Using the large model in the vertical domain to perform semantic reasoning on the structured text to generate local implicit ability features and preliminary confidence scores;

[0070] According to the preliminary confidence scores, using the attention mechanism to adjust the weights of the local implicit ability features, combining the time series information between the structured texts, and using a time decay function to calculate the final aggregated weights to generate the implicit ability features with the confidence scores.

[0071] Specifically, a tokenizer customized for the vertical domain is used to ensure the integrity of professional terms (such as "Spring Boot", "ScrumMaster") and avoid incorrect segmentation. The input structured text is converted into a sequence of Token IDs, and special tokens (such as [CLS] and [SEP]) are added at the beginning and end of the sequence. This sequence is input into the large model in the vertical domain to obtain a context-related word embedding sequence with a high dimension (such as 768 or 1024 dimensions). E CLS The output embedding for the [CLS] Token, used as the aggregated representation of the entire text block, is input E CLS into one or more feed-forward neural networks, which form the classification head. Each classification head corresponds to a local implicit ability feature c, and the output layer uses the Sigmoid function to calculate the probability P(c|E of the text block embodying c CLS) The probability value output by the classification head is used as the initial confidence score.

[0072] Construct an input vector x for each time step t (corresponding to a structured text) t . Using the self-attention layer in the Transformer, automatically calculate the contribution weight a of each time step t to the aggregated representation t . Define the time decay function decay(t), expressed as: decay(t) = exp(-λ * time_diff_from_present); Combine the attention weight and time decay to obtain the final aggregated weight w for each time step t , expressed as: w t = normalize(a t * decay(t)); where exp() is the exponential function, λ is the decay rate, time_diff_from_present is the time difference between time step t and the current time, and normalize() is the normalization function.

[0073] Use the final aggregated weight w t to perform weighted averaging on the local latent ability features of each time step to obtain the global confidence, expressed as:

[0074] AvgConf = ∑w t · c t ;

[0075] VarConf = ∑w t · (c t - AvgConf) 2 ;

[0076] c global = f(AvgConf, VarConf);

[0077] where AvgConf is the weighted average of the local confidence, c t is the initial confidence score for time step t, VarConf is the weighted variance, and c global is the confidence score; f() can represent the Sigmoid function or the Logistic function.

[0078] By combining time series information to perform weighted aggregation on local features in different periods, it is possible to evaluate the development trend of talent capabilities and provide a more dynamic and comprehensive perspective. Integrating multiple evidence sources and assigning different weights avoids over-inference due to single and one-sided descriptions, making the global feature evaluation results more robust and reliable. The multi-stage confidence calculation combines the semantic relevance of evidence to provide a more refined and trustworthy confidence score for the final latent ability evaluation, thereby improving the depth and accuracy of the evaluation.

[0079] Step 3: Figure 2 This is a schematic diagram of the process for generating a job requirement portrait and a professional ability portrait for the present invention. As Figure 2 shown, connect the explicit qualification features and the implicit ability features to the talent knowledge graph, perform entity linking and querying, and generate a job requirement portrait and a professional ability portrait.

[0080] Among them, the talent knowledge graph refers to storing talent-related entities (such as skills, positions, industries, and companies, etc.) and their relationships (such as "Java is a programming skill" and "Backend developers need Java") in the form of a graph; the job requirement portrait refers to a structured description generated based on enterprise requirements (such as job descriptions); the professional ability portrait refers to a structured description generated based on the candidate's professional background (such as a resume).

[0081] Further, the process of the entity linking and querying includes:

[0082] Screen out an entity list from the explicit qualification features, such as Java, Spring Boot, and backend developers, etc.

[0083] For each entity string in the entity list, call the entity linking API provided by the talent knowledge graph. The API will use methods such as string matching and context disambiguation to return a matching graph node and a graph confidence level; for example, the API returns: node "Java (programming language)", confidence level 0.98.

[0084] Select the best graph node according to the matching graph node and the graph confidence level. Usually, select the matching with the highest confidence level and the type meeting the expectation as the "best graph node" of this entity. Store the ID of this best graph node and associate it with the original explicit feature. Execute a graph query (such as SPARQL query) on the ID of the best graph node to obtain graph association information, such as the association information of the entity "Java" being "belongs to the commonly used skills for backend development and is strongly related to Spring Boot".

[0085] According to the pre-defined Schema of the "job requirement portrait" or "professional ability portrait", supplement the graph association information into the corresponding explicit qualification features, and combine with the implicit ability features to generate the job requirement portrait and the professional ability portrait.

[0086] Table 3 Job Requirement Portrait

[0087] Feature Type Label Feature Description Explicit Feature Skills Java (Programming language, strongly related to Spring Boot) Explicit Feature Skills Spring Boot (Framework, commonly used in microservices development) Explicit Feature Experience 5 years as a back - end development engineer (requiring Java, Spring Boot) Explicit Feature Education Bachelor's degree in Computer Science, Tsinghua University (cultivating solid programming skills) Explicit Feature Responsibilities Developed microservices system (microservices development) Explicit Feature Responsibilities Maintained microservices system Implicit Feature Problem - solving ability Confidence level: 0.90 Implicit Feature Team collaboration ability Confidence level: 0.91

[0088] Specifically, the job requirement portrait is shown in Table 3. Through entity linking, the terms extracted from the text (such as "Spring Boot") are mapped to the standard concept nodes in the knowledge graph, thereby enriching the connotation of the features. For example, understanding and mastering Spring Boot means having relevant knowledge of microservice development. Through knowledge graph query, the implicit associated knowledge that originally required expert experience to judge is made explicit and injected into the portrait, enhancing the information volume and intelligence level of the portrait. The entire process (screening entities, linking, querying, and integrating) forms a closed loop, combining explicit and implicit features with structured domain knowledge, making the generated job requirement portrait and professional ability portrait more comprehensive and clearer in structure, providing high-quality data support for accurate matching.

[0089] Step 4: Based on the job requirement portrait, generate a candidate portrait set from the professional ability portrait using a hybrid retrieval strategy;

[0090] Among them, the candidate portrait set refers to a relatively small subset selected from the entire "professional ability portrait" database.

[0091] Furthermore, the implementation process of the hybrid retrieval strategy includes:

[0092] For the job requirement portrait, analyze the hard constraint conditions and extract the semantic core to generate a demand vector representation;

[0093] Among them, the hard constraint conditions refer to the conditions that must be met, such as "5 years of experience" and "Java skills". The semantic core includes text or features such as key skill requirements, core responsibility descriptions, and the emphasis on target implicit capabilities in the portrait. Through the domain Sentence-BERT model to encode the key text descriptions, or combined with structured feature encoding, the semantic core of the job requirements is transformed into a demand vector representation with a fixed dimension, for example, the demand vector: [0.22, 0.64,...].

[0094] Then, by constructing a database query, apply the hard constraint conditions to the corresponding structured fields of the professional ability portrait database. Use the efficient index of the database or search engine to execute this query, and output a list of candidate IDs that meet all the hard constraint conditions, that is, the "preliminary subset of candidate portraits".

[0095] Take the demand vector as the query vector, obtain the professional vector representation of the preliminary subset of candidate portraits, perform KNN search, calculate the similarity score (such as cosine similarity) between the demand vector representation and each professional vector representation, and screen out the top N candidates with the highest scores (N is the preset recall number, such as 500), or candidates with scores higher than the preset threshold to obtain a list of candidate IDs;

[0096] Retrieve the corresponding, complete, and structured "professional ability portraits" in batches from the database according to the candidate ID list, and output the candidate portrait set.

[0097] After quickly narrowing down the scope through hard constraint filtering, semantic similarity retrieval is then performed. By using vector representation to capture deep semantic information, candidates who do not use the same keywords but have highly matching abilities and experiences can be identified, effectively improving the recall rate of relevant candidates and reducing omissions. This hybrid strategy combines the precision of structured data filtering and the comprehensiveness of semantic retrieval, achieving a better balance among efficiency, accuracy, and recall rate compared to a single retrieval method.

[0098] Step Five: Input the job requirement portrait and the candidate portrait set into the SFT-LLM model (supervised fine-tuning large language model), output the job suitability, and based on the attention mechanism, output a person-job matching analysis report to obtain a candidate list;

[0099] Among them, the candidate list is the result obtained after in-depth analysis, precise scoring, and ranking of the candidate portrait set.

[0100] Furthermore, the generation process of the candidate list includes:

[0101] Encode the job requirement portrait and the candidate portrait set into a multi-modal input structure suitable for the SFT-LLM model, including Pointwise input, Pairwise input, and Listwise input; for the Pointwise input, directly output the job suitability; for the Pairwise input and the Listwise input, output the job suitability of the relative ranking. As shown in Table 4, the job suitability output by the SFT-LLM model inference is given.

[0102] Table 4 Job Suitability

[0103]

[0104] Obtain the attention distribution of each layer (especially the cross-attention layer or the top self-attention layer) when the model calculates the suitability score. These weights α ijIndicates the degree of attention of the model to other elements j in the input when processing a certain element i. Analyze the internal attention weight distribution of the SFT-LLM model, find the parts with high cross-attention weights between the demand portrait and the candidate portrait, and identify the parts with high attention and strong correlation with the matching judgment in the model as positive matching evidence or mismatch points. Among them, positive matching evidence means that the key skills of the demand are highly relevant to the candidate's description, while the lack of strong evidence support or the existence of contradictory information in the candidate's description for the key requirements of the demand is a mismatch point. By extracting these positive matching evidence and mismatch points, a detailed person-job matching analysis report is generated to help more intuitively understand the matching results. As shown in Table 5, the candidates are prioritized according to the job suitability and stored in association with the person-job matching analysis report to obtain the candidate list.

[0105] Table 5 Generate candidate list

[0106]

[0107] The person-job matching analysis report based on the attention mechanism reveals the key basis for the model's decision-making and makes the decision-making process of the "black box" model transparent. This not only enhances the user's trust in the results but also saves the user the cumbersome process of screening, comparing, and judging from a large number of candidates, improving the efficiency and quality of talent screening and matching decisions.

[0108] Step Six: Calculate the user's preference feedback on the candidate list through the interaction interface, and map the preference feedback to a reward signal based on direct rules.

[0109] Among them, preference feedback refers to the information about which results are better and which results are worse conveyed by the user's behavior or explicit expression when viewing and operating the candidate list.

[0110] Furthermore, Figure 3 is a schematic flow diagram of the calculation process of the reward signal of the present invention. As Figure 3 shown, the calculation process of the reward signal includes:

[0111] Collect the user's sorting adjustment operations, explicit score inputs, and employment status markings for candidates from the interaction interface using the event-driven mechanism, and quantify them as preference interaction data;

[0112] Among them, the sorting adjustment operation means that the user moves candidate C i from the initial position pos init to the final position pos final . When the user adjusts the position of the candidate, calculate the moving distance Δpos i = pos init - pos final ; the explicit score input means that the user gives a score for Ci Allocate rating s i (Range from the minimum rating s min to the maximum rating s max , such as 1 - 5 stars), normalize the user rating s i to the interval [0, 1]; The employment status flag refers to the status status i of the user flag C i (such as "Accepted for interview" and "Pending for irrelevance") and optionally with a reason reason i (such as "Skill mismatch" and "Salary reason").

[0113] Perform single - feedback mapping on the preference interaction data one by one to generate sorting adjustment rewards, explicit rating rewards, and employment status rewards, expressed as:

[0114] r pos,i = γ·Δpos i ;

[0115]

[0116] where r pos,i is the i - th sorting adjustment reward, γ is the reward coefficient for each movement (such as 0.1); r rating,i is the i - th explicit rating reward, is the scaling factor (such as 0.8), β is the neutral benchmark (such as 0.5), rating(C i ) is the quantized explicit rating input; r status,i is the i - th employment status reward, status flag Status(C i ) = status i , reason flag Reason(C i ) = reason i , ∧ is a logical symbol representing the "and" operation.

[0117] Perform multi - feedback aggregation on the sorting adjustment reward, the explicit rating reward, and the employment status reward, and adjust according to the demand priority and candidate scarcity to obtain the reward signal.

[0118] Among them, multi - feedback aggregation can perform weighted summation or take the maximum value of the sorting adjustment reward, the explicit rating reward, and the employment status reward to obtain the multi - feedback aggregation result r total,i , and adjust according to the demand priority p j and candidate scarcity q i , expressed as:

[0119]

[0120] where r adj,i is the reward signal, θ and are adjustment coefficients (such as 0.1); the demand priority refers to the urgency of the current recruitment demand among all pending demands, and the acquisition methods include: manual setting and marking and automatic judgment based on rules (such as the job vacancy time being greater than 90 days); the candidate scarcity refers to the relative abundance or scarcity of candidates who meet the core conditions of the demand in the talent pool or the market. The acquisition process includes: if in the internal "professional ability profile" database, the number of candidates returned by the query is much lower than the number usually returned when processing other demands, or lower than a preset threshold, then it is determined that the candidate scarcity of this demand is high. On the contrary, if the returned number is large, the scarcity is low.

[0121] By collecting and mapping multiple interaction behaviors (such as sorting, scoring, and status marking), rather than a single signal, it is possible to more comprehensively capture the multi-dimensional preferences of users and provide richer and more accurate guiding information for model optimization. In addition, by allowing different aggregation methods and adjustment factors, the final reward signal can be flexibly shaped, and the influence degree of feedback information and business context on model learning can be finely regulated.

[0122] Step 7: Optimize the internal parameters of the SFT-LLM model according to the reward signal using a reinforcement learning algorithm.

[0123] Among them, the internal parameters refer to all the learnable numerical values that make up the SFT-LLM deep neural network model itself, including weights, biases, etc.

[0124] The present invention provides a more accurate matching effect by integrating a large model in a vertical domain, a knowledge graph, and a reinforcement learning feedback mechanism. First, through refined preprocessing and feature extraction, a comprehensive "job demand profile" and "professional ability profile" are constructed, which overcomes the defects of traditional methods that rely on keywords, cannot capture implicit qualities, and lack understanding of the domain background, and lays a high-quality data foundation for accurate matching. Secondly, the strategy of "hard filtering and semantic retrieval" is adopted to efficiently recall relevant candidates, and deep sorting is performed through the optimized SFT-LLM to output accurate fitness and interpretable matching analysis reports. Finally, a reinforcement learning optimization closed-loop based on the real preference feedback of users is introduced, enabling the model to continuously learn and adapt to changes in user needs, ensuring that the current market demand for efficient and accurate talent matching is met.

[0125] Example 2:

[0126] Based on the method steps described in Embodiment 1, a large technology company applied the talent feature extraction and matching method described in the present invention when recruiting senior artificial intelligence researchers (focused on specific cutting-edge fields, such as Transformer models for protein folding). The company connected its internal ATS system and external professional network data sources to this method and adopted a large model for vertical fields for in-depth training and an internal talent knowledge graph. Since the skill requirements for this research position are highly specialized and evolving rapidly, traditional keyword screening and manual review are difficult to effectively identify truly top-notch candidates and their implicit abilities such as research potential, and a large number of applications need to be processed. Therefore, there is an urgent need for the deep semantic understanding, precise matching and ranking, and continuous optimization capabilities provided by this method. The talent feature extraction and matching method based on a large model for vertical fields includes:

[0127] Perform text preprocessing and block segmentation on the input enterprise requirements and career backgrounds to obtain structured text; use the large model for vertical fields to extract explicit qualification features, and perform context reasoning on the structured text to generate implicit ability features with confidence scores; connect the explicit qualification features and the implicit ability features to the talent knowledge graph, perform entity linking and querying, and generate a job requirement portrait and a professional ability portrait;

[0128] Based on the job requirement portrait, adopt a hybrid retrieval strategy to generate a candidate portrait set from the professional ability portrait; input the job requirement portrait and the candidate portrait set into the SFT-LLM model, output the job suitability, and output a person-job matching analysis report based on the attention mechanism to obtain a candidate list;

[0129] Calculate the preference feedback of the user on the candidate list through the interaction interface, map the preference feedback to a reward signal based on direct rules; use the reinforcement learning algorithm to optimize the internal parameters of the SFT-LLM model according to the reward signal.

[0130] Specifically, sample a batch of experience segments of (state, action, reward signal) from the accumulated feedback data. The state can be understood as a pair of <job requirement portrait, candidate portrait>, and the action is the ranking or scoring given by the model. In this embodiment, the PPO algorithm is used to evaluate the reward r obtained by taking a certain action (such as ranking candidate A first) in a specific state adj,AHow much of an "advantage" it has compared to the average expected reward in this state. An action that receives a positive reward (e.g., candidate B is selected for an interview) has a positive advantage, and an action that receives a negative reward (e.g., candidate D is rejected) has a negative advantage. PPO calculates the reward signal. At the same time, a KL divergence penalty term is used to ensure that the new policy (updated model parameters) does not deviate too far from the old policy (original parameters) to maintain the stability and generalization ability of the model. Finally, an optimizer (such as Adam) is used to update all the internal parameters of the SFT-LLM model (including the weights and biases of the attention layer, feed-forward layer, etc.). After this update iteration, a new version of the SFT-LLM model is obtained. When dealing with future requirements, this new model will be more likely to rank candidates like candidate B, who may be recognized by the user, higher, and candidates like D lower.

[0131] As shown in Table 6, the candidate list after model update is given.

[0132] Table 6 Updated Candidate List

[0133]

[0134] Although the embodiments of the present invention have been shown and described, it will be understood by those of ordinary skill in the art that various changes, modifications, substitutions, and variations can be made to these embodiments without departing from the principles and spirit of the present invention. The scope of the present invention is defined by the appended claims and their equivalents.

Claims

1. A method for talent feature extraction and matching based on a large model in a vertical domain, characterized in that include: Perform text preprocessing and block segmentation on the input enterprise needs and professional background to obtain structured text; Use the vertical domain big model to extract explicit qualification features, and perform contextual reasoning on the structured text to generate implicit ability features with confidence scores; connect the explicit qualification features and the implicit ability features to the talent knowledge graph, perform entity linking and query, and generate job requirement portraits and professional ability portraits; Based on the job requirement profile, a hybrid retrieval strategy is used to generate a candidate profile set from the professional ability profile; Input the job requirement profile and the candidate profile set into the SFT-LLM model, output the job suitability, and output a person-job matching analysis report based on the attention mechanism to obtain a candidate list; The user's preference feedback on the candidate list is calculated through an interactive interface, and the preference feedback is mapped into a reward signal based on a direct rule; and the internal parameters of the SFT-LLM model are optimized using a reinforcement learning algorithm according to the reward signal.

2. The talent feature extraction and matching method based on the vertical domain large model according to claim 1, characterized in that The process of extracting the explicit qualification features includes: Define the target feature list to specify explicit feature types and explicit feature formats; Designing dedicated prompt templates for different explicit feature types to instruct the vertical domain big model to extract the explicit qualification features; wherein, skill features use enumeration extraction prompts, and experience features use relational extraction prompts; Inputting the dedicated prompt template and the structured text into the vertical field big model to obtain the explicit qualification feature; The explicit qualification features are aligned with the vertical domain ontology library, and a manual terminology review queue is triggered for unlisted terms.

3. The talent feature extraction and matching method based on the vertical domain large model according to claim 1, wherein The extraction process of the implicit capability features includes: Using the vertical domain big model to perform semantic reasoning on the structured text, generating local implicit capability features and preliminary confidence scores; According to the preliminary confidence score, the weight of the local implicit ability feature is adjusted using the attention mechanism, and the final aggregation weight is calculated using the time decay function in combination with the time series information between the structured texts to generate the implicit ability feature with the confidence score.

4. The talent feature extraction and matching method based on the vertical domain large model according to claim 1, wherein The entity linking and querying process includes: Filtering an entity list from the explicit qualification features; for each entity string in the entity list, calling the entity linking API provided by the talent knowledge graph, and returning a matching graph node and a graph confidence; The best graph node is selected according to the matching graph node and the graph confidence, and a graph query is performed on the best graph node to obtain graph association information; the graph association information is supplemented to the corresponding explicit qualification characteristics, and combined with the implicit ability characteristics to generate the job requirement portrait and the professional ability portrait.

5. The talent feature extraction and matching method based on the vertical domain large model according to claim 1, wherein The implementation process of the hybrid search strategy includes: For the job requirement profile, analyze the hard constraints, extract the semantic core, and generate a requirement vector representation; According to the hard constraints, query the professional ability profile to obtain a preliminary subset of candidate profiles; Obtain the occupational vector representation of the preliminary subset of candidate portraits, perform KNN search, calculate the similarity score between the requirement vector representation and the occupational vector representation, and filter out the candidate ID list; Output the candidate portrait set according to the candidate ID list.

6. The talent feature extraction and matching method based on the vertical domain large model according to claim 1, wherein The generation process of the candidate list includes: Encode the job requirement portrait and the candidate portrait set into a multi-modal input structure adapted to the SFT-LLM model, including Pointwise input, Pairwise input, and Listwise input; for the Pointwise input, directly output the job suitability; for the Pairwise input and the Listwise input, output the relative sorted job suitability; Analyze the internal attention weight distribution of the SFT-LLM model, extract positive matching evidence and mismatch points, generate the human-job matching analysis report; sort the candidates according to the job suitability and store them in association with the human-job matching analysis report to obtain the candidate list.

7. The method for talent feature extraction and matching based on a vertical domain large model according to claim 1, wherein The calculation process of the reward signal includes: Collect the user's sorting adjustment operations, explicit score inputs, and employment status markings for candidates from the interaction interface, and quantify them into preference interaction data; Perform single-feedback mapping on the preference interaction data one by one to generate sorting adjustment rewards, explicit score rewards, and employment status rewards; Aggregate the sorting adjustment rewards, the explicit score rewards, and the employment status rewards through multi-feedback, and adjust them according to the requirement priority and candidate scarcity to obtain the reward signal.

Citation Information

Cited By

  • Human resource enhancement recommendation method and device based on explicit feedback

    CN120634494A

  • Intelligent resume screening method and system based on dynamic rule fusion

    CN120996157A

  • Post talent portrait automatic generation system and method

    CN121352750A

  • AI-based multi-dimensional enterprise talent matching method and system, and storage medium

    CN121414312A

  • Talent ability evaluation and post matching system based on knowledge graph

    CN122066292A