Method and device for selecting supervised fine-tuning training data of man-post matching model

By employing multi-dimensional evaluation and business feedback calibration methods, the problem of training data selection in the supervised fine-tuning stage of the person-job matching model was solved, improving the model's practicality and training efficiency, and ensuring the scientific nature and business value orientation of data selection.

CN121329362AActive Publication Date: 2026-01-13ENGLISH SHI INTERCONNECTION BEIJING INFORMATION TECH CO +1
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
CN202511874269.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-12
Publication Date
2026-01-13
Estimated Expiration
2045-12-12

AI Technical Summary

Technical Problem

Existing technologies neglect the differences in matching difficulty between job descriptions and candidate resumes during the supervised fine-tuning stage of the job matching model. This leads to overfitting of the model to simple patterns, insufficient learning of complex cases, lack of business feedback loop mechanism, and blind selection of training data, which limits the model's generalization ability and robustness.

Method used

By evaluating the matching difficulty of job description-candidate resume pairs from multiple dimensions and combining the real behavior of recruitment decision-makers for data calibration, a supervised fine-tuning training set with a specific reasoning style is constructed to ensure the scientific nature and business relevance of the training data.

Benefits of technology

It achieves accurate assessment of matching difficulty, improves the model's practicality and business relevance, optimizes training efficiency and effectiveness, and enhances the model's generalization ability and robustness.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121329362A_ABST
    Figure CN121329362A_ABST
Patent Text Reader

Abstract

The invention discloses a method and a device for selecting supervised fine-tuning training data of a man-post matching model, and relates to the technical field of human resource data processing based on artificial intelligence. The method comprises the steps of obtaining an initial difficulty score of a post-resume pair through multi-dimensional difficulty assessment; generating a matching label based on the real behavior of the recruitment decision maker; a comprehensive difficulty score is obtained through the comparison model preliminary judgment and matching label consistency calibration; and selecting data adaptive to the difficulty level according to the comprehensive difficulty score to construct a training set of a specific reasoning style. According to the method, accurate screening and optimal configuration of the training data are realized, and the accuracy, generalization ability and training efficiency of the man-post matching model are effectively improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of human resource data processing based on artificial intelligence, in particular to a selection method and device for supervised fine-tuning training data of a human-post matching model. BACKGROUND

[0002] In the current field of artificial intelligence-driven recruitment technology, human-post matching models based on natural language processing have become the core tool for improving the efficiency of resume screening and job recommendation. The performance of such models, especially in the supervised fine-tuning phase, is highly dependent on the quality of the training data used. Existing technologies mainly have the following defects when constructing supervised fine-tuning training sets: (1) The inherent difficulty differences in matching judgments of different position descriptions and resume combinations are ignored. Simple matching tasks (such as direct comparison of education and work experience) and complex matching tasks (such as skill transferability and achievement correlation) are mixed, resulting in overfitting of the model to simple patterns and insufficient learning of complex cases, which restricts the generalization ability and robustness of the model; (2) There is a lack of business feedback loop mechanism. The labeling of training data is usually based on offline static labels and does not incorporate real behavior feedback from recruitment decision makers (such as enterprise HR), resulting in a disconnect between data selection and actual business needs, making it difficult for the model to learn key decision factors.

[0003] (3) It cannot guide the construction of training sets. There is a specific synergistic effect between different styles of reasoning process data and different source data of different judgment difficulty, but existing methods do not systematically analyze this correlation, resulting in blind selection of training data and limiting the further improvement of model training efficiency and effectiveness.

[0004] Therefore, there is an urgent need for a new technology that can accurately assess the matching judgment difficulty of position description-candidate resume pairs, dynamically calibrate by incorporating real business feedback, and scientifically guide the selection of training data and the construction of training sets in the supervised fine-tuning phase. SUMMARY

[0005] In view of the above-mentioned defects or shortcomings in the prior art, the present application provides a selection method and device for supervised fine-tuning training data of a human-post matching model, which can effectively solve all the technical problems mentioned in the background art.

[0006] In one aspect of the present application, a selection method for supervised fine-tuning training data of a human-post matching model is provided, comprising the following steps: The matching judgment difficulty of the position description-candidate resume pair is evaluated by the artificial intelligence model in multiple dimensions to obtain the difficulty score of each dimension, and the difficulty scores of each dimension are weighted and summed to obtain the initial difficulty score of the position description-candidate resume pair; The matching label generated based on the real behavior of a recruitment decision maker, the real behavior of the recruitment decision maker including positive feedback operations and negative feedback operations on candidate resumes; a preliminary matching judgment is made on the position description-candidate resume pair by an artificial intelligence model to obtain a preliminary matching result, if the preliminary matching result is consistent with the matching label, the initial difficulty score is down-weighted, if the preliminary matching result is inconsistent with the matching label, the initial difficulty score is up-weighted, and a comprehensive difficulty score of the position description-candidate resume pair is obtained; According to the comprehensive difficulty score, a position description-candidate resume pair with an adaptive difficulty level is selected to construct a supervised fine-tuning training set of a specific reasoning style; wherein the reasoning style is differentiated supervised fine-tuning training data in the detail, logical structure or language style of the reasoning process.

[0007] Another aspect of the present application also provides a selection device of supervised fine-tuning training data of a human-post matching model, comprising: The model scoring module is used for multi-dimensional evaluation of the matching difficulty of the position description-candidate resume pair by an artificial intelligence model to obtain a difficulty score in each dimension, and the difficulty scores in each dimension are weighted and summed to obtain an initial difficulty score of the position description-candidate resume pair; The scoring calibration module is used for obtaining a matching label generated based on the real behavior of a recruitment decision maker, the real behavior of the recruitment decision maker including positive feedback operations and negative feedback operations on candidate resumes; a preliminary matching judgment is made on the position description-candidate resume pair by an artificial intelligence model to obtain a preliminary matching result, if the preliminary matching result is consistent with the matching label, the initial difficulty score is down-weighted, if the preliminary matching result is inconsistent with the matching label, the initial difficulty score is up-weighted, and a comprehensive difficulty score of the position description-candidate resume pair is obtained; The training data screening module is used for selecting a position description-candidate resume pair with an adaptive difficulty level to construct a supervised fine-tuning training set of a specific reasoning style according to the comprehensive difficulty score; wherein the reasoning style is differentiated supervised fine-tuning training data in the detail, logical structure or language style of the reasoning process.

[0008] The selection method and device of supervised fine-tuning training data of a human-post matching model provided by the present application have the following beneficial effects: (1) Accurate evaluation of the matching difficulty of the position description-candidate resume pair Through the multi-dimensional difficulty evaluation mechanism, the matching judgment difficulty of the position description-candidate resume pair is systematically analyzed from three aspects of explicit information, implicit information and fuzzy information, ensuring the comprehensiveness and accuracy of the difficulty score, and providing a scientific basis for the selection of subsequent training data.

[0009] (2) Data calibration oriented by business value Innovatively introducing the real behavior of the recruitment decision maker as a calibration factor, the algorithm evaluation is directly linked with the business effect, ensuring that the high-difficulty data screened out has more training value, effectively improving the practicability and business fitting degree of the model.

[0010] (3) Optimization of training efficiency and effect By establishing the adaptability relationship of the comprehensive difficulty score of the reasoning style and the position description-candidate resume pair, the source data with the most learning value can be selected for the specific supervised fine-tuning training target, significantly improving the efficiency and final performance of model training, and providing differentiated data selection strategies for models with different parameter sizes. BRIEF DESCRIPTION OF DRAWINGS

[0011] Other features, objects and advantages of the present application will become more apparent from the following detailed description of non-limiting embodiments made with reference to the accompanying drawings: Figure 1 is a flowchart of a method for selecting supervised fine-tuning training data of a person-job matching model according to an embodiment of the present application; Figure 2 is a structural schematic diagram of a device for selecting supervised fine-tuning training data of a person-job matching model according to an embodiment of the present application; Figure 3 is a structural schematic diagram of an electronic device according to an embodiment of the present application. DETAILED DESCRIPTION

[0012] To make the objectives, technical solutions and advantages of the embodiments of the present application clearer, the technical solutions in the embodiments of the present application will be described clearly and completely below with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are some but not all of the embodiments of the present application. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative work fall within the protection scope of the present application.

[0013] An embodiment of the present application provides a method for selecting supervised fine-tuning training data of a person-job matching model, and the specific implementation process of the method will be described in detail below. Referring to Figure 1 , the method comprises the following steps: Step S101: The matching judgment difficulty of the position description-candidate resume pair is evaluated by the artificial intelligence model in multiple dimensions to obtain an initial difficulty score.

[0014] The traditional method ignores the matching difficulty difference of the position description-candidate resume pair as the training data of the supervised fine-tuning stage, resulting in the performance decline of the model in complex scenarios. This step establishes a multi-dimensional evaluation system to ensure that the training data covers different matching difficulty levels, thereby improving the robustness of the model.

[0015] Specifically, different position description-candidate resume pairs have inherent difficulty differences in matching judgment, which directly affects the training effect of the subsequent model supervised fine-tuning stage. By designing structured evaluation dimensions, this difficulty difference can be quantified systematically. In specific implementation, first, a large language model (such as GPT-4, ChatGPT, Deepseek, etc.) is used to perform multi-dimensional matching difficulty analysis on the input position description-candidate resume pair, obtaining difficulty scores for each dimension. The difficulty scores for each dimension are weighted and summed to obtain the initial difficulty score of the position description-candidate resume pair.

[0016] The multi-dimensions in this step include explicit information judgment difficulty, implicit information judgment difficulty, and fuzzy information judgment difficulty. Among them, the explicit information judgment difficulty evaluation quantifies the direct comparison difficulty, including education, work experience, professional requirements, etc. The implicit information judgment difficulty evaluation requires information derived from industry knowledge and judgment, including skill transfer difficulty, achievement correlation difficulty, and work experience completion difficulty; among them, skill transfer difficulty is used to evaluate the feasibility and matching degree of the candidate's skill across domains and tools; achievement correlation difficulty is used to evaluate the correlation strength between the candidate's achievements derived from their project experience and the requirements of the target position; work experience completion difficulty is used to evaluate the complexity of relying on common sense and industry experience to complete the judgment when there is missing or ambiguous information in the candidate's resume. Fuzzy information judgment difficulty is used to evaluate the basic judgment difficulty caused by ambiguous text description or missing information.

[0017] When the difficulty scores for each dimension are weighted and summed, the weight of the implicit information judgment difficulty is greater than that of the explicit information judgment difficulty, and the weight of the explicit information judgment difficulty is greater than that of the fuzzy information judgment difficulty. The preferred weight distribution is: implicit information judgment difficulty 55%, explicit information judgment difficulty 25%, and fuzzy information judgment difficulty 20%. This weight distribution reflects the importance difference of different information dimensions in recruitment decision-making, ensuring the business reasonableness of the difficulty score.

[0018] Step S102: Obtain the matching label generated based on the real behavior of the recruitment decision maker, and perform preliminary matching judgment through an artificial intelligence model. According to the consistency of the preliminary matching result with the matching label, the initial difficulty score is calibrated to obtain the comprehensive difficulty score.

[0019] Since the offline labeling of the difficulty of the matching judgment of the job description-candidate resume pair by the artificial intelligence model (such as a large model) is somewhat disconnected from the real decision, if only the difficulty score labeled by the artificial intelligence model is used to screen the training data, the model after supervised fine-tuning will have a learning bias. Therefore, the real behavior of the recruitment decision maker (such as downloading the resume, marking as unsuitable) is introduced, which reflects the real matching judgment in the actual business scenario and is the most valuable supervision signal, which can ensure that the data selection meets the actual recruitment needs. In specific implementation, the user behavior log can be obtained from the back-end database of the recruitment platform, and behaviors such as “downloading the resume” and “saving the resume” are parsed as positive feedback operations, and the corresponding matching label is “match”; behaviors such as “marking as unsuitable” are parsed as negative feedback operations, and the corresponding matching label is “unmatch”. At the same time, the same or different artificial intelligence model (such as a large language model) is used to make a preliminary matching judgment on the same job description-candidate resume pair, and a preliminary matching result of “match” or “unmatch” is obtained.

[0020] Further, the initial difficulty score is calibrated. When the preliminary judgment of the artificial intelligence model is consistent with the real behavior of the recruitment decision maker, it means that the judgment is relatively easy, and when the two are inconsistent, it means that there is a cognitive difference or a complex situation, indicating that the judgment is more challenging. Therefore, if the preliminary matching result is consistent with the matching label, the initial difficulty score is multiplied by a weight coefficient less than 1 (preferably 0.8) to reduce the initial difficulty score; if not, the initial difficulty score is multiplied by a weight coefficient greater than 1 (preferably 1.2) to increase the initial difficulty score. This calibration mechanism ensures that those cases that are truly challenging in business practice can obtain a higher difficulty score, thereby being given priority consideration in subsequent training data selection. The calibrated initial difficulty score is referred to as the comprehensive difficulty score, and using the comprehensive difficulty score to screen suitable training data for the supervised fine-tuning stage can recommend fewer obviously mismatched positions.

[0021] Step S103: According to the comprehensive difficulty score, select the job-resume pair with an appropriate difficulty level to construct a supervised fine-tuning training set with a specific reasoning style.

[0022] “Reasoning style” refers to the difference in the level of detail, logical structure, or language style reflected in the model reasoning process data. This difference may be caused by the generation of large language models with different parameter sizes (such as 7B and 32B) or different architectures, or by the difference in prompt words (Prompt), such as different instruction prompts guiding generation, some styles emphasizing the complete presentation and comparison of evidence, and some styles emphasizing step-by-step reasoning according to the importance weight of evidence.

[0023] A "supervised fine-tuning training set with a specific reasoning style" refers to a pre-prepared dataset with inputs and outputs in a specific format for performing supervised fine-tuning (SFT). Its core characteristic is that it requires not only that the model provide a "match / non-match" answer, but also that the model demonstrate the entire logical process of reasoning to arrive at that answer in a fixed, readable, and specific reasoning style.

[0024] For example: Suppose we want to construct a training set for "Style A: Emphasizing stepwise reasoning based on the importance of evidence". Each training data point in this set would take the following form: 1. Input [Job Description]: Seeking one Senior Java Engineer. Requirements: 5+ years of backend development experience, proficient in Spring Cloud microservice architecture, experience in high-concurrency system design, Bachelor's degree or above in Computer Science or related field.

[0025] [Resume Text]: Zhang San has 6 years of Java development experience. He led the microservice transformation of an e-commerce platform, using Spring Boot and Dubbo technologies. He is familiar with MySQL and Redis. He holds a Bachelor's degree in Information Management.

[0026] 2. Expected output (i.e., reasoning style) [Reasoning process]: Analysis of primary core requirements: The job requires "proficiency in Spring Cloud microservice architecture" and "experience in designing high-concurrency systems." These are the most critical technical skills for this position.

[0027] Evidence matching assessment: Candidates must have 6 years of experience, meeting the strict requirement of "5 years or more".

[0028] The candidate has experience leading microservices transformations, but their technology stack is Dubbo rather than Spring Cloud. Although the underlying concepts are similar, the specific technical differences represent the main risk point of this mismatch.

[0029] The candidate's resume did not explicitly mention specific projects or metrics related to "high-concurrency system design," indicating insufficient experience matching.

[0030] The major and academic qualifications meet the requirements, which is a non-critical deduction item.

[0031] Overall assessment: Since the candidate failed to provide direct and compelling evidence supporting their claims regarding the two most critical technical capabilities (Spring Cloud and high concurrency), therefore... [Final Conclusion]: Mismatch.

[0032] As can be seen from the example above, the entire "input-output" pair constitutes one training data point in the "style A training set".

[0033] The technical principle behind this step is that different reasoning styles (i.e., the characteristics of how the model outputs the reasoning process) have a specific fit relationship with training data of different matching difficulty. By optimally matching the two, the training effect can be maximized.

[0034] The specific implementation process includes two sub-steps: Step 1: Establish a mapping relationship between "reasoning style and difficulty level". Keeping the pre-trained model, model hyperparameters, and test set unchanged, multiple training sets are constructed using job description-candidate resume pairs with different difficulty levels for the same reasoning style. The models are then fine-tuned under supervision, and the performance metrics of each fine-tuned model are evaluated on a unified test set. The difficulty level of the training set corresponding to the model with the best performance metrics is determined as the difficulty level that is suitable for the reasoning style.

[0035] Step 2: Selecting Training Data Based on the specific reasoning style used in the current training task, job description-candidate resume pairs with overall difficulty scores falling within the corresponding difficulty level range are selected to construct a supervised fine-tuning training set.

[0036] Furthermore, the method in this embodiment also includes: Step S104: Based on the comprehensive difficulty score, select job description-candidate resume pairs with different appropriate difficulty levels to construct training sets for the models to be trained with different parameter scales.

[0037] Specifically, matching job descriptions to candidate resumes at different difficulty levels allows for the matching of models with different parameter sizes. This is because models with different parameter sizes have varying learning capabilities and capacity limitations. Small-parameter models have limited resources and require simple data for rapid convergence, while large-parameter models are more capable and can learn complex patterns. Therefore, datasets with higher difficulty and more complex content may be better suited to models with larger parameter sizes (e.g., 3B) due to their stronger learning and representation capabilities; conversely, datasets with lower difficulty and clearer patterns may be better suited to models with smaller parameter sizes (e.g., 0.5B), helping them achieve rapid convergence and effective learning under resource-constrained conditions. This provides a basis for differentiated data selection based on different computing resources and performance requirements in practical applications.

[0038] In practice, based on the overall difficulty score, different levels of job descriptions and candidate resumes can be selected to construct supervised fine-tuning training sets for models with different parameter sizes. For models with smaller parameter sizes (e.g., 0.5B), data with lower overall difficulty scores is prioritized; for models with larger parameter sizes (e.g., 3B), data with higher overall difficulty scores is prioritized. This allows the accuracy of the 0.5B model with small parameter sizes to approach that of larger models on suitable, simple data, while the accuracy of the 3B model with large parameter sizes is greatly improved on suitable, highly difficult data, fully leveraging the performance potential of larger models.

[0039] The following describes the complete implementation process of this embodiment in detail, using a real-world application of a large recruitment platform as an example: (1) Data preparation stage We extracted 100,000 job description-candidate resume pairs from the platform's historical logs, including combinations of various industries, job levels, and experience levels.

[0040] (2) Difficulty assessment stage Using a large language model combined with structured prompts, the matching difficulty of each job description-candidate resume pair is assessed across three dimensions.

[0041] For example, when comparing the resume of a senior Java engineer with that of a candidate with 5 years of experience, the difficulty score for judging explicit information is 6.5 (the years of work experience match but the major is slightly different), the difficulty score for judging implicit information is 8.2 (a deep analysis of the transferability of the technology stack and the relevance of project results is required), and the difficulty score for judging fuzzy information is 3.1 (the information is basically complete). After weighted summation, the initial difficulty score is 7.1.

[0042] (3) Business calibration phase The actual feedback on the platform regarding the job description and candidate resumes was that the company's HR downloaded the resumes (the matching tag was "matched"), while the initial judgment result of the large language model was "not matched". Because the initial matching result is inconsistent with the matching tag, the initial difficulty score is weighted up, that is: initial difficulty score 7.1 × weight coefficient 1.2 = overall difficulty score 8.5, which belongs to the "difficult to match judgment" level.

[0043] (4) Training set construction stage Assume the current training task adopts "Style C: Step-by-Step Core Responsibility Matching Analysis," and that this style is known to be best suited for data at the "Match Judgment Difficulty" level through preliminary experiments. Therefore, all job description-candidate resume pairs with an overall difficulty score in the range of 7.5-10.0 are selected and used specifically to generate supervised fine-tuning training data for "Style C."

[0044] This embodiment uses a multi-dimensional difficulty assessment system to accurately quantify the complexity of matching job descriptions with candidate resumes. It introduces a business feedback closed-loop mechanism to dynamically calibrate the algorithm evaluation with real recruitment decisions, ensuring the business value orientation of training data selection. It establishes the adaptability relationship between reasoning style and the difficulty of matching job descriptions with candidate resumes, optimizes data allocation for different training objectives, and implements differentiated data selection strategies based on the characteristics of model parameter scales. This fully leverages the learning potential of models with different parameter scales, ultimately achieving a comprehensive improvement in the accuracy, generalization ability, and practicality of the job matching model, and significantly optimizing training efficiency and resource utilization.

[0045] See Figure 2 In another embodiment of the present invention, a supervised fine-tuning training data selection device 200 for a job-person matching model is provided, including a model scoring module 201, a scoring calibration module 202, and a training data filtering module 203. The supervised fine-tuning training data selection device 200 for the job-person matching model is capable of executing the supervised fine-tuning training data selection method for the job-person matching model in the method embodiment.

[0046] Specifically, the device 200 for selecting supervised fine-tuning training data for the person-job matching model includes: The model scoring module 201 is used to evaluate the difficulty of matching the job description-candidate resume pair through an artificial intelligence model in multiple dimensions, obtain the difficulty score under each dimension, and perform a weighted summation of the difficulty scores under each dimension to obtain the initial difficulty score of the job description-candidate resume pair. The scoring calibration module 202 is used to obtain matching tags generated based on the actual behavior of the recruitment decision-maker, including positive feedback operations and negative feedback operations on candidate resumes; to perform preliminary matching judgment on the job description-candidate resume pair through an artificial intelligence model to obtain a preliminary matching result; if the preliminary matching result is consistent with the matching tag, the initial difficulty score is downweighted; if the preliminary matching result is inconsistent with the matching tag, the initial difficulty score is upweighted to obtain a comprehensive difficulty score for the job description-candidate resume pair. The training data filtering module 203 is used to select job descriptions and candidate resumes that match the difficulty level based on the comprehensive difficulty score to construct a supervised fine-tuning training set for a specific reasoning style; wherein, the reasoning style is supervised fine-tuning training data that is differentiated in terms of the level of detail of the reasoning process, logical structure, or language style.

[0047] It should be noted that the selection device 200 for supervised fine-tuning training data of the job matching model provided in this embodiment corresponds to the technical solution that can be used to execute various method embodiments. Its implementation principle and technical effect are similar to the method, and will not be repeated here.

[0048] See Figure 3 Another embodiment of the present invention provides a schematic diagram of an electronic device 300, which is used to implement the method for selecting supervised fine-tuning training data of the person-job matching model in the method embodiment. The electronic device 300 in the embodiments of the present invention may include, but is not limited to, a PC, a server, a laptop computer, and a smart terminal. Figure 3 The electronic device 300 shown is merely an example and should not be construed as limiting the functionality and scope of use of the embodiments of the present invention.

[0049] like Figure 3 As shown, the electronic device 300 may include a processing device (e.g., a central processing unit, a graphics processing unit, etc.) 301, which can perform various appropriate actions and processes to implement the methods of the embodiments described herein, based on a program stored in a read-only memory (ROM) 302 or a program loaded from a storage device 308 into a random access memory (RAM) 303. The RAM 303 also stores various programs and data required for the operation of the electronic device 300. The processing device 301, ROM 302, and RAM 303 are interconnected via a bus 304. An input / output (I / O) interface 305 is also connected to the bus 304.

[0050] Typically, the following devices can be connected to I / O interface 305: input devices 306 including, for example, touchscreens, touchpads, keyboards, mice, cameras, microphones, accelerometers, gyroscopes, etc.; output devices 307 including, for example, liquid crystal displays (LCDs), speakers, vibrators, etc.; storage devices 308 including, for example, magnetic tapes, hard disks, etc.; and communication devices 309. Communication device 309 allows electronic device 300 to communicate wirelessly or wiredly with other devices to exchange data. Although Figure 3 An electronic device 300 with various devices is shown; however, it should be understood that it is not required to implement or possess all of the devices shown. More or fewer devices may be implemented or possessed alternatively.

[0051] The above description is merely a preferred embodiment of the present invention. Those skilled in the art should understand that the scope of disclosure in this invention is not limited to the specific combination of the above-described technical features, but should also cover other technical solutions formed by any combination of the above-described technical features or their equivalents without departing from the above-described concept. For example, technical solutions formed by substituting the above features with (but not limited to) technical features with similar functions disclosed in this invention.

Claims

1. A method for selecting supervised fine-tuning training data for a person-job matching model, characterized in that, Includes the following steps: The difficulty of matching job description-candidate resume pairs is evaluated from multiple dimensions using an artificial intelligence model. A difficulty score is obtained for each dimension, and the difficulty scores for each dimension are weighted and summed to obtain the initial difficulty score for the job description-candidate resume pair. Obtain matching tags generated based on the actual behavior of recruitment decision-makers, including positive and negative feedback actions on candidate resumes; The job description-candidate resume pair is initially matched and judged by an artificial intelligence model to obtain an initial matching result. If the initial matching result is consistent with the matching tag, the initial difficulty score is downweighted. If the initial matching result is inconsistent with the matching tag, the initial difficulty score is upweighted to obtain the comprehensive difficulty score of the job description-candidate resume pair. Based on the comprehensive difficulty score, a supervised fine-tuning training set for a specific reasoning style is constructed by selecting job descriptions and candidate resumes that match the difficulty level; wherein, the reasoning style is supervised fine-tuning training data that differs in the level of detail of the reasoning process, logical structure, or language style.

2. The method for selecting supervised fine-tuning training data for a person-job matching model according to claim 1, characterized in that, Based on the comprehensive difficulty score, the steps for constructing a supervised fine-tuning training set for a specific reasoning style, including selecting job descriptions and candidate resumes that match the difficulty level, are as follows: Keeping the pre-trained model, model hyperparameters, and test set unchanged, multiple training sets are constructed using job description-candidate resume pairs with different difficulty levels for the same reasoning style. The model is then fine-tuned under supervision, and the performance metrics of each fine-tuned model are evaluated on a unified test set. The difficulty level of the training set corresponding to the model with the best performance metrics is determined as the difficulty level that is suitable for the reasoning style. Based on the specific reasoning style used in the current training task, job descriptions and candidate resumes whose overall difficulty scores fall within the corresponding difficulty level range are selected to form a supervised fine-tuning training set for constructing the specific reasoning style.

3. The method for selecting supervised fine-tuning training data for a person-job matching model according to claim 1, characterized in that, The steps of reducing the weight of the initial difficulty score if the preliminary matching result matches the matching label, and increasing the weight of the initial difficulty score if the preliminary matching result does not match the matching label, include: If the preliminary matching result is consistent with the matching label, the initial difficulty score is multiplied by a weight coefficient less than 1; if the preliminary matching result is inconsistent with the matching label, the initial difficulty score is multiplied by a weight coefficient greater than 1.

4. The method for selecting supervised fine-tuning training data for a person-job matching model according to claim 1, characterized in that, The dimensions include the difficulty of judging explicit information, the difficulty of judging implicit information, and the difficulty of judging fuzzy information; wherein, the weight of the difficulty of judging implicit information is greater than that of judging explicit information, and the weight of the difficulty of judging explicit information is greater than that of judging fuzzy information.

5. The method for selecting supervised fine-tuning training data for a person-job matching model according to claim 4, characterized in that, The difficulty in judging implicit information includes at least one of the following: the difficulty of transferring the job seeker's skills, the difficulty of relating the job seeker's achievements, and the difficulty of completing the job seeker's work experience.

6. The method for selecting supervised fine-tuning training data for a person-job matching model according to claim 1, characterized in that, The positive feedback operation is to download candidate resumes, and the negative feedback operation is to mark candidate resumes as unsuitable.

7. The method for selecting supervised fine-tuning training data for a person-job matching model according to claim 1, characterized in that, Also includes: Based on the comprehensive difficulty score, job description-candidate resume pairs with different appropriate difficulty levels are selected to construct training sets for the models to be trained with different parameter scales.

8. A device for selecting supervised fine-tuning training data for a person-job matching model, characterized in that, include: The model scoring module is used to evaluate the difficulty of matching job description-candidate resume pairs in multiple dimensions using an artificial intelligence model, obtain a difficulty score for each dimension, and then perform a weighted summation of the difficulty scores for each dimension to obtain the initial difficulty score for the job description-candidate resume pair. The scoring calibration module is used to obtain matching tags generated based on the actual behavior of the hiring decision-makers, including positive feedback and negative feedback on candidate resumes. The job description-candidate resume pair is initially matched and judged by an artificial intelligence model to obtain an initial matching result. If the initial matching result is consistent with the matching tag, the initial difficulty score is downweighted. If the initial matching result is inconsistent with the matching tag, the initial difficulty score is upweighted to obtain the comprehensive difficulty score of the job description-candidate resume pair. The training data filtering module is used to select job descriptions and candidate resumes that match the difficulty level based on the comprehensive difficulty score to construct a supervised fine-tuning training set for a specific reasoning style; wherein, the reasoning style is supervised fine-tuning training data that is differentiated in terms of the level of detail of the reasoning process, logical structure, or language style.

9. The device for selecting supervised fine-tuning training data for a person-job matching model according to claim 8, characterized in that, The training data filtering module is specifically used for: Keeping the pre-trained model, model hyperparameters, and test set unchanged, multiple training sets are constructed using job description-candidate resume pairs with different difficulty levels for the same reasoning style. The model is then fine-tuned under supervision, and the performance metrics of each fine-tuned model are evaluated on a unified test set. The difficulty level of the training set corresponding to the model with the best performance metrics is determined as the difficulty level that is suitable for the reasoning style. Based on the specific reasoning style used in the current training task, job descriptions and candidate resumes whose overall difficulty scores fall within the corresponding difficulty level range are selected to form a supervised fine-tuning training set for constructing the specific reasoning style.

10. The device for selecting supervised fine-tuning training data for a person-job matching model according to claim 8, characterized in that, The dimensions include the difficulty of judging explicit information, the difficulty of judging implicit information, and the difficulty of judging fuzzy information; wherein, the weight of the difficulty of judging implicit information is greater than that of judging explicit information, and the weight of the difficulty of judging explicit information is greater than that of judging fuzzy information.

Citation Information

Patent Citations

  • Intelligent auditing method, auditing system, electronic equipment and storage medium

    CN117764540A

  • Position resume intelligent matching method and system based on multi-dimensional analysis

    CN119313304A

  • Myopia occurrence risk assessment method based on multi-source data

    CN120727301A

  • Apparatus and methods for customization and utilization of target profiles

    US20250054068A1

  • Automated Talent Acquisition and Management System Using AI

    US20250232262A1