Resume analysis method and system based on multiple large language models
By employing parallel parsing of multiple language models and a conflict re-decision mechanism, the problem of insufficient accuracy in parsing by a single model is solved, achieving high-precision resume parsing and system robustness, adapting to various resume formats, and facilitating industrial applications.
Patent Information
- Application Number
- CN202511178381.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-22
- Publication Date
- 2025-11-21
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Existing resume parsing systems rely on a single large language model, which results in insufficient parsing accuracy. In particular, they are prone to misjudgment or omission in key fields, making it difficult to meet the needs of high-precision recruitment.
We employ parallel parsing using multiple language models, integrate the final results through consistency checks and conflict re-decision mechanisms, and combine them with a large language model for adjudication. This dynamically optimizes the model pool and improves parsing accuracy and robustness.
It significantly improves parsing accuracy, enhances system robustness, can identify and correct erroneous results of individual models, adapts to various resume formats, and is easy to industrialize.
Smart Images

Figure CN120996024A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of large language model technology, specifically involving a resume parsing method and system based on multiple language models. Background Technology
[0002] Existing resume parsing systems typically rely on a single large language model to extract resume information. Due to limitations in the model's generalization ability and the diversity of training data, the accuracy of parsing is generally insufficient. In particular, key fields such as actual expertise, job functions, and industry experience are prone to misjudgment or omission, making it difficult to meet the needs of high-precision recruitment applications. Summary of the Invention
[0003] The technical problem to be solved by the present invention is to overcome the shortcomings of the prior art and provide a resume parsing method and system based on multiple language models.
[0004] The technical solution adopted to solve the above technical problems is: a resume parsing method based on a multi-language model, including the following specific steps: Step 1: Parallel Model Parsing Step: Input an unstructured resume text data to be parsed into at least two heterogeneous large language models simultaneously through the application programming interface; each large language model independently extracts and parses information according to its own model structure and training data, and outputs structured data results containing multiple preset fields respectively. Step 2: Consistency Detection Step: Align the structured data output by all large language models in Step 1. For each preset field, compare the field values output by different models one by one. If all models output the same value for a certain field, then the consistent value is determined as the final parsed value of that field. Step 3: Conflict re-decision step: If in Step 2 it is found that the output values of different large language models are inconsistent for a certain preset field, the conflict resolution mechanism is triggered; Step 4: Result Fusion and Output Step: The final parsed values of all conflict-free fields that reached a consensus in Step 2 are integrated with the optimal values of all conflicting fields determined by the judge's big language model in Step 1 to generate a complete, unified, high-precision structured resume data as the final output of the system. Step 5: Dynamic Optimization of the Model Pool: During continuous operation, the system automatically collects and analyzes the output results of each large language model in multiple parsing tasks; by comparing the output of each model with the final result after the judgment in Step 3, the system calculates the performance indicators of each model in terms of accuracy, recall, or F1 score; based on the historical data of these performance indicators, the system periodically performs dynamic scoring and ranking of each model in the model pool; according to the scoring results, the model with the worst performance is automatically removed from the currently available model pool, and a new large language model is simultaneously introduced into the model pool.
[0005] The above technical solutions significantly improve the accuracy and reliability of the analysis and enhance the robustness of the system.
[0006] Furthermore, the conflict resolution mechanism includes: automatically aggregating all candidate values of the conflict field, generating the original query command and complete response context of the large language model corresponding to each candidate value, and the text segment in the original resume text where the conflict field is located and its surrounding context information; The aggregated information is used as input and submitted to a large language model for adjudication. The large language model performs comprehensive analysis, reasoning and judgment on the input information, outputs the optimal value of the conflict field, and provides a confidence score that reflects the degree of certainty of its judgment.
[0007] Through the above technical solution, even if one or more models output incorrect results due to their own defects, training data bias, or temporary malfunctions, the system can identify and correct them through subsequent consistency detection and conflict resolution mechanisms. The overall performance of the system does not depend on the worst performance of a single model, but rather on the collective wisdom of the models and the highest level of the arbitration model, making it extremely resistant to interference.
[0008] Furthermore, the large language model includes the Deepseek series model, the GPT series model, the Claude series model, the Wenxin Yiyan series model, the Tongyi Qianwen series model, the ChatGLM series model, and the LLaMA series model.
[0009] Furthermore, the referee's large language model in step three includes: A standalone, general-purpose enhanced large language model that did not participate in the initial parallel parsing; A dedicated model for conflict adjudication, finely trained based on historical judgment data; Select the highest-performing model instance from the multiple large language models that participated in the initial parsing in Step 1.
[0010] The above technical solutions provide multiple implementation methods, allowing implementers to choose according to their own resources and technical conditions.
[0011] Furthermore, the preset fields include the following information categories: basic personal information, contact information, work experience information, skills and abilities information, educational background information, and salary and benefits information.
[0012] The above technical solution can output a very complete and in-depth candidate profile, meeting the requirements for data granularity in modern recruitment.
[0013] Furthermore, the basis for introducing new models in step five includes: introducing new models based on public evaluation rankings, introducing new models based on performance on specific internal test sets, and introducing new models based on parsing quality scores from user feedback.
[0014] A system for a resume parsing method based on multiple language models includes a data receiving and distribution module, a multi-model parsing scheduling module, a consistency detection and conflict identification module, a conflict re-decision scheduling module, a data fusion and output module, and a model performance monitoring and optimization module. The data receiving and distribution module is responsible for receiving resume data submitted by users and distributing it in parallel to multiple large language model parsing services; the multi-model parsing scheduling module is responsible for managing and scheduling a model pool containing multiple heterogeneous large language models, sending parsing requests to each model and receiving the structured results returned by each model.
[0015] Furthermore, the consistency detection and conflict identification module is used to receive the results returned by each model and perform field-level consistency comparison to identify fields with output conflicts; the conflict re-decision scheduling module is used to aggregate the context information required for the judgment of the identified conflict fields and call the specified referee big language model to perform arbitration judgment; the data fusion and output module is used to integrate the values of non-conflicting fields and the values of conflicting fields after arbitration, assemble them into the final structured resume data and output it; the model performance monitoring and optimization module is used to continuously monitor the parsing performance of each model in the model pool, and perform model scoring, elimination and new model introduction operations according to preset strategies to realize dynamic life cycle management of the model pool.
[0016] The above technical solutions enable the system to have high scalability and versatility, and adapt to different environments.
[0017] The beneficial effects of this invention are as follows: This invention effectively reduces the risk of misjudgment by a single model through multi-model parallel parsing and conflict re-decision mechanism, improves the overall system robustness, significantly improves the accuracy of resume parsing, and realizes automatic optimization and updating of the model pool through dynamic scoring mechanism, solving the problem of difficult model selection and updating. The system has strong versatility and scalability, can adapt to resume data of various types and formats, and is easy to industrialize. Attached Figure Description
[0018] Figure 1 This is a flowchart of the method of the present invention. Detailed Implementation
[0019] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the invention.
[0020] like Figure 1 As shown in this embodiment, a resume parsing method and system based on a multi-language model includes the following specific steps: Step 1: Parallel Model Parsing Step: Input an unstructured resume text data to be parsed into at least two heterogeneous large language models simultaneously through the application programming interface; each large language model independently extracts and parses information according to its own model structure and training data, and outputs structured data results containing multiple preset fields respectively. Step 2: Consistency Detection Step: Align the structured data output by all large language models in Step 1. For each preset field, compare the field values output by different models one by one. If all models output the same value for a certain field, then the consistent value is determined as the final parsed value of that field. Step 3: Conflict re-decision step: If in Step 2 it is found that the output values of different large language models are inconsistent for a certain preset field, the conflict resolution mechanism is triggered; Step 4: Result Fusion and Output Step: The final parsed values of all conflict-free fields that reached a consensus in Step 2 are integrated with the optimal values of all conflicting fields determined by the judge's big language model in Step 1 to generate a complete, unified, high-precision structured resume data as the final output of the system. Step 5: Dynamic Optimization of the Model Pool: During continuous operation, the system automatically collects and analyzes the output results of each large language model in multiple parsing tasks. By comparing the output of each model with the final result after the decision in Step 3, the system calculates the performance metrics of each model in terms of accuracy, recall, or F1 score. Based on the historical data of these performance metrics, the system periodically performs dynamic scoring and ranking of each model in the model pool. According to the scoring results, the worst-performing model is automatically removed from the currently available model pool, and new large language models are introduced into the model pool, significantly improving parsing accuracy and reliability, and enhancing the robustness of the system.
[0021] The conflict resolution mechanism includes: automatically aggregating all candidate values of the conflict field, generating the original query command and complete response context of the large language model corresponding to each candidate value, and the text segment in the original resume text where the conflict field is located and its surrounding context information; The aggregated information is used as input and submitted to a dedicated adjudication language model. This model comprehensively analyzes, reasons, and judges the input information, outputting the optimal value for the conflict field and providing a confidence score reflecting the degree of certainty in its judgment. Even if one or more models output incorrect results due to their own defects, training data bias, or temporary malfunctions, the system can identify and correct them through subsequent consistency checks and conflict resolution mechanisms. The overall system performance does not depend on the worst performance of a single model, but rather on the collective wisdom of the models and the highest level of the arbitration model, exhibiting extremely strong anti-interference capabilities.
[0022] The large language model includes the Deepseek series model, the GPT series model, the Claude series model, the Wenxin Yiyan series model, the Tongyi Qianwen series model, the ChatGLM series model, and the LLaMA series model.
[0023] The large language model for the judges in step three includes: A standalone, general-purpose enhanced large language model that did not participate in the initial parallel parsing; A dedicated model for conflict adjudication, finely trained based on historical judgment data; The most powerful model instance selected from the multiple large language models that participated in the initial parsing in Step 1 is presented, along with several implementation schemes, allowing implementers to choose according to their own resources and technical conditions.
[0024] The preset fields include the following information categories: basic personal information, contact information, work experience information, skills and abilities information, educational background information, and salary and benefits information. They can output a very complete and in-depth candidate profile, meeting the requirements for data granularity in modern recruitment.
[0025] The basis for introducing new models in step five includes: introducing new models based on public evaluation rankings, introducing new models based on performance on specific internal test sets, and introducing new models based on parsing quality scores from user feedback.
[0026] A system for a resume parsing method based on multiple language models includes a data receiving and distribution module, a multi-model parsing scheduling module, a consistency detection and conflict identification module, a conflict re-decision scheduling module, a data fusion and output module, and a model performance monitoring and optimization module. The data receiving and distribution module is responsible for receiving resume data submitted by users and distributing it in parallel to multiple large language model parsing services; the multi-model parsing scheduling module is responsible for managing and scheduling a model pool containing multiple heterogeneous large language models, sending parsing requests to each model and receiving the structured results returned by each model.
[0027] The consistency detection and conflict identification module receives the results returned by each model and performs field-level consistency comparison to identify fields with output conflicts. The conflict re-decision scheduling module aggregates the context information required for the decision for the identified conflict fields and calls the specified referee big language model for arbitration. The data fusion and output module integrates the values of non-conflicting fields and the values of conflicting fields after arbitration, assembles them into the final structured resume data, and outputs it. The model performance monitoring and optimization module continuously monitors the parsing performance of each model in the model pool, performs model scoring, elimination, and new model introduction operations according to preset strategies, realizes dynamic lifecycle management of the model pool, and makes the system highly scalable and versatile, adaptable to different environments.
[0028] The above description is merely a preferred embodiment of the present invention and is not intended to limit the scope of protection of the present invention.
Claims
1. A resume parsing method based on a multi-language model, characterized in that, The specific steps include the following: Step 1: Parallel Model Parsing Step: Input an unstructured resume text data to be parsed into at least two heterogeneous large language models simultaneously through the application programming interface; each large language model independently extracts and parses information according to its own model structure and training data, and outputs structured data results containing multiple preset fields respectively. Step 2: Consistency Detection Step: Align the structured data output by all large language models in Step 1, and compare the field values output by different models for each of the preset fields. If all models output the same value for a certain field, then that consistent value is determined as the final parsed value for that field. Step 3: Conflict re-decision step: If in Step 2 it is found that the output values of different large language models are inconsistent for a certain preset field, the conflict resolution mechanism is triggered; Step 4: Result Fusion and Output Step: The final parsed values of all conflict-free fields that reached a consensus in Step 2 are integrated with the optimal values of all conflicting fields determined by the judge's big language model in Step 1 to generate a complete, unified, high-precision structured resume data as the final output of the system. Step 5: Dynamic Optimization of the Model Pool: During continuous operation, the system automatically collects and statistically analyzes the output results of each large language model in multiple parsing tasks; by comparing the output of each model with the final result after the decision in Step 3, the system calculates the performance metrics of each model in terms of accuracy, recall, or F1 score; based on the historical data of these performance metrics, the system periodically performs dynamic scoring and ranking of each model in the model pool. Based on the scoring results, the worst-performing model is automatically removed from the currently available model pool, and a new large language model is added to the model pool at the same time.
2. The resume parsing method based on a multi-language model according to claim 1, characterized in that, The conflict resolution mechanism includes: automatically aggregating all candidate values of the conflict field, generating the original query command and complete response context of the large language model corresponding to each candidate value, and the text segment in the original resume text where the conflict field is located and its surrounding context information; The aggregated information is used as input and submitted to a large language model for adjudication. The large language model performs comprehensive analysis, reasoning and judgment on the input information, outputs the optimal value of the conflict field, and provides a confidence score that reflects the degree of certainty of its judgment.
3. The resume parsing method based on a multi-language model according to claim 2, characterized in that, The large language model includes the Deepseek series model, the GPT series model, the Claude series model, the Wenxin Yiyan series model, the Tongyi Qianwen series model, the ChatGLM series model, and the LLaMA series model.
4. The resume parsing method based on a multi-language model according to claim 3, characterized in that, The large language model for the judges in step three includes: A standalone, general-purpose enhanced large language model that was not involved in the initial parallel parsing; A dedicated model for conflict adjudication, finely trained based on historical judgment data; Select the highest-performing model instance from the multiple large language models that participated in the initial parsing in Step 1.
5. The resume parsing method based on a multi-language model according to claim 4, characterized in that, The preset fields include the following information categories: basic personal information, contact information, work experience information, skills and abilities information, educational background information, and salary and benefits information.
6. The resume parsing method based on a multi-language model according to claim 5, characterized in that, The basis for introducing new models in step five includes: introducing new models based on public evaluation rankings, introducing new models based on performance on specific internal test sets, and introducing new models based on parsing quality scores from user feedback.
7. The system for a resume parsing method based on a multi-language model according to claim 6, characterized in that, It includes a data receiving and distribution module, a multi-model parsing and scheduling module, a consistency detection and conflict identification module, a conflict re-decision scheduling module, a data fusion and output module, and a model performance monitoring and optimization module; The data receiving and distribution module is responsible for receiving resume data submitted by users and distributing it in parallel to multiple large language model parsing services; the multi-model parsing scheduling module is responsible for managing and scheduling a model pool containing multiple heterogeneous large language models, sending parsing requests to each model and receiving the structured results returned by each model.
8. The system for a resume parsing method based on a multi-language model according to claim 7, characterized in that, The consistency detection and conflict identification module receives the results returned by each model and performs field-level consistency comparison to identify fields with output conflicts. The conflict re-decision scheduling module aggregates the context information required for the decision for the identified conflict fields and calls the specified referee big language model for arbitration. The data fusion and output module integrates the values of non-conflicting fields and the values of conflicting fields after arbitration, assembles them into the final structured resume data, and outputs it. The model performance monitoring and optimization module continuously monitors the parsing performance of each model in the model pool, performs model scoring, elimination, and new model introduction operations according to preset strategies, and realizes dynamic lifecycle management of the model pool.
Citation Information
Cited By
Knowledge graph query method and device based on multi-language model result fusion
CN121636662A
Data processing method and device, equipment, storage medium and program product
CN121724032A
A data processing method, device, apparatus, storage medium, and program product
CN121724032B
Data verification and screening method and system for output results of multi-source large model
CN121936634A