Large model-based labeler hiring test evaluation system

CN122222560BActive Publication Date: 2026-09-04LINGBO WEIBU (BEIJING) TECHNOLOGY CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202610316315.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2026-03-16
Publication Date
2026-09-04
Estimated Expiration
2046-03-16

AI Technical Summary

Technical Problem

[0005]本发明旨在提供基于大模型的标注人员招聘测试评估系统,以解决现有技术中静态测试题库难以适应不同标注任务特性、无法动态调整评估策略、导致对标注人员真实技能水平评估准确性不足的技术矛盾

Benefits of technology

[0018]1.本发明通过任务特性解析模块实现了对标注任务多维特性的精准量化,为测试内容的动态生成提供了精确的输入条件。动态测试生成模块基于大语言模型与任务参数的深度结合,能够生成与真实工作场景高度一致的专项测试题目,有效避免了通用测试与具体任务脱节的问题。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122222560B_ABST
    Figure CN122222560B_ABST
Patent Text Reader

Abstract

The application relates to the technical field of artificial intelligence, and particularly discloses an annotation personnel recruitment test evaluation system based on a large model. The system comprises a task characteristic analysis module, a dynamic test generation module, a multi-modal interaction test module, a cognitive behavior analysis module and a comprehensive evaluation decision module. Through analyzing task characteristic parameters, the system dynamically generates adaptive test questions, collects and analyzes multi-modal operation behaviors of candidates, and performs weighted evaluation and recruitment recommendation by fusing multi-dimensional features. The application realizes the precision and self-adaptation of annotation personnel skill evaluation, and significantly improves the matching degree of recruitment test and post requirements.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of artificial intelligence technology, specifically relating to a recruitment, testing, and evaluation system for labeled personnel based on large models. Background Technology

[0002] The development of artificial intelligence technology relies heavily on large-scale, high-quality labeled data, and the recruitment and evaluation of professional labelers is a core aspect of ensuring data quality. Traditional recruitment processes typically employ a combination of standardized aptitude tests and interviews to comprehensively assess candidates' professional skills, comprehension abilities, and work attitude.

[0003] Intelligent assessment systems based on large models are increasingly being applied in the human resources field, aiming to improve recruitment efficiency through automated testing and data analysis. These systems quantify the basic annotation abilities of personnel by building a general test question bank, thereby providing a reference for recruitment decisions.

[0004] Existing technologies primarily rely on pre-set static test question banks, which are often designed for general annotation scenarios and struggle to adapt to the subtle differences in data formats, annotation standards, and domain knowledge across various annotation tasks. Due to a lack of dynamic awareness of task characteristics, the system cannot adjust its evaluation strategy according to specific task requirements, resulting in a high rate of misjudgment of the true skill level of annotation personnel. Furthermore, general tests fail to effectively identify the professional strengths or potential weaknesses of annotation personnel in specific task types, thus affecting the accurate matching of talent with positions. Therefore, how to construct an intelligent testing and evaluation system that can dynamically adapt to task characteristics and accurately assess the professional skills of annotation personnel has become a pressing technical challenge in this field. Summary of the Invention

[0005] This invention aims to provide a large-model-based recruitment testing and evaluation system for annotation personnel, in order to solve the technical contradictions in the existing technology where static test question banks are difficult to adapt to the characteristics of different annotation tasks, cannot dynamically adjust evaluation strategies, and result in insufficient accuracy in evaluating the true skill level of annotation personnel.

[0006] The technical solution of this invention is to construct a complete system capable of dynamically perceiving task characteristics, adaptively generating test content, and performing multi-dimensional in-depth evaluation. This system includes a task characteristic analysis module, a dynamic test generation module, a multimodal interactive testing module, a cognitive behavior analysis module, and a comprehensive evaluation and decision-making module.

[0007] The task feature analysis module performs in-depth analysis of the raw data of the target annotation task in the access system, extracting four core task dimension parameters: task domain features, data format specifications, annotation rule complexity, and quality acceptance criteria. The dynamic test generation module, built on a large language model, receives the four core task dimension parameters output from the task feature analysis module and generates specialized test questions and corresponding standard reference answer sets that are highly adapted to the target annotation task through parameterized conditions.

[0008] The multimodal interaction testing module presents candidates with specialized test questions output by the dynamic test generation module, supporting four interaction modes: text annotation, region selection, polygon drawing, and category label selection. It also records the candidate's complete operation sequence and timestamp data in real time. The cognitive behavior analysis module performs multi-granular analysis on the operation sequences recorded by the multimodal interaction testing module, extracting four types of behavioral feature vectors: operation accuracy indicators, operation efficiency indicators, rule compliance indicators, and abnormal behavior pattern indicators.

[0009] The comprehensive evaluation and decision-making module receives a set of standard reference answers from the dynamic test generation module and four types of behavioral feature vectors from the cognitive behavior analysis module. It calculates the candidate's comprehensive ability score through a weighted fusion algorithm and generates the final recruitment recommendation decision based on a preset job suitability threshold.

[0010] Furthermore, the specific parsing process executed by the task characteristic parsing module is as follows. First, statistical analysis is performed on the historical labeled sample set of the target labeling task to calculate the labeling category distribution entropy value and the labeling boundary ambiguity index. The two together constitute the task domain feature vector.

[0011] Secondly, the structured constraints in the annotation specification document are analyzed, extracting three data format specification parameters: annotation tool compatibility requirements, annotation file output format specifications, and annotation attribute field integrity rules. Next, natural language processing techniques are used to analyze the logical dependencies in the annotation rule document, constructing a rule dependency graph and calculating its node degree variance and average path length; these two parameters jointly characterize the annotation rule complexity. Finally, the error classification system and corresponding deduction weights in the quality acceptance standard document are analyzed, constructing a quality assessment matrix as a quantitative expression of the quality acceptance standard.

[0012] Furthermore, the dynamic test generation module's working mechanism includes the following steps. The large language model receives task domain feature vectors as semantic background conditions for generating test questions, data format specification parameters as constraints on the output structure, annotation rule complexity as a control parameter for question difficulty, and quality acceptance criteria as a reference benchmark for scoring. Based on the above four conditions, the large language model generates a question set containing 2-5 test questions. Each question includes a complete question stem description, interactive operation requirements, and a standard reference answer. The standard reference answer not only includes the final correct annotation result but also the sequence of key operation steps in the annotation process and warnings about common error types.

[0013] Furthermore, the interaction modes of the multimodal interaction testing module strictly correspond to the data types of the target annotation tasks. For image-based annotation tasks, two interaction modes are enabled: region selection and polygon drawing, with real-time capture of the candidate's coordinate positioning accuracy and contour fitting degree. For text-based annotation tasks, two interaction modes are enabled: text annotation and classification label selection, with recording of the candidate's entity boundary delineation accuracy and classification logic consistency. All interaction operations are recorded with millisecond-level timestamps, forming a complete behavior log including operation type, operation parameters, and operation sequence.

[0014] Furthermore, the multi-granularity parsing process executed by the cognitive behavior analysis module specifically includes the following steps: Operational accuracy is calculated by comparing the spatial overlap or semantic matching degree between candidate operation results and standard reference answers. Operational efficiency is quantified by analyzing the number of effective operations completed per unit time and the optimality of the operation path. Rule compliance is evaluated by detecting the number and severity of violations of annotation rule constraints in the operation sequence. Abnormal behavior pattern is identified by extracting features from three typical abnormal behaviors in the operation sequence: random clicks, frequent undoings, and long pauses.

[0015] Furthermore, the weighted fusion algorithm used in the comprehensive evaluation decision module is implemented as follows: A weight of 0.4 is assigned to the operational accuracy indicator, 0.25 to the operational efficiency indicator, 0.2 to the rule compliance indicator, and 0.15 to the abnormal behavior pattern indicator. The weighted sum is used to obtain the comprehensive ability score, which is normalized to 0-100. The system presets three tiers of job suitability thresholds: scores above 80 recommend advanced-labeled positions, scores between 60 and 80 recommend standard-labeled positions, and scores below 60 are not recommended. The recruitment recommendation decision, along with the detailed comprehensive ability score, is output to the recruitment management system.

[0016] Furthermore, the system also includes a continuous learning and optimization module. This module periodically collects annotation quality data from recruited annotators in real projects and performs correlation analysis with the evaluation results from the recruitment testing phase. Based on the analysis results, it dynamically adjusts the feature extraction strategy in the cognitive behavior analysis module and the weight allocation scheme in the comprehensive evaluation decision module, achieving continuous alignment between the system's evaluation model and actual job requirements.

[0017] Compared with the prior art, the beneficial effects of the present invention are as follows:

[0018] 1. This invention achieves precise quantification of the multidimensional characteristics of labeled tasks through a task characteristic analysis module, providing accurate input conditions for the dynamic generation of test content. The dynamic test generation module, based on a deep integration of a large language model and task parameters, can generate specialized test questions highly consistent with real-world work scenarios, effectively avoiding the problem of general testing being disconnected from specific tasks.

[0019] 2. The multimodal interaction testing module fully records the details of candidates' operational behaviors, providing a rich data foundation for in-depth analysis. The cognitive behavior analysis module extracts behavioral features from multiple dimensions, revealing the cognitive strategies and work habits hidden behind the candidates' test results. The comprehensive evaluation and decision-making module achieves a precise mapping from test data to recruitment decisions through scientific weighted fusion and threshold judgment.

[0020] 3. The entire system forms a complete technical closed loop from task understanding, test generation, behavior collection to comprehensive evaluation, which significantly improves the accuracy, relevance and matching degree with job requirements of the annotation personnel's skill assessment, and fundamentally solves the technical problem of insufficient adaptability and accuracy of traditional static testing systems. Attached Figure Description

[0021] Figure 1 This is a schematic diagram of the overall technical architecture of the labeling personnel recruitment testing and evaluation system based on a large model proposed in this invention;

[0022] Figure 2 This is a schematic diagram of the core principle framework for task characteristic analysis and dynamic test generation in this invention;

[0023] Figure 3 This is a logical flowchart of the multimodal interaction testing and cognitive behavior analysis in this invention;

[0024] Figure 4 This is a schematic diagram illustrating the weighted fusion algorithm principle of the comprehensive evaluation and decision-making module in this invention;

[0025] Figure 5 This is a schematic diagram of the interaction relationship and data flow between the continuous learning and optimization module and other modules of the system in this invention. Detailed Implementation

[0026] Please refer to the attached document. Figures 1 to 5 This embodiment details the technical implementation of a large-scale model-based labeled personnel recruitment testing and evaluation system. The system consists of a task characteristic analysis module, a dynamic test generation module, a multimodal interactive testing module, a cognitive behavior analysis module, a comprehensive evaluation and decision-making module, and a continuous learning and optimization module. These modules are interconnected through strictly defined data interfaces and logical flows, forming an automated evaluation closed loop from task input to recruitment decision output.

[0027] The task feature parsing module is the system's input, responsible for deep parsing of the raw data from target annotation tasks accessing the system. This module receives raw data including historical annotated sample sets, annotation specification documents, annotation rule documents, and quality acceptance standard documents. Internally, the module contains four parallel parsing subunits, each responsible for extracting four core task-dimensional parameters: task domain features, data format specifications, annotation rule complexity, and quality acceptance standards.

[0028] In the task domain feature extraction subunit, the system first loads the historical labeled sample set and performs statistical calculations on the category distribution of all labeled samples. The system calculates the label category distribution entropy value, which reflects the degree of category balance in the labeling task; a higher entropy value indicates that the task involves more categories and the distribution is more uniform. Simultaneously, the system calculates the label boundary ambiguity index, which is derived by analyzing the feature differences between adjacent category samples at the label boundary; a higher index indicates a more severe boundary ambiguity problem in the task.

[0029] The entropy value of the annotation category distribution and the ambiguity index of the annotation boundary together constitute a two-dimensional vector, serving as the quantitative output of the task domain features. In the data format specification extraction subunit, the system parses the structured constraints in the annotation specification document. The system extracts annotation tool compatibility requirements, specifically including supported annotation software versions, a list of plugin dependencies, and operating system environment configurations. The system extracts annotation file output format specifications, including file storage format, coordinate system standards, attribute field naming rules, and metadata encapsulation structure. The system extracts annotation attribute field integrity rules, defining the number of attributes each annotation object must contain, the range of optional attributes, and the dependencies between attributes.

[0030] These three parameters together constitute the data format specification vector. In the annotation rule complexity extraction subunit, the system uses natural language processing technology to parse the annotation rule document. The system constructs a rule dependency graph, where nodes represent individual annotation rules and edges represent logical dependencies between rules. The system calculates the node degree variance of the rule dependency graph, reflecting the degree of imbalance in connections between rules. The system calculates the average path length of the rule dependency graph, reflecting the depth of the rule chain.

[0031] Node degree variance and average path length together constitute the annotation rule complexity vector. In the quality acceptance standard extraction sub-unit, the system analyzes the error classification system in the quality acceptance standard document. The system constructs a quality assessment matrix, where rows correspond to error types, columns correspond to severity levels, and matrix element values ​​are the deduction weights for each error type at its corresponding severity level. This matrix serves as a quantitative expression of the quality acceptance standard. The task characteristic parsing module packages these four dimensions into a structured data package and transmits it to the dynamic test generation module.

[0032] Please refer to the attached document. Figure 2 The dynamic test generation module is built upon a large language model and receives structured data packets from the task feature parsing module. Internally, the module includes a parameter preprocessing unit, a large language model inference unit, and a question post-processing unit. The parameter preprocessing unit converts the input four-dimensional parameters into conditional control instructions understandable by the large language model. Task domain feature vectors are mapped to semantic background conditions to limit the topic scope and scenario features of the test questions. Data format specification parameters are converted into output structure constraints to ensure that the generated questions meet the tool and format requirements of the target task. Annotation rule complexity parameters are quantified as difficulty control parameters, directly affecting the number of nested rules and logical complexity within the questions. The quality acceptance criterion matrix is ​​parsed into a scoring reference benchmark to ensure that the scoring criteria of the standard reference answer are consistent with the actual project.

[0033] The large language model inference unit employs a pre-trained language model based on the Transformer architecture, with a model parameter scale of no less than tens of billions. The model runs under a specific instruction template, which explicitly requires generating a set of 2-5 test questions. Each question must include a complete stem description, interactive operation requirements, and a standard reference answer. The stem description must clearly define the annotation objects, annotation rules, and output requirements. The interactive operation requirements must explicitly specify the available interaction modes and their operational specifications.

[0034] The standard reference answer not only includes the final correct annotation result, but also the sequence of key operational steps in the annotation process and warnings about common error types. The question post-processing unit performs format validation and logical consistency checks on the raw content generated by the model. This unit ensures that all questions meet the constraints set by the parameter preprocessing unit and structures the standard reference answer into machine-readable JSON data. The dynamic test generation module transmits the final generated question set and the standard reference answer to the multimodal interactive testing module.

[0035] Please refer to the attached document. Figure 3The multimodal interaction testing module is responsible for presenting test questions to candidates and recording their interactions. This module includes a question rendering engine, an interaction capture engine, and a behavior log management unit. The question rendering engine automatically selects the interface layout and interactive components based on the question type. For image annotation tasks, the engine enables two interaction modes: region selection and polygon drawing. The region selection mode provides three basic shape tools: rectangle, circle, and ellipse, supporting operations such as coordinate dragging, size adjustment, and angle rotation. The polygon drawing mode provides functions such as continuous click to generate vertices, Bézier curve fitting, and boundary smoothing.

[0036] For text annotation tasks, the engine enables two interaction modes: text annotation and category tag selection. The text annotation mode offers four granularity options: character-level, word-level, phrase-level, and sentence-level, supporting visual markers such as highlighting, underlining, and background color changes. The category tag selection mode offers three interaction methods: single selection, multiple selection, and hierarchical selection, supporting tag search, filtering, and batch operations. The interaction capture engine records all user actions with millisecond-level precision.

[0037] Each operation record contains three core fields: operation type, operation parameters, and operation timestamp. The operation type field is encoded as a 4-digit integer, corresponding to four basic operations and their combinations: text annotation, region selection, polygon drawing, and category label selection. The operation parameter field stores the specific numerical value of the operation: a two-dimensional floating-point array for coordinate operations, a character offset for text operations, and a label identifier for category operations. The operation timestamp field records the precise number of milliseconds from the start of the test to the current operation.

[0038] The behavior log management unit receives data streams from the interaction capture engine in real time and performs preliminary preprocessing on the operation sequences. The unit verifies the logical rationality of the operation sequences, detects and marks obviously abnormal operations, such as coordinate points exceeding image boundaries, mismatched label types, and sequences that violate the operation order. The unit compresses the preprocessed behavior logs into binary format and transmits them to the cognitive behavior analysis module.

[0039] The cognitive behavior analysis module performs multi-granularity parsing on the received behavior logs, extracting four types of behavioral feature vectors. The module includes an operation accuracy analysis unit, an operation efficiency analysis unit, a rule compliance analysis unit, and an abnormal behavior pattern analysis unit. The operation accuracy analysis unit evaluates operation accuracy through two methods: spatial overlap calculation and semantic matching degree calculation.

[0040] For image annotation tasks, the unit calculates the overlap ratio between candidate annotated regions and standard reference answer regions, i.e., the intersection-union ratio (IU). For text annotation tasks, the unit uses an edit distance-based similarity algorithm to calculate the character-level matching degree between candidate annotated segments and standard reference answers. The operation efficiency analysis unit evaluates operation efficiency from both time and path dimensions.

[0041] The unit calculates the number of valid operations completed per unit time. A valid operation is defined as an operation that is not flagged as anomaly by the system and ultimately contributes to the annotation results. The unit analyzes the optimality of operation paths, quantifying this by comparing the difference between the actual operation path and the theoretically optimal path. The theoretically optimal path is defined by the sequence of key operation steps in the standard reference answer. The rule compliance analysis unit detects behaviors in operation sequences that violate annotation rule constraints. The unit loads the annotation rule set provided by the dynamic test generation module and checks each operation sequence against the rule requirements.

[0042] For each rule, the unit records the number of violations and a severity score. The severity score is divided into three levels based on the impact of the violation on the final annotation quality. The abnormal behavior pattern analysis unit identifies three typical abnormal behaviors. The unit identifies random click patterns using a clustering algorithm, characterized by uniform distribution of operation coordinates and no semantic relationship. The unit detects frequent undo patterns using a state machine, characterized by a continuous alternating sequence of operations and undoes. The unit detects long pause patterns using a time threshold, characterized by intervals between adjacent operations exceeding a preset 5-second threshold.

[0043] These four analysis units output operational accuracy indicators, operational efficiency indicators, rule compliance indicators, and abnormal behavior pattern indicators, which together constitute a set of behavioral feature vectors and are transmitted to the comprehensive evaluation and decision-making module.

[0044] Please refer to the attached document. Figure 4 The comprehensive evaluation and decision-making module receives a set of standard reference answers from the dynamic test generation module and a set of behavioral feature vectors from the cognitive behavior analysis module. Internally, the module includes a feature normalization unit, a weighted fusion calculation unit, and a decision output unit. The feature normalization unit maps the four behavioral feature indicators to a standardized range of 0 to 1. The operation accuracy indicator directly uses the raw values ​​of the intersection-union ratio or edit distance similarity. The operation efficiency indicator is normalized by comparing the candidate's efficiency value with the preset benchmark efficiency value. The rule compliance indicator is calculated by subtracting the percentage of violation scores from Formula 1.

[0045] The abnormal behavior pattern index is calculated by subtracting the percentage of abnormal behaviors from Formula 1. The weighted fusion calculation unit uses a linear weighted model to calculate the comprehensive capability score. The unit assigns a weight of 0.4 to the operational accuracy index, 0.25 to the operational efficiency index, 0.2 to the rule compliance index, and 0.15 to the abnormal behavior pattern index. The comprehensive capability score is calculated by weighted summation:

[0046]

[0047] in Represents overall ability score, This represents an indicator of operational accuracy. This represents an operational efficiency indicator. Indicators representing rule compliance This represents an indicator of abnormal behavior patterns. The calculated score is normalized to 0-100 points through a linear transformation. The decision output unit generates a recruitment recommendation decision based on a preset job suitability threshold. The system sets three gradient thresholds: a score of 80 or above generates a high-level labeled job recommendation instruction; a score of 60-80 generates a standard labeled job recommendation instruction; and a score below 60 generates a no-recommendation instruction. The decision output unit encapsulates the recruitment recommendation instruction and comprehensive ability score details into a structured message and transmits it to an external recruitment management system via an application programming interface.

[0048] Please refer to the attached document. Figure 5 The continuous learning and optimization module, as an adaptive component of the system, periodically executes the model optimization process. This module includes a data collection unit, a correlation analysis unit, and a parameter adjustment unit. Every 30 days, the data collection unit extracts annotation quality data from the recruitment management system, based on the work of recruited annotators in real projects. The quality data includes four indicators: annotation accuracy, annotation efficiency, rule violation rate, and project evaluation level.

[0049] The unit also retrieves the evaluation records of these personnel from the system database when they participated in the recruitment test. The correlation analysis unit calculates the correlation coefficient between the test evaluation results and actual work performance. The unit calculates the Pearson correlation coefficient between operational accuracy and annotation accuracy, the Spearman correlation coefficient between operational efficiency and annotation efficiency, the negative correlation between rule compliance and rule violation rate, and the grade correlation between abnormal behavior pattern and project evaluation level.

[0050] The parameter adjustment unit dynamically adjusts system parameters based on correlation analysis results. When the correlation coefficient between a certain behavioral characteristic and actual work performance remains below 0.6, the unit initiates a feature extraction strategy optimization process, adjusting the feature calculation method of the corresponding analysis unit in the cognitive behavior analysis module. When there is a significant deviation between the weight allocation and the actual importance, the unit initiates a weight optimization process, using the gradient descent algorithm to adjust the weight allocation scheme in the comprehensive evaluation decision module. Through this closed-loop feedback mechanism, the continuous learning optimization module ensures that the system evaluation model remains continuously aligned with actual job requirements.

[0051] During system implementation, each module is deployed on a distributed computing cluster. The task feature parsing module requires at least 8 CPU cores and 16 gigabytes of memory to process large-scale labeled sample sets. The dynamic test generation module requires a dedicated graphics processing unit (GPU) accelerator card with at least 16 gigabytes of video memory to support efficient inference for large language models.

[0052] The multimodal interaction testing module adopts a front-end and back-end separation architecture. The front-end is implemented based on web browser technology, while the back-end service requires load balancing and high concurrency processing capabilities. The cognitive behavior analysis module adopts a streaming processing architecture, capable of processing millisecond-level behavior log data in real time. The comprehensive evaluation and decision-making module is deployed as a highly available microservice to ensure the real-time performance and reliability of recruitment decisions. The continuous learning and optimization module executes automatically according to a preset cycle, requiring no manual intervention.

[0053] The entire system communicates between modules through a unified service bus and employs a two-way authentication mechanism based on digital certificates to ensure secure data transmission. The system provides standardized application programming interfaces (APIs) and supports seamless integration with mainstream recruitment management systems, human resource systems, and annotation project management platforms.

[0054] In a specific deployment case, the complete workflow of the system for handling image object detection and annotation tasks is as follows: The task characteristic analysis module analyzes the historical sample set of the task, calculating an annotation category distribution entropy of 2.3 and an annotation boundary ambiguity index of 0.7. After parsing the annotation specifications, it is determined that a rectangular selection tool should be used, with the output format being Pascal's visual object classification standard, and the attribute fields required to include three items: object category, occlusion degree, and truncation degree. After analyzing the annotation rule document, the node degree variance is calculated to be 1.8, and the average path length is 4.2. The quality assessment matrix extracted from the quality acceptance criteria contains 5 error types and 3 severity levels.

[0055] The dynamic test generation module generates three test questions based on these parameters: the first question requires labeling vehicles in a traffic scene image with rectangular bounding boxes; the second question requires special processing of occluded vehicles; and the third question requires identifying and labeling vehicles with blurred boundaries. The multimodal interactive test module presents these questions to candidates, who use the rectangular bounding box selection tool to perform operations. The system records all bounding box coordinates, adjustment operations, and time series.

[0056] The cognitive behavior analysis module analyzed the recorded data and calculated the operational accuracy index to be 0.85, the operational efficiency index to be 0.72, the rule compliance index to be 0.9, and the abnormal behavior pattern index to be 0.95. The comprehensive evaluation and decision-making module calculated the comprehensive ability score to be 76.5 points and generated a standard-labeled job recommendation instruction based on the threshold. The continuous learning and optimization module collected the candidate's actual work data in subsequent cycles and found that the correlation between the operational efficiency index and actual work efficiency was low. Therefore, the optimization process of the operational efficiency analysis unit was initiated, and the criteria for judging effective operations and the calculation method for path optimality were adjusted.

[0057] The system employs a multi-layered security protection mechanism at the data processing level. All labeled data entering the system is anonymized before storage, removing personally identifiable information and sensitive business information. Behavioral log data is encrypted using Advanced Encryption Standard (AES) algorithms during transmission and stored in a distributed, fragmented manner. Model parameters and weight files are protected using Digital Rights Management (DRM) technology to prevent unauthorized access and copying.

[0058] The system access control employs role-based access management, distinguishing between three roles: system administrator, recruitment manager, and candidate. Each role can only access functions and data within its authorized scope. The system audit module records all critical operations, including parameter modifications, model updates, and decision queries, maintaining complete operation logs for auditing purposes.

[0059] The system performance has undergone rigorous testing. Under standard hardware configuration, the task characteristic analysis module processes typical annotation tasks in an average of 12 seconds. The dynamic test generation module generates a question set in an average of 8 seconds. The multimodal interactive testing module's interaction response time is less than 100 milliseconds. The cognitive behavior analysis module processes complete test behavior logs in an average of 5 seconds. The comprehensive evaluation decision module generates recruitment decisions in an average of 1 second. The system supports concurrent processing of testing and evaluation tasks for 100 candidates, maintaining resource utilization below 75% and ensuring stable system operation. Through this highly integrated, automated, and intelligent technology, the system of this invention can accurately, efficiently, and fairly complete the entire process of recruitment testing and evaluation for annotation personnel.

Claims

1. A recruitment testing and evaluation system for labeled personnel based on a large model, characterized in that, include: The task feature analysis module is used to perform in-depth analysis of the original data of the target annotation task in the system, and extract four core task dimension parameters: task domain features, data format specifications, annotation rule complexity, and quality acceptance standards. The dynamic test generation module, built on a large language model, receives four core task dimension parameters from the task feature parsing module and generates specialized test questions and corresponding standard reference answer sets that are highly adapted to the target annotation task through parameterized conditions. The multimodal interactive testing module presents candidates with specialized test questions output by the dynamic test generation module. It supports four interaction modes: text annotation, region selection, polygon drawing, and category label selection, and records the candidate's complete operation sequence and timestamp data in real time. The cognitive behavior analysis module performs multi-granular analysis on the operation sequences recorded by the multimodal interaction testing module, and extracts four types of behavioral feature vectors: operation accuracy index, operation efficiency index, rule compliance index, and abnormal behavior pattern index. The comprehensive evaluation and decision-making module receives a set of standard reference answers from the dynamic test generation module and four types of behavioral feature vectors from the cognitive behavior analysis module. It calculates the candidate's comprehensive ability score through a weighted fusion algorithm and generates the final recruitment recommendation decision based on a preset job suitability threshold. The task characteristic analysis module includes: The task domain feature extraction unit is used to perform statistical analysis on the historical labeled sample set of the target labeling task, calculate the labeling category distribution entropy value and the labeling boundary ambiguity index, which together constitute the task domain feature vector. The data format specification extraction unit is used to parse the structured constraints in the annotation specification document and extract three data format specification parameters: annotation tool compatibility requirements, annotation file output format specifications, and annotation attribute field integrity rules. The annotation rule complexity extraction unit is used to parse the logical dependencies in the annotation rule document using natural language processing technology, construct the rule dependency graph and calculate its node degree variance and average path length, which together characterize the annotation rule complexity. The quality acceptance standard extraction unit is used to analyze the error classification system and corresponding deduction weights in the quality acceptance standard document, and to construct a quality assessment matrix as a quantitative expression of the quality acceptance standard. The cognitive behavior analysis module includes: The operation accuracy analysis unit calculates the accuracy by comparing the spatial overlap or semantic matching degree between the candidate's operation results and the standard reference answer. The operation efficiency analysis unit quantifies the efficiency by analyzing the number of effective operations completed per unit time and the optimality of the operation path. The rule compliance analysis unit evaluates the rule compliance by detecting the number and severity of violations of annotation rule constraints in the operation sequence; The abnormal behavior pattern analysis unit extracts features by identifying three typical abnormal behaviors in the operation sequence: random clicks, frequent undoings, and long pauses.

2. The large-model-based labeled personnel recruitment testing and evaluation system according to claim 1, characterized in that, The dynamic test generation module includes: The parameter preprocessing unit is used to convert the four core task dimension parameters into conditional control instructions that the large language model can understand. The large language model reasoning unit is used to generate a set of 2-5 test questions based on the conditional control instructions output by the parameter preprocessing unit. Each question is accompanied by a complete question stem description, interactive operation requirements and standard reference answer. The question post-processing unit is used to perform format validation and logical consistency checks on the original content generated by the model, and to encapsulate the standard reference answer in a structured manner.

3. The labeling personnel recruitment testing and evaluation system based on a large model according to claim 2, characterized in that, The conditional control instruction mapping process executed by the parameter preprocessing unit is as follows: Map the task domain feature vectors to semantic background conditions for generating test questions; Convert data format specification parameters into constraints for the output structure; The complexity of annotation rules is quantified into a control parameter for the difficulty of the questions; The quality acceptance standards are analyzed as a reference benchmark for scoring.

4. The labeling personnel recruitment testing and evaluation system based on a large model according to claim 1, characterized in that, The multimodal interaction testing module includes: The question rendering engine automatically selects the interface layout and interactive components based on the question type. The interaction capture engine records all user operations with millisecond-level precision. Each operation record contains three core fields: operation type, operation parameters, and operation timestamp. The behavior log management unit receives data streams from the interaction capture engine in real time, performs preliminary preprocessing on the operation sequences, and transmits the preprocessed behavior logs to the cognitive behavior analysis module.

5. The large-model-based labeled personnel recruitment testing and evaluation system according to claim 4, characterized in that, The task rendering engine activates the corresponding interaction mode based on the data type of the target annotation task: For image annotation tasks, two interaction modes are enabled: region selection and polygon drawing. For text annotation tasks, two interaction modes are enabled: text annotation and category label selection.

6. The large-model-based labeling personnel recruitment testing and evaluation system according to claim 1, characterized in that, The specific implementation of the weighted fusion algorithm used in the comprehensive evaluation and decision-making module is as follows: Assign a weight of 0.4 to the operational accuracy indicator, a weight of 0.25 to the operational efficiency indicator, a weight of 0.2 to the rule compliance indicator, and a weight of 0.15 to the abnormal behavior pattern indicator. The comprehensive ability score is obtained by weighted summation, and the score range is normalized to 0-100. Recruitment recommendation decisions are generated based on three preset job suitability thresholds: those scoring 80 or above are recommended for advanced-labeled positions, those scoring 60-80 are recommended for standard-labeled positions, and those scoring below 60 are not recommended.

7. The large-model-based labeled personnel recruitment testing and evaluation system according to claim 1, characterized in that, It also includes a continuous learning and optimization module, which regularly collects annotation quality data from recruited annotators in real projects, performs correlation analysis with the evaluation results of the recruitment test phase, and dynamically adjusts the feature extraction strategy in the cognitive behavior analysis module and the weight allocation scheme in the comprehensive evaluation decision module based on the analysis results.

8. The labeling personnel recruitment testing and evaluation system based on a large model according to claim 7, characterized in that, The continuous learning optimization module includes: The data collection unit is used to periodically extract real work quality data of recruited and marked personnel from the recruitment management system; The correlation analysis unit is used to calculate the correlation coefficient between test evaluation results and actual work performance. The parameter adjustment unit is used to dynamically adjust system parameters based on the correlation analysis results. When the correlation coefficient between a certain behavioral feature and the actual work performance is consistently lower than 0.6, the feature extraction strategy optimization process is initiated. When the weight allocation deviates significantly from the actual importance, the weight optimization process is initiated.

Citation Information

Patent Citations

  • Personalized intelligent interview system

    CN120806899A

  • Intelligent AI interview system and method based on multi-model fusion

    CN121169339A