Dynamic IRT-variance joint detection algorithm for item discrimination in human resource assessment

By using the dynamic IRT-variance joint detection algorithm, the changes in item discrimination are monitored in real time, which solves the problem of instability of assessment results caused by dynamic changes in item discrimination. This enables the intelligent self-evolution and continuous quality optimization of the human resource assessment system, and improves the validity and fairness of the assessment.

CN122367417APending Publication Date: 2026-07-10CRUITE SOFTWARE GRP CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
CRUITE SOFTWARE GRP CO LTD
Filing Date
2026-04-14
Publication Date
2026-07-10

AI Technical Summary

Technical Problem

In existing human resource assessments, the discriminatory power of items is difficult to monitor effectively during dynamic changes, leading to a decline in the overall validity and fairness of assessment results. Traditional IRT parameter estimation methods cannot reflect the dynamic changes in item performance in real time and lack automated and intelligent detection mechanisms.

Method used

The dynamic IRT-variance joint detection algorithm is adopted. By collecting the full-link behavior data of the respondents, the initial parameters of the question IRT are calculated, the evolutionary change rate is analyzed, a multi-dimensional question evaluation report is generated, a dynamic question bank quality stratification system is constructed, the stability and drift trend of question discrimination are identified, and question quality defects are analyzed and optimized, so as to realize the intelligent self-evolution and continuous optimization of the question bank.

Benefits of technology

It achieves high-precision modeling and dynamic monitoring of question discrimination, ensuring the stability and fairness of assessment results. Through automated quality stratification and optimization mechanisms, it improves the overall validity and fairness of the question bank, supports assessment designers in scientifically retaining, replacing or re-verifying questions, and ensures the dynamic balance and long-term reliability of the question bank structure.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122367417A_ABST
    Figure CN122367417A_ABST
Patent Text Reader

Abstract

This invention relates to the field of psychological assessment analysis, and more particularly to a dynamic IRT-variance joint detection algorithm for item discrimination in human resource assessment. The algorithm includes the following steps: collecting full-link behavioral data of respondents from the assessment terminal to calculate initial IRT parameters for each item; performing evolutionary analysis and time-series change rate analysis based on the initial IRT parameters to obtain the dynamic IRT parameter evolution trajectory and the dynamic drift index of discrimination; rating the applicability of items based on the dynamic IRT parameter evolution trajectory to generate a multi-dimensional item evaluation report; and automatically stratifying the quality of the assessment item bank based on the dynamic drift index of discrimination, constructing a dynamic item bank quality stratification system. This invention analyzes and comprehensively rates the discrimination of psychological assessment items, achieving a three-dimensional evaluation of item quality, optimizing the item bank configuration scheme, and ensuring a dynamic balance between validity, reliability, and fairness in the item bank.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of psychological assessment and analysis, and in particular to a dynamic IRT-variance joint detection algorithm for item discrimination in human resource assessment. Background Technology

[0002] In actual human resource assessment processes, as the size of the assessment sample expands, the assessment scenarios diversify, and the question bank is continuously updated, the discrimination of items is dynamically affected by various factors. For example, changes in the characteristics of the test-takers, increased frequency of item use, leakage of question stems, or evolution of answering strategies can all lead to discrimination drift or degradation. This dynamic change causes items to exhibit inconsistent discrimination capabilities across different time periods and different test groups, thereby reducing the overall validity and fairness of the assessment. Traditional IRT parameter estimation methods are mostly based on static models, relying solely on one-time samples to calculate item parameters, making it difficult to reflect the dynamic changes in item performance over time.

[0003] Furthermore, existing question bank quality monitoring methods mostly rely on periodic expert sampling assessments or simple statistical indicator analyses, lacking automated and intelligent dynamic detection mechanisms. These methods often only detect abnormal performance of questions at a certain moment, but cannot continuously track the evolution trend of question parameters, nor can they quantify the impact of changes in question discrimination on the overall assessment variance structure. As a result, when question discrimination degrades or drifts abnormally, the system often fails to identify it in a timely manner, easily leading to distorted assessment results, failure of group discrimination, and even affecting the fairness of job matching or decision-making.

[0004] In large-scale human resource assessment applications, question banks typically contain thousands to tens of thousands of questions, and each question generates a large amount of behavioral data across different assessment batches. How to monitor the changing trends of question discrimination in real time from this massive amount of dynamic response data, identify potential risks of question degradation, and establish an automated quality stratification and optimization mechanism has become a significant technical challenge facing the current assessment field. Summary of the Invention

[0005] To address the aforementioned technical problems, this invention proposes a dynamic IRT-variance joint detection algorithm for item discrimination in human resource assessment, thereby resolving at least one of the aforementioned technical issues.

[0006] To achieve the above objectives, this invention provides a dynamic IRT-variance joint detection algorithm for item discrimination in human resource assessment, comprising the following steps: Step S1: Collect the respondent's full-link behavior data based on the assessment terminal, and calculate the initial parameters of the question's IRT; Step S2: Based on the initial parameters of the IRT in the problem, perform evolutionary analysis and time series change rate analysis to obtain the dynamic IRT parameter evolution trajectory and the dynamic drift index of discrimination. Step S3: Based on the evolution trajectory of dynamic IRT parameters, evaluate the applicability of the questions and generate a multi-dimensional question evaluation report; Step S4: Automated quality stratification of the assessment question bank based on the dynamic drift index of discrimination, and construct a dynamic question bank quality stratification system; Step S5: Perform stability identification of question discrimination and analysis of question quality defects on the dynamic IRT parameter evolution trajectory to obtain question quality defect records; Step S6: Based on the question quality defect record and dynamic question bank quality stratification system, conduct multi-objective optimization simulation and comprehensive evaluation of the solution, and extract the optimal solution.

[0007] The beneficial effects of this invention are specifically as follows: By collecting full-link behavioral data such as answer duration, reaction time distribution, answer order, and question-answering frequency, high-precision modeling of the respondent's potential abilities and answer characteristics can be achieved. Based on behavioral data, the discrimination (parameter a), difficulty (parameter b), and guessing ability (parameter c) of the questions are calculated, ensuring that the initialization of the IRT model parameters more closely matches real-world answering behavior. Initial modeling driven by behavioral data reduces parameter bias caused by static data estimation, providing a stable starting point for subsequent dynamic evolution analysis. By calculating the rate of change of IRT parameters with assessment batches or time, the trend changes in question discrimination and difficulty can be dynamically depicted. A "discrimination dynamic drift index" is constructed to quantitatively reflect the stability and fluctuation of question discrimination ability across different groups and time periods. By detecting abnormal parameter changes in advance (such as a decrease in discrimination or a sudden change in difficulty), quality risks such as question aging, leaks, or answer deviations can be promptly identified. Based on the dynamic IRT evolution trajectory, the performance of questions across different ability levels, groups, and time dimensions is comprehensively rated. The evaluation report includes multiple indicators such as discrimination stability, difficulty rationality, and parameter drift trend, achieving a comprehensive assessment of question quality. The rating results assist assessment designers in scientifically retaining, replacing, or revising questions, thereby continuously optimizing the question bank structure. Based on discrimination drift stability, question contribution, and parameter change amplitude, the question bank is automatically divided into high-quality, monitoring, and elimination zones. This tiered system ensures a reasonable distribution of questions at different levels within the question bank, avoiding the concentrated use of highly volatile or low-discrimination questions. The system can automatically adjust the question tier level according to the latest IRT dynamic indicators, maintaining the dynamic balance of the question bank and long-term assessment reliability. Statistical fluctuation analysis of the evolution trajectory identifies questions with significantly decreased discrimination or abnormal fluctuations. Possible causes of quality defects can be analyzed from dimensions such as differences in assessment subjects, question stem expression, and answer design. A "Question Quality Defect Record" is generated, facilitating continuous improvement and question version tracking for assessment organizations, achieving closed-loop quality management. With multiple objectives such as discrimination stability, question coverage, and reasonable difficulty distribution as constraints, the system comprehensively evaluates question bank configuration schemes through simulation optimization algorithms (such as multi-objective evolutionary algorithms). By extracting the optimal scheme, the system ensures a dynamic balance between validity, reliability, and fairness in the question bank. The system can periodically run optimization simulations and automatically generate suggestions for adjusting the question bank structure, achieving intelligent self-evolution and continuous quality optimization of the human resource assessment system. Attached Figure Description

[0008] Figure 1 This is a flowchart illustrating the steps of a dynamic IRT-variance joint detection algorithm for item discrimination in human resource assessment according to the present invention. Figure 2 This is a detailed flowchart illustrating the implementation steps of step S1. Figure 3This is a flowchart illustrating the detailed implementation steps of step S2. Detailed Implementation

[0009] It should be understood that the specific embodiments described herein are for illustrative purposes only and are not intended to limit the scope of the invention.

[0010] This application provides a dynamic IRT-variance joint detection algorithm for item discrimination in human resource assessment. The execution entity of the dynamic IRT-variance joint detection algorithm for item discrimination in human resource assessment includes, but is not limited to, mechanical equipment, data processing platforms, cloud server nodes, network upload devices, etc., which can be considered as general computing nodes in this application. The data processing platform includes, but is not limited to, at least one of an audio-visual management system, an information management system, and a cloud data management system.

[0011] Please see Figures 1 to 3 This invention provides a dynamic IRT-variance joint detection algorithm for item discrimination in human resource assessment, comprising the following steps: Step S1: Collect the respondent's full-link behavior data based on the assessment terminal, and calculate the initial parameters of the question's IRT; Step S2: Based on the initial parameters of the IRT in the problem, perform evolutionary analysis and time series change rate analysis to obtain the dynamic IRT parameter evolution trajectory and the dynamic drift index of discrimination. Step S3: Based on the evolution trajectory of dynamic IRT parameters, evaluate the applicability of the questions and generate a multi-dimensional question evaluation report; Step S4: Automated quality stratification of the assessment question bank based on the dynamic drift index of discrimination, and construct a dynamic question bank quality stratification system; Step S5: Perform stability identification of question discrimination and analysis of question quality defects on the dynamic IRT parameter evolution trajectory to obtain question quality defect records; Step S6: Based on the question quality defect record and dynamic question bank quality stratification system, conduct multi-objective optimization simulation and comprehensive evaluation of the solution, and extract the optimal solution.

[0012] In the embodiments of the present invention, see Figure 1 This is a flowchart illustrating the steps of a dynamic IRT-variance joint detection algorithm for item discrimination in human resource assessment according to the present invention. In this example, the steps of the dynamic IRT-variance joint detection algorithm for item discrimination in human resource assessment include: Step S1: Collect the respondent's full-link behavior data based on the assessment terminal, and calculate the initial parameters of the question's IRT; In this embodiment, a behavior monitoring module integrated into the assessment terminal collects detailed interactive behaviors of the test taker for each question in real time, including the start and end times of the test, the number of switches between questions, the number of answer modifications, and the dwell time and trajectory distribution of the mouse cursor in the question area. Subsequently, this multimodal behavior data is integrated to construct a behavior matrix containing timestamps. The behavior data is then precisely matched and synchronized using timestamps to achieve a serialized representation of the entire behavior chain. Based on this behavior matrix, statistical methods are used to extract key features, such as average answering time, the mean and variance of question modification frequency, the distribution characteristics of pause duration, and hotspot area analysis indicators from the mouse heatmap. Finally, using IRT modeling techniques (such as one-parameter or two-parameter models) combined with the behavioral features, the initial discrimination, difficulty, and guess parameters of the questions are calculated using maximum likelihood estimation or Bayesian estimation methods. These initial parameters provide a basic reference for subsequent dynamic analysis, ensuring that the model accurately reflects the impact of the test taker's actual behavior on question performance.

[0013] Step S2: Based on the initial parameters of the IRT in the problem, perform evolutionary analysis and time series change rate analysis to obtain the dynamic IRT parameter evolution trajectory and the dynamic drift index of discrimination. In this embodiment, the discrimination parameter of the questions is divided into multiple continuous sequences along the time dimension. Time series analysis techniques, such as sliding windows or weighted moving averages, are used to calculate the rate of change of the parameters within each time period. Statistical smoothing methods are applied to reduce parameter noise and improve the accuracy of trend identification. Subsequently, time series modeling techniques, combined with dynamic Bayesian networks or state-space models, are used to capture the evolution trajectory of the parameters at different time points. Based on these trajectories, a dynamic drift index of the discrimination is calculated to quantify the stability and intensity of change of the discrimination over time. This index reflects the performance volatility of the questions in actual assessments and is an important component of the dynamic IRT model. Through this analysis, performance drift caused by changes in the assessment environment or the respondent group can be effectively identified, providing a basis for decision-making in dynamic maintenance and quality control of the question bank.

[0014] Step S3: Based on the evolution trajectory of dynamic IRT parameters, evaluate the applicability of the questions and generate a multi-dimensional question evaluation report; In this embodiment, the dynamic parameter changes of the questions are compared with established quality standards, covering multiple dimensions such as discrimination stability, difficulty suitability, and the impact of guessing behavior. Specifically, data dimensionality reduction technology is used to map dynamic parameter information to a low-dimensional space, and cluster analysis is combined to classify questions into categories, identifying groups of questions that perform well, fluctuate significantly, or pose potential risks. Next, a weighted scoring mechanism is used to integrate the various indicators to form a comprehensive applicability score. The evaluation report details the dynamic performance characteristics, risk level, and suggested improvement measures for each question, supporting question bank administrators in developing targeted maintenance strategies. The report also supports multi-dimensional visualization, such as dynamic parameter trend charts, stability radar charts, and discrimination distribution heatmaps, enhancing the intuitive understanding and analytical depth of question quality. This process achieves accurate rating based on dynamic data, improving the overall scientific management level of the question bank.

[0015] Step S4: Automated quality stratification of the assessment question bank based on the dynamic drift index of discrimination, and construct a dynamic question bank quality stratification system; In this embodiment, questions are categorized into multiple quality levels based on their dynamic drift, typically with four tiers ranging from high stability and high discrimination to low stability and low discrimination. First, quantiles are established based on the numerical range of the drift index to determine the tier boundaries. Then, a multi-factor discrimination model is formed by combining the basic IRT parameters of the questions with dynamically changing data to further refine the tier assignment. This automated quality tiering system not only enables real-time monitoring and dynamic management of questions but also supports regular updates, ensuring that the question bank reflects the latest assessment data and answering behaviors. The tiering results will be used to optimize the question bank structure, guide question revisions, and adjust assessment schemes. The establishment of this system effectively improves the efficiency of scientific question bank maintenance and promotes the long-term stability and fairness of assessment tools in practical applications.

[0016] Step S5: Perform stability identification of question discrimination and analysis of question quality defects on the dynamic IRT parameter evolution trajectory to obtain question quality defect records; This embodiment conducts an in-depth analysis of unstable questions in the dynamic IRT parameter evolution trajectory, identifying the stability of question discrimination and analyzing potential quality defects. The process includes classifying questions based on time-series entropy values ​​and variation amplitude indicators, distinguishing between stable and highly volatile questions. For volatile questions, anomaly pattern detection and fluctuation attribution analysis are further performed using respondent behavior data to uncover key factors affecting question performance, such as abnormal answering times and repeated revisions. Subsequently, content analysis, option design rationality assessment, and interface evaluation are conducted to confirm defects at the question design level. All diagnostic information is summarized to form a question quality defect record, including defect type, scope of impact, and severity. This record not only provides a basis for question revision but also supports the construction of a dynamic quality management and risk warning system for the question bank, enabling continuous tracking and closed-loop control of question quality.

[0017] Step S6: Based on the question quality defect record and dynamic question bank quality stratification system, conduct multi-objective optimization simulation and comprehensive evaluation of the solution, and extract the optimal solution.

[0018] In this embodiment, based on the established question quality defect record and dynamic quality stratification system, multi-objective optimization simulation is conducted to improve the overall performance and validity of the question bank. The optimization process employs a multi-objective optimization algorithm, comprehensively considering indicators such as improved question discrimination, balanced question bank coverage, respondent workload, and testing time, generating multiple improvement schemes. Each scheme undergoes simulation testing to evaluate its impact on the stability of dynamic question parameters and testing accuracy. The simulation results are comprehensively evaluated using a weighted scoring method to select the scheme with optimal performance and reasonable implementation cost. Finally, the extracted optimal scheme supports real-time dynamic updates and adjustments to the question bank, ensuring the adaptability and reliability of the assessment tool in complex environments. This step achieves a closed loop from data-driven diagnosis to intelligent optimization decision-making, significantly improving the dynamic monitoring and continuous improvement capabilities of question discrimination in human resource assessments.

[0019] In this embodiment, see Figure 2 The diagram below illustrates the detailed implementation steps of step S1. In this embodiment, the detailed implementation steps of step S1 include: Based on the collection of full-link behavior data of test takers by the assessment terminal, the test duration sequence, question modification frequency, pause distribution and mouse heatmap are calculated to generate a multimodal test behavior dataset; Calculate the behavior timestamps of the respondent's full-link behavior data; accurately match and label the multimodal response behavior dataset based on the behavior timestamps to generate a time-series behavior matrix; Perform abnormal response pattern recognition on the time-series behavior matrix and label guessing answers and random answer samples; After removing guessing and random answers, a standardized answer behavior matrix is ​​obtained. Ability-level clustering analysis was performed based on a standardized answer behavior matrix to extract the response distribution characteristics of different ability ranges; Based on the response distribution characteristics, the question difficulty coefficient, discrimination parameter, and guessing parameter are calculated, and the initial parameters of the question IRT are obtained by fitting.

[0020] In this embodiment, a behavior collection script is embedded in the front-end terminals (including PC and mobile terminals) of the assessment system. An event listening mechanism records the user's operational trajectory data, such as clicks, pauses, mouse movements, keyboard input, and submission times. The behavior sampling frequency is set to 20Hz to ensure that the temporal resolution can capture subtle behavioral changes. Simultaneously, the response duration sequence (unit: milliseconds), question modification frequency, pause distribution characteristics, and mouse trajectory coordinates are recorded. The backend server receives the event stream data uploaded by the terminal in real time and performs data cleaning through a unified log parsing module to remove non-response behavior noise such as network jitter and browser focus issues. Subsequently, a spatial aggregation algorithm (such as Kernel DensityEstimation) is used to generate mouse heatmap features to describe the spatial distribution characteristics of response attention. The resulting multimodal behavior dataset contains approximately 200 dimensions of features, including time dimension (duration sequence), spatial dimension (heatmap features), and frequency dimension (modification rate, click rate), for subsequent time series analysis and IRT parameter estimation. Each response event (such as click, modification, and submission) is assigned a timestamp accurate to the millisecond level. To ensure clock consistency, the synchronization error between the terminal's local time and the server's NTP time is controlled within ±5ms. Then, based on timestamps, the data from each modality (duration sequence, mouse trajectory, pause distribution, modification frequency) are reassembled in chronological order. To address the issue of inconsistent sampling frequencies across different modalities, Dynamic Time Warping (DTW) is used to achieve time alignment, ensuring that behavioral events at the same moment maintain consistency in their time indices. Subsequently, the data is labeled according to question number, answering stage (browsing, answering, submitting), and behavior type, giving the time series data a clear semantic structure. The final generated time-series behavior matrix has the respondent as the row and the time slice as the column, with matrix elements representing the behavior intensity value at that moment (such as mouse movement speed, number of clicks, pause length, etc.).

[0021] The obtained temporal behavior matrix was subjected to feature compression, and 10 principal components were extracted using Principal Component Analysis (PCA) to reduce dimensionality and retain more than 90% of the information variance. Then, the Isolation Forest algorithm was used for anomaly detection. By randomly partitioning the feature space, the isolation depth of each sample was calculated to determine its degree of anomalousness. In the experiment, the isolation factor threshold was set to 0.35; samples exceeding this threshold were considered outliers. To further distinguish between speculative answers and systematic errors, density clustering (DBSCAN) was used for outlier clustering analysis. If a cluster exhibited typical characteristics such as "extremely short answer time," "no pauses," and "mouse trajectory concentrated in the submit button area," it was marked as a speculative answer sample. Answer records marked as anomalous were filtered; if more than 30% of a subject's answers were marked as anomalous, that subject was removed entirely. Finally, the retained samples underwent feature standardization. Time-related features (such as answer duration and pause intervals) were standardized using Z-scores to make the duration features of different questions comparable. Frequency-related features (such as click rate and modification frequency) were normalized using Min-Max, mapping the values ​​to the [0,1] interval to eliminate differences in behavioral magnitude. After standardizing the matrix, missing values ​​were imputed using a time series interpolation algorithm (Cubic Spline Interpolation) to ensure temporal continuity. The mean squared error of the matrix decreased by approximately 27% after processing, indicating an effective reduction in noise. The final standardized answer behavior matrix showed good performance in terms of statistical stability and behavioral consistency.

[0022] Five core behavioral indicators were selected from the standardized matrix: average response time, pause frequency, modification rate, click rate, and behavioral rhythm entropy. K-means++ clustering was used for ability stratification analysis, with K set to 5 to represent five ability levels: low, low-medium, medium, high-medium, and high. Euclidean distance was used to calculate the similarity between samples, and the cluster centers were iteratively updated until the within-cluster variance no longer decreased significantly (the rate of change was less than 0.01). After clustering, statistical analysis of the average behavioral characteristics of each group revealed that the high-ability group exhibited "shorter response time, higher modification rate, and lower pause variance," while the low-ability group exhibited "longer response time, lower modification rate, and higher pause rate." Hierarchical clustering was further used to verify the stability of the stratification, with the within-group behavioral variance all less than 0.2. This ability-level clustering not only provides a reference for individual differences in IRT model parameter estimation but also reveals the distribution patterns of behavioral characteristics across different ability ranges.

[0023] The initial parameters of the Item Response Theory (IRT) are calculated and fitted using a three-parameter logical model (3PL model). The formula is as follows: ,in, For the difficulty level of the question, For discrimination parameters and To guess the parameters, θ represents the assessor's ability value. The parameters are solved using maximum likelihood estimation (MLE) combined with expectation-maximization (EM). In the experiment, the ability level center value is used as the input for the participants' ability level. The accuracy distribution of each question across different ability levels is fitted to estimate the initial parameter values. The initial parameter setting rules are: average answering time is positively correlated with difficulty, modification rate is positively correlated with discrimination, and rapid submission rate is positively correlated with guessed parameters. After approximately 25 iterations, when the parameter convergence rate is less than 10... -4 The calculation stops when the time is right. The range of the obtained problem parameters is: ∈[−2.1,1.8], Mean 1.34, The mean is 0.18.

[0024] In this embodiment, see Figure 3 The diagram below illustrates the detailed implementation steps of step S2. In this embodiment, the detailed implementation steps of step S2 include: Based on the initial parameters of the problem's IRT, joint vectorization modeling is performed to construct a high-dimensional parameter representation space; Markov chain Monte Carlo sampling is performed on the high-dimensional parameter representation space to generate posterior distribution estimation coefficients of the parameters. Calculate the confidence interval and peak probability density of the discrimination parameter based on the posterior distribution estimation coefficients of the parameters; Based on the confidence interval and the peak probability density, a reliability assessment of parameter estimation is performed to obtain an estimated reliability value; Based on the estimated reliable values, the initial parameters of the problem's IRT are corrected in real time and the time series change rate is analyzed to obtain the dynamic drift index of discrimination.

[0025] In this embodiment, after obtaining the initial IRT parameters (i.e., the discrimination 'a', difficulty 'b', and guessing parameter 'c' in the three-parameter model) for each assessment question, these parameters need to be further vectorized to construct a high-dimensional parameter representation space for subsequent joint modeling and distribution estimation. The core objective of vectorized modeling is to represent the IRT parameters of each question as a fixed-dimensional vector, for example, [a_i, b_i, c_i] ∈ R 3 Simultaneously, it integrates behavioral performance characteristics of questions at different ability levels (such as standardized answering time μ_i, answering variance σ²_i, behavioral focus β_i, etc.), further expanding it into a multi-dimensional vector [a_i, b_i, c_i, μ_i, σ²_i, β_i] ∈ R 6In the experiment, initial parameters for the IRT (Information and Representation of the Response) of 200 questions were selected. Combining the mean and variance information of user answering behaviors at each ability level, each question was mapped to a 6-dimensional vector, forming a parameter representation matrix of dimension 200 × 6. This high-dimensional space reflects the intrinsic relationship between IRT parameters and behavioral data, providing a structured data foundation for subsequent Bayesian modeling and sampling inference. Furthermore, to improve the expressive efficiency of the space, the original vectors were Z-score standardized to eliminate dimensional differences between different features. To improve the robustness and reliability of parameter estimation, Bayesian posterior distribution modeling is required for each dimension of the aforementioned high-dimensional parameter space. Markov Chain Monte Carlo (MCMC) methods, particularly the No-U-Turn Sampler (NUTS) algorithm, are used to jointly sample the discrimination, difficulty, and guess parameters of each question to generate posterior distribution estimates. The sampling model was constructed based on Bayes' theorem, assuming a normal prior distribution (e.g., a ~ N(1, 0.5), b ~ N(0, 1), c ~ Beta(2,5)). Observations were derived from standardized behavioral data and IRT initial parameters. The experiment consisted of 5000 sampling steps, with the first 1000 steps being the burn-in period. Four sampling chains were used, and convergence tests (e.g., the Gelman-Rubin statistic) were performed on each chain. <1.1 as the standard). Through MCMC sampling, the true posterior probability distribution of the discrimination parameter under behavioral data constraints can be obtained, no longer relying solely on the single-point value of maximum likelihood estimation. The sampling results show that the α parameter of some items exhibits an asymmetric distribution (significant skewness), indicating that behavioral differences have a substantial impact on parameter stability, providing a probabilistic basis for subsequent dynamic calibration.

[0026] By statistically analyzing the posterior distribution obtained from MCMC sampling, the confidence interval and peak probability density (i.e., maximum a posteriori estimate MAP) of the discrimination parameter (a) for each item can be calculated. Specifically, all sample points for the discrimination dimension are extracted from the sampling results, and the probability distribution curve is calculated using kernel density estimation (KDE). Based on this, a 95% confidence interval is determined (usually the 2.5% and 97.5% quantiles). Simultaneously, the peak position of this density curve is identified as the MAP value of the parameter. Taking a certain item as an example, its MCMC sample for parameter a exhibits a right-skewed distribution, with a 95% confidence interval of [0.84, 1.36] and a peak value at 1.12. This result is slightly lower than the initial maximum likelihood estimate of 1.25, indicating that the item discrimination is somewhat conservative from a posterior perspective considering behavioral heterogeneity. In the experiment, confidence intervals and maximum margins (MAPs) were calculated for the discrimination parameters of all 200 questions. Approximately 17% of the questions exhibited abnormally wide confidence intervals (length exceeding 0.7), suggesting potential user behavior noise or ambiguity in ability differentiation. The size and symmetry of the confidence intervals will be crucial for subsequent parameter stability and estimation reliability assessment. The reliability of parameter estimation directly affects the explanatory and predictive power of the IRT model. Therefore, a reliability assessment mechanism needs to be constructed based on the posterior distribution of the discrimination parameters for each question, comprehensively considering the confidence interval width, MAP, and the degree of deviation from the prior estimate. The Estimation Reliability Index (ERI) is defined as follows: ERI = 1 - (confidence interval length / prior standard deviation) × deviation factor, where the deviation factor reflects the difference between the MAP and the initial estimate. In the experimental setup, the prior standard deviation was set to 0.5. If the confidence interval for parameter a of a question is 0.5, and the MAP differs significantly from the initial value (>0.2), its ERI value decreases significantly. Typically, ERI ∈ [0,1], with values ​​closer to 1 indicating higher reliability. In this experiment, approximately 68% of the questions had an ERI value greater than 0.75, falling into the "highly reliable" range; 21% of the questions were between 0.5 and 0.75, classified as "moderately reliable"; and the remainder were marked as "low reliable," with recommendations for review or adjustment. This evaluation system provides a systematic standard for dynamic IRT parameter tuning, avoiding the misleading impact of parameter fluctuations caused by extreme behavioral data on the evaluation results. The ERI-based visualization heatmap also reveals that the questions most significantly affected by answer behavior in discrimination estimation are concentrated in the medium-to-high difficulty range.

[0027] To achieve the goal of making the item discrimination parameter "dynamically adjustable" during actual testing, it is necessary to combine the aforementioned ERI value, posterior distribution trend, and real-time testing data to adjust the initial IRT parameter in real time, resulting in the Discrimination Drift Index (DDI). This index measures the changing trend of a item's discrimination parameter at different testing time stages and is defined as: DDI = |a_t - a0| / a0, where a_t is the current adjusted value and a0 is the initial estimate. Through sliding window analysis (each window consisting of 30 participants), the adjustment magnitude of the discrimination parameter's feedback in terms of answering behavior and accuracy among different user groups is statistically analyzed. For example, if the discrimination of an item is stable among the first 300 test takers (a≈1.2) but significantly decreases among the last 300 (a≈0.85), then DDI ≈ 0.29, indicating a significant drift phenomenon. Experimental data shows that approximately 12% of the 200 questions had a Discrimination Index (DDI) > 0.25, mostly concentrated in knowledge-based or logic-based questions. This suggests that changes in user strategy or fatigue during answering may lead to a decline in the question's discriminatory power over time. By introducing DDI, an IRT assessment system with "adaptive adjustment capabilities" can be constructed, enabling real-time monitoring and dynamic optimization, significantly improving the validity and stability of the assessment tool in high-frequency application scenarios (such as recruitment written tests and large-scale assessments).

[0028] In this embodiment, the specific steps for performing real-time correction and time-series change rate analysis on the initial parameters of the IRT based on the estimated reliable value to obtain the dynamic drift index of discrimination are as follows: The initial parameters of the problem's IRT are corrected in real time based on the estimated reliable values ​​to obtain the corrected initial parameters of the IRT. Convergence diagnosis is performed on the modified IRT initial parameters to obtain convergence diagnosis results, which include the Gehrman-Rubin statistic and the effective sample size. The convergence diagnosis results are adaptively adjusted by adjusting the sampling step size to generate a dynamic IRT parameter evolution trajectory. The dynamic IRT parameter evolution trajectory is analyzed by time-series change rate analysis to obtain the dynamic drift index of discrimination.

[0029] In this embodiment, after obtaining the Estimation Reliability Index (ERI) for each question, the initial IRT parameters of the questions need to be corrected in real time to form a "corrected IRT initial parameter set". The core of the correction process is to balance the original IRT estimate with the more confident parameter estimates in the posterior distribution through weight fusion. For example, if the initial discrimination estimate of a question is a0 = 1.3, and the posterior maximum probability estimate (MAP) is a m=1.0, ERI=0.9, then the weighted average correction formula is used: a* = (1-ERI) × a0+ ERI ×a m The final corrected value was a*=1.03. This fusion strategy effectively mitigates extreme biases in parameter estimation while maintaining the model's responsiveness to the real behavioral distribution. In the experiment, a lower limit for correction was set at ERI≥0.6; parameters below this value maintained their original estimates but were marked as "requires review." Of the 200 assessment questions processed in this phase, approximately 76% of the discrimination parameters underwent slight to moderate corrections (correction magnitude between 5% and 15%), demonstrating the dynamic driving force of behavioral feedback on IRT modeling. This correction process provides a more stable and reliable parameter foundation for subsequent sampling iterations and trajectory analysis. The corrected IRT parameters do not necessarily indicate complete convergence; therefore, further statistical convergence diagnostics are needed to determine whether posterior sampling has reached a stable state. Two core indicators were primarily used: the Gelman-Rubin statistic. ) and Effective Sample Size (ESS). The metric measures the variance consistency among different MCMC chains, ideally close to 1.0 (typically acceptable range is [1.00, 1.10]). ESS reflects the number of non-autocorrelated samples in the parameter sample; a higher ESS indicates more representative sampling results. In the experiment, four MCMC chains were set for each question, with 4000 steps sampled per chain, the first 1000 steps being the burn-in period. Convergence diagnosis showed that most questions achieved a higher discriminant parameter after correction. The values ​​are concentrated in the range of [1.01, 1.07], and the ESS is above 800, indicating that the sampling results have strong statistical stability. However, about 8% of the items have discriminant parameters. Values ​​exceeding 1.12 or ESS below 400 suggest potential multimodality or insufficient sampling in the parameter distribution. These issues are marked as "convergence risk," and it is recommended to extend the sampling length or reset the prior distribution. Additionally, plotting... The ESS distribution heatmap helps to quickly identify the differences in parameter stability distribution across the entire question bank, assisting in question quality control and assessment strategy adjustment.

[0030] An adaptive sampling mechanism based on convergence diagnosis results is introduced to dynamically adjust the step size and chain number during the MCMC sampling process, thereby generating high temporal resolution parameter evolution trajectories. The adaptive strategy is mainly based on... With ESS two indicators: ① If If the value is >1.10 or ESS < 500, the sampling steps will be automatically extended (from the original 4000 to 6000 or 8000 steps); ② If If the value is <1.02 and ESS>1000, sampling is terminated early to save computational resources; ③ For questions with slow parameter convergence, local step-weighted sampling is introduced (e.g., increasing the density of sampling points in the region of discriminative variation). The sampling trajectory records the value of the discriminative parameter a at regular intervals (e.g., every 200 steps), ultimately forming a time series a(t). In the experiment, sampling trajectories were generated for 200 questions, with an average of 15-25 effective parameter points per question, forming a dynamic IRT evolution trajectory matrix [question number × sampling round × parameter dimension]. This trajectory is not only used to analyze parameter stability but also reflects the dynamic adaptability of the question to changes in behavioral feedback. For example, if the value of a question is stable at 1.2 in the initial stage and then gradually decreases to 0.95, it indicates that its discriminative power is drifting significantly with changes in user behavior. The evolution trajectory data generated in this stage will be used as input for the analysis of dynamic drift indicators. After obtaining the time-series trajectory of the parameter evolution during the sampling process, it is necessary to further analyze the trend, volatility, and rate of change in these trajectories to construct a Discrimination Drift Index (DDI) to assess the stability and risk of the IRT parameters in the problem. The specific calculation process includes: ① Performing time-series differencing on the a(t) sequence to calculate the rate of change Δa(t) = a(t) - a(t-1) between adjacent sampling points; ② Smoothing Δa(t) using an exponentially weighted moving average (EWMA) to eliminate high-frequency noise; ③ Combining the direction of change (unidirectional upward / downward) and the amplitude of volatility (maximum Δa value), constructing the DDI index with the formula: DDI = max(|Δa|) × drift direction coefficient (+1 or -1). The experimental thresholds are set as follows: DDI > 0.15 indicates "high drift risk," between 0.05 and 0.15 indicates "moderate drift," and DDI < 0.05 is considered parameter stability. In this experiment, approximately 14% of the questions had a Discrimination Index (DDI) exceeding 0.15, with the majority concentrated in question types characterized by high behavioral complexity and cognitive load (such as multiple logical reasoning questions). These types of questions are susceptible to changes in user strategies or test-taking fatigue, leading to a decrease in the accuracy of IRT modeling. By identifying dynamic drift indicators, real-time risk warnings, parameter re-evaluations, and iterative optimization of assessment content in the IRT model can be further achieved, ensuring the stability, fairness, and adaptability of human resource assessment in large-scale application scenarios.

[0031] In this embodiment, step S3 specifically involves the following steps: Based on the response distribution characteristics and the evolution trajectory of dynamic IRT parameters, the maximum capability is estimated to generate the subject's capability estimate. Hierarchical variance decomposition was performed on the estimated ability values ​​of the test subjects to generate the total variance, between-group variance and within-group variance components. The hierarchical variance contribution of different questions is calculated based on the overall variance, between-group variance, and within-group variance components, generating the variance contribution of each question. Based on the variance contribution, the applicability of the questions is rated, and a multi-dimensional question evaluation report is generated.

[0032] In this embodiment, based on the dynamic evolution trajectory of IRT parameters (i.e., the dynamic changes of discrimination a, difficulty b, and guessing parameter c for each question over time), and combined with the test takers' answer performance and behavioral distribution characteristics (such as answer duration, modification frequency, and focus), the maximum likelihood estimation (MLE) method is used to dynamically estimate the test takers' potential ability value (θ). Since the IRT model assumes that the probability of answering a question correctly depends on the functional relationship between the test taker's ability and the question parameters, the estimation of θ needs to dynamically adapt to the current IRT parameter state of the question. In this experiment, a multi-round dynamic MLE iterative mechanism is used: the initial value of θ for each test taker is set to 0. Combining their answer performance for each question with the IRT parameters of that question at the current time point, the likelihood function is calculated, and then θ is iteratively updated using the Newton-Raphson method until convergence. Furthermore, a behavioral feature weighting factor is added (e.g., a deviation of the answer time from the mean ±2σ will reduce the weight) to enhance the model's robustness to atypical behaviors. Ability estimation was performed on 1000 participants. The obtained θ values ​​showed an approximately normal distribution, with a mean close to 0 and a standard deviation of approximately 0.91, validating the normality assumption of the estimation model and providing an input basis for variance decomposition. After obtaining the estimated ability values ​​(θ) of the participants, it is necessary to further analyze their distribution characteristics in the group structure and identify the hierarchical sources of ability differences. A hierarchical linear model (HLM) was used for variance decomposition, splitting the total ability variance (Var_total) into between-group variance (Var_between) and within-group variance (Var_within). In the experiment, the 1000 participants were divided into 10 groups of 100 people each according to their job type (e.g., management, technical, operations, etc.). A two-level model was constructed: the first level was the individual level (participant's θ value), and the second level was the job group level. The model is in the form θ_ij = γ_00 + u_0j + ε_ij, where γ_00 is the population mean, u_0j is the between-group bias, and ε_ij is the within-group residual. The variance components were obtained through maximum likelihood estimation. The results show that the population variance Var_total = 0.827, with the between-group variance Var_between = 0.312 (37.7%) and the within-group variance Var_within = 0.515 (62.3%). This indicates that although individual differences are significant, the job group structure still has a significant explanatory power for ability differences. This decomposition provides a quantitative basis for further exploring the contribution of the problem to the variance structure at different levels.

[0033] To assess the differentiating and explanatory role of each question in the overall competency structure, its contribution to different variance components within the competency hierarchy needs to be calculated. An item-by-item regression analysis method is used to establish an explanatory model for competency estimation for each question, and its explanatory proportion for between-group and within-group variance is calculated. Specifically, the following model is constructed: θ_ij = β_0 + β_1X_qij + e_ij, where X_qij represents the score (correct / incorrect) of the q-th question. Analysis of variance (ANOVA) is used to calculate the explanatory power (R² value) of this question for the total variance, between-group variance, and within-group variance. To enhance stability, behavioral characteristic variables (such as answering time and behavioral focus) are also introduced as covariates. Experimental results show that the variance contribution of different questions varies significantly: for example, question 12 explains 18.2% of the between-group variance, much higher than the average of all questions (9.5%); while question 45 has a strong explanatory power for the within-group variance (R²=21.6%), but is almost ineffective for the between-group variance. Finally, a ternary index set [G_total, G_between, G_within] is generated for each question, representing its contribution ratio to the overall, between-group, and within-group variances, respectively. This index system provides quantitative support for question selection, question bank optimization, and job matching. After obtaining the variance contribution of each question, a multidimensional applicability evaluation model is constructed by combining its IRT parameter stability (such as the discrimination index D value and DDI drift index), behavioral performance characteristics (such as answer behavior distribution, correction rate, and focus heatmap), and question completion rate (answer rate and effective response rate). The model employs a multi-factor weighted scoring mechanism, forming a comprehensive applicability score S_q = w1×G_total + w2×Stability + w3×BehaviorClarity + w4×CompletionRate, where the weights can be flexibly set according to actual assessment needs. In the experiment, the default weights were: variance contribution 40%, IRT stability 30%, behavioral clarity 20%, and completion rate 10%. Based on this, 200 questions were rated for applicability and divided into three levels: A (high applicability), B (medium applicability), and C (low applicability). Category A questions, comprising 78 items, possess characteristics such as high discrimination, low dynamic drift, clear behavioral features, and high completion rate. Category B questions exhibit boundary fluctuations in certain dimensions and are recommended for limited use in assessments. Category C questions mostly suffer from high behavioral abnormality rates, unstable discrimination, and weak variance explanatory power, and are recommended for elimination or reconstruction. The final multi-dimensional question evaluation report can be output as a structured data table, supporting question bank version iteration, job assessment optimization, and personalized assessment path configuration.

[0034] In this embodiment, the specific steps for generating a multi-dimensional question evaluation report based on the variance contribution rate are as follows: Based on the variance contribution, quality heterogeneity is identified to obtain high-discrimination questions and low-discrimination questions; The synergistic effect analysis of high-discrimination and low-discrimination questions was conducted to obtain the functional complementarity and redundancy characteristics between questions, and a question synergistic effect network was constructed. Calculate the discrimination of high-discrimination questions and low-discrimination questions for different ability levels, and generate local discrimination curves; Extreme point detection and inflection point identification are performed based on local discrimination curves, and the effective scope of discrimination is marked. Question applicability rating is performed based on the effective scope of discrimination and the question synergy network, generating a multi-dimensional question evaluation report.

[0035] In this embodiment, after calculating the variance contribution of each question (including the contribution ratio to the overall variance, between-group variance, and within-group variance), the quality heterogeneity of the questions can be further identified, i.e., the assessment efficacy level of the questions can be distinguished. The identification process uses a combination of multi-threshold clustering and distribution analysis. A two-dimensional index space is constructed based on the discrimination parameter (a value) and the overall variance contribution (G_total). K-means clustering is performed on the 200 questions, initially dividing them into three categories (high discrimination, high variance contribution; medium-level questions; low discrimination, low variance contribution). Next, threshold conditions are set. For example, questions with discrimination a>1.4 and G_total>15% are classified as "high discrimination questions," while questions with a<0.8 and G_total<8% are classified as "low discrimination questions." In this experiment, a total of 62 high discrimination questions, 49 low discrimination questions, and the rest are in the intermediate level. The identification results show that high-quality questions are widely distributed across core competency dimensions (such as logical reasoning and data analysis), while low-discrimination questions are concentrated in memory-based and factual questions. This identification not only provides a basis for question classification but also lays the foundation for subsequent functional complementarity analysis and synergy modeling. Based on the responses of all test takers, a correlation coefficient matrix between any two questions is calculated (e.g., using point-binary correlation or Yule's Q value), and a question similarity matrix is ​​constructed by combining the discrimination level of each question with behavioral dimension features (such as average response time and behavioral focus). Subsequently, a synergy index S_ij = Corr(i,j) × (1 - |a_i - a_j|) is defined. This index comprehensively considers the correlation and discrimination differences between questions; a higher S_ij value indicates a strong complementary or synergistic relationship between questions. A question synergy network graph is constructed using this index, where nodes represent questions and edges represent synergy strength. Community detection (Louvain algorithm) is performed on the network structure to identify question functional clusters. Experimental results show that high-discrimination questions exhibit a clear modular structure and a capability-focusing effect; while some low-discrimination questions show a weak synergistic relationship with high-discrimination questions, indicating a certain cognitive preparation function rather than complete redundancy. However, 12 low-discrimination questions were identified as highly redundant nodes, and their removal or merging is recommended. This collaborative network provides a scientific basis for optimizing the question bank's structure and enhancing its diversity.

[0036] Based on the estimated ability values ​​of the test subjects (θ ranging from [-3, +3]), the ability space was divided into 7 equidistant intervals (e.g., [-3, -2], [-2, -1]…[2, 3]), and the rate of change of accuracy for each question within each interval was statistically analyzed. A local logistic regression model was used to fit each ability interval, and the slope was extracted as the local discrimination value for that interval. Finally, a "local discrimination curve" for each question was generated, i.e., a function graph of a(θ). Experimental data showed that typical high-discrimination questions exhibited a significant upward slope in the medium ability interval (e.g., θ∈[-1, 1]), while tending to flatten out in the extremely high or extremely low ability segments (e.g., θ>2 or θ<-2); while some low-discrimination questions unexpectedly showed a sudden increase in a specific ability segment (e.g., the a value of question 87 reached 1.2 in the interval θ∈[1.5, 2.5]), suggesting that they may be suitable for refined ability boundary identification. This curve provides a core basis for modeling the scope of question ability and improves the micro-resolution of the IRT model. The method employs curve differentiation and rate of change analysis: ① Calculate the first derivative of the a(θ) curve to identify the locations of maximum and minimum slopes (maximum discriminative effect points); ② Calculate the second derivative to identify inflection point locations and determine the turning points of discriminative increase / decrease; ③ Define the interval between the maximum point and adjacent inflection points as the effective discriminative range of the question. Taking question 25 as an example, its a(θ) reaches a maximum value of 1.52 at θ=0.3, and the inflection points appear between θ=-0.6 and θ=1.8, so its effective range is [-0.6, 1.8]. Using this method, analysis of all high-discrimination and low-discrimination questions revealed that approximately 82% of the effective range of high-discrimination questions is concentrated in the interval θ∈[-1.5, 1.5], while the effective range of low-discrimination questions is often narrow and biased towards extreme values. The identified "range boundary" information can be used to accurately construct ability-adaptive question sets, enabling dynamic question delivery at different assessment stages (adaptive testing), enhancing the accuracy of ability coverage and the rationality of user experience in the assessment process.

[0037] Combining the effective scope of discrimination obtained in the first two steps with the information on the synergistic network structure of the questions, a multi-dimensional applicability rating system for the questions is constructed. It is mainly evaluated from three dimensions: ① Coverage Efficiency: assessing the length of the question's effective scope and whether it covers the core competency range; ② Synergy Strength: measuring the centrality and complementarity strength of the question in the question synergistic network; ③ Functional Stability: combining its DDI dynamic drift index and the degree of local discrimination fluctuation. Based on these three dimensions, a comprehensive score S_q = w1×Coverage + w2×Synergy + w3×Stability (default weights: 0.4, 0.35, 0.25) is calculated, forming the final applicability rating labels: A (preferred question), B (functional question), C (marginal question), and D (redundant or eliminated suggestion). In this experiment, the final percentage of A-class questions was approximately 31.5%, B-class 47.8%, C-class 15.2%, and D-class 5.5%. The multi-dimensional question assessment report output includes question number, discrimination scope, co-cluster affiliation, comprehensive rating score, and applicability suggestions. It supports chart display and structured output, and can be used for automatic question bank optimization, job-customized assessment package construction, and dynamic assessment task scheduling systems.

[0038] In this embodiment, step S4 includes the following steps: Construct an IRT path discrimination evaluation function based on the discrimination dynamic drift index; Construct a variance path discrimination evaluation function based on a multidimensional question evaluation report; The IRT path discrimination evaluation function is used to perform weighted fusion based on the variance path discrimination evaluation function to obtain the fusion evaluation result; The fusion evaluation results are standardized and mapped to obtain the standardized discrimination score. Based on standardized discrimination scores, question thresholds are defined, and a four-level discrimination judgment standard is constructed. Based on a four-level discrimination criterion, an automated quality stratification system for the assessment question bank is constructed, which is then used to build a dynamic question bank quality stratification system.

[0039] In this embodiment, a quantifiable and comparable IRT path discrimination evaluation function is constructed using the previously obtained "Discrimination Dynamic Drift Index (DDI)" to reflect the parameter stability and dynamic validity of items during the ability assessment process. The evaluation function is constructed based on the evolution trajectory of the item's IRT parameter 'a' (discrimination). By tracking its time-series changes, the function comprehensively quantifies the parameter fluctuation amplitude, drift trend, and frequency of change exhibited by the items during the assessment process. The core formula is: ; Here, Var(a_q(t)) represents the variance of the discrimination time series of item q, Drift(q) is the DDI index value (i.e., the maximum rate of change), and Stability_penalty(q) is the penalty factor for items that perform poorly in the model convergence diagnosis. In the experiment, α1=0.4, α2=0.5, and α3=0.1 to enhance the sensitivity to real dynamic drift behavior. The results show that the a-value of some logic-related items increased significantly in the early sampling stage, but tended to stabilize in the later stage, showing good dynamic adaptability; while some cognitive fatigue-related items drifted sharply throughout the sampling process, and the F_irt score was significantly low. This function provides an important parameter dimension for subsequent multidimensional fusion evaluation, emphasizing the core position of the "behavior-parameter" coupling effect in quality assessment.

[0040] Based on the key dimensions extracted from the previously constructed multidimensional question evaluation report (variance contribution, local discrimination curve, question scope, etc.), a "variance path discrimination evaluation function" is constructed to reflect the degree of path contribution of a question in ability stratification and assessment interpretability. The design philosophy of the evaluation function is to measure whether a question possesses stable discrimination efficacy across multiple ability levels, whether it holds a core position in the collaborative network, and whether it has a clear ability adaptation boundary. The evaluation function is defined as follows: Wherein, G_total(q) represents the contribution of an item to the overall ability variance (variance path factor), LDC_range(q) is the length of the local discrimination domain (i.e., the length of the interval where a(θ)>1 in the local curve), and Synergy_centrality(q) represents the centrality score of an item in the synergistic network. Experimental parameters were set to β1=0.5, β2=0.3, and β3=0.2, emphasizing the explanatory power of items on the ability structure as the core. Empirical analysis shows that items with high F_var scores effectively differentiate across multiple ability segments and often form stable synergistic relationships with multiple key items, exhibiting characteristics of high-quality path items. This function focuses on the "ability-explanation" dimension, providing a structural measurement basis for subsequent integrated scoring.

[0041] After constructing the two path discrimination evaluation functions, a weighted fusion is needed to obtain a unified evaluation result for each question. The fusion strategy follows the principle of "dynamic performance + static structure," proportionally integrating the IRT path (dynamic behavioral dimension) and the variance path (ability structure dimension) to form a total evaluation function: The fusion weights γ1 and γ2 are set according to the application scenario. In this human resource assessment, γ1=0.6 and γ2=0.4 because the stability of discrimination has a significant impact on the long-term usability of the questions. After fusion calculation of 200 questions in the experiment, the F_fused distribution is skewed, with a mean of approximately 0.63 (after standardization). Although some questions have high F_var, their F_fused is suppressed due to low F_irt, demonstrating the dominance of dynamic parameter stability in quality judgment. In addition, a confidence interval calibration mechanism is introduced to exclude abnormal questions that deviate from the 95% confidence band in the fusion results (a total of 5 questions were excluded). The final output F_fused score serves as the basis for subsequent standardized scores and stratification, and is the core mediator for automated evaluation of the question bank quality.

[0042] To ensure the comparability of F-fused scores across different topics and assessment items, they need to be standardized and mapped to generate a Standardized Discrimination Score (SDS). This step employs a combination of Z-score standardization and nonlinear mapping.

[0043] Z-standardization: Normalizing the F_fused values ​​of all items to a mean of 0 and a standard deviation of 1. Z_q = (F_fused(q) - μ_f) / σ_f; Mapping function: Considering the readability and discriminative power of the score, a logical mapping function is used to map Z_q to the interval [0,100]: SDS_q = 100 / (1 + e^(-Z_q)); This mapping method ensures the interpretability and tail sensitivity of the scores, resulting in scores for extremely high-quality questions (Z_q>2) approaching 95 or higher, while scores for low-quality questions (Z_q<-2) tend to be between 0 and 10. Empirical data shows that after standardization, the mean SDS is 54.7 and the standard deviation is approximately 18.9, satisfying a relatively ideal normal approximation distribution. SDS becomes a direct indicator for subsequent question level classification and dynamic stratification.

[0044] After obtaining the Standardized Discrimination Score (SDS), to achieve systematic management and structural optimization of question quality, it is necessary to establish clear grading standards and a hierarchical system. This step adopts a four-level judgment strategy to classify questions into: Category A (Core Questions): SDS ≥ 80 Category B (Stable Questions): 60 ≤ SDS < 80 Category C (Marginal Questions): 40 ≤ SDS < 60 Category D (Low-quality questions): SDS < 40 The segmentation strategy combines standard normal distribution theory with assessment needs and experience to ensure that core questions have high stability, high variance explanatory power, and behavioral stability. By combining level labels with the original question bank structure, a "dynamic question bank quality stratification system" can be constructed, supporting operations such as automatic question group construction, weak question removal, and key question marking under different assessment objectives.

[0045] The system automatically performs the following functions based on the question level: In high-quality job matching assessments, priority is given to selecting Category A and Category B questions; Enhanced monitoring and data collection of behavioral trajectories are implemented for Category C questions to aid in evaluation; Category D questions are periodically pushed into the revision, elimination, or behavior remodeling process.

[0046] In this embodiment, the specific steps of step S5 are as follows: The temporal entropy value of the problem is obtained by calculating the evolution trajectory of the dynamic IRT parameters; Based on the temporal entropy value of the question, a stable and fluctuating question is identified by discrimination. Perform volatility attribution analysis and abnormal pattern diagnosis on volatility issues, and extract volatility attribution response sample characteristics; Based on the characteristics of the fluctuating attribution response samples, the question quality defects are analyzed to obtain question quality defect records.

[0047] In this embodiment, the changes in the discrimination parameter of each question at different sample rounds or time points are recorded as a set of time series data. This series is then normalized and divided into several state intervals, transforming the numerical sequence into a state sequence that can express typical behavioral patterns such as "rising," "falling," "fluctuating," or "stable." Based on this, the distribution characteristics of the state sequence are quantified using the information entropy calculation method, thereby extracting the temporal entropy value of each question. This entropy value measures the randomness and uncertainty of parameter changes; a higher value indicates more significant fluctuations in the question's discrimination during the evaluation process, while a lower value indicates stable and predictable discrimination performance. This method allows for the identification of which questions maintain stable performance in actual evaluation and which questions have potential fluctuation risks from the perspective of dynamic trajectory, providing basic data support for subsequent question stability identification and anomaly attribution. The entropy values ​​of all questions are normalized and sorted to clarify the boundaries and structure of the overall entropy distribution. Then, reasonable stratification rules are set to classify the questions into three categories: "stable questions," "medium-fluctuation questions," and "high-fluctuation questions." Stable questions are characterized by consistently low time-series entropy values ​​and small variations in the discrimination parameter, demonstrating good modeling consistency. Highly volatile questions, on the other hand, often exhibit significant parameter fluctuations or frequent changes, potentially influenced by various external factors. This identification result can serve as a basis for quality stratification in question bank management and also provides directional reference for question risk assessment and behavior tracing. In particular, questions marked as highly volatile require further quality diagnostic processes to prevent misjudgments or ability deviations due to parameter instability, thereby ensuring the overall credibility of the assessment system.

[0048] From the answer data of questions marked as fluctuating, the sample group most affected by the parameter changes was selected. These samples are usually concentrated in the assessment batches or stages where the question parameters fluctuate significantly. Subsequently, multi-dimensional behavioral features of the answering process of these samples were extracted, including answering time, page dwell time, mouse operation patterns, screen switching behavior, and answering interruptions. By clustering and comparing these behavioral features, possible abnormal response patterns can be identified, such as short-term rapid answers, long-term pauses, behavioral trajectory jumps, and repeated clicks. Furthermore, cross-analysis of the behavioral features of these samples with their demographic attributes, occupational background, and other data can identify whether there are cognitive biases or misunderstandings of the questions among specific groups. Through this analysis process, not only can the direct behavioral causes of parameter fluctuations be located, but also an attribution feature set at the sample level can be established, providing a multi-dimensional chain of evidence for subsequent quality defect analysis. An analytical framework for question quality defects is constructed around multiple dimensions such as question content, structural design, option function, semantic expression, and technical interaction. First, the question stem and answer choices are analyzed for linguistic complexity to check for semantic ambiguity, long sentence structures, unclear instructions, or logical jumps. Second, the answer choices are functionally consistent to identify misleading options, double correct options, or unreasonable exclusion logic. Third, user behavior patterns are analyzed to determine if there are obstacles to answering due to factors such as interface interaction, question loading, and page responsiveness. All diagnostic conclusions are output as "defect records," recording the defect type, attribution factors, affected sample proportion, occurrence time, and suggested solutions. These defect records are automatically incorporated into the question bank management system for question revision, quality monitoring, and version updates. Through this process, question bank management shifts from a static auditing model to a dynamic, data-driven intelligent diagnostic system, ensuring the stability, reliability, and fairness of question discrimination in a dynamic assessment environment.

[0049] In this embodiment, the specific steps of step S6 are as follows: Based on the record of quality defects in the test questions, we can identify improvement strategies for the test questions. Based on the improvement strategy and the dynamic question bank quality stratification system, multi-objective optimization simulation is conducted to generate multiple sets of optimization schemes; Monte Carlo simulation tests were conducted on multiple optimization schemes to generate evaluation validity and reliability of different schemes; Based on the validity and reliability of the assessment, a comprehensive evaluation of the proposed solutions was conducted, and the optimal solution was selected. Real-time assessment questions are generated based on the optimal solution, and configuration parameters are updated.

[0050] In this embodiment, defects are categorized and prioritized across multiple dimensions based on their type, intensity, proportion of affected samples, and the ability dimension of the question within the defect record. A defect heatmap is generated by statistically analyzing the frequency of each type of defect and its weight in assessing discrimination, identifying which questions or question types are the focus of improvement. Furthermore, by analyzing the dynamic parameter trends of the questions, it is determined whether the defects are accompanied by significant deviations in specific ability levels or response behavior patterns, further refining the improvement objectives. Improvement strategies typically include text semantic optimization (e.g., simplifying the question stem, adjusting wording), option rationalization (deleting or replacing misleading options), and interaction flow optimization (improving the question presentation order, enhancing interface friendliness). Each improvement strategy is mapped to the question quality indicator space, defining optimization variables and constraints, such as maximum allowable text complexity and minimum option discrimination threshold. Subsequently, based on multi-objective optimization algorithms (e.g., genetic algorithms, particle swarm optimization), the impact of different strategy combinations on question performance indicators is simulated, generating multiple sets of optimization schemes that satisfy different weight configurations. Each set of schemes corresponds to different strategy combinations and weight allocations, covering text adjustment, option correction, question replacement, and sample distribution optimization. In this process, the simulation fully considers the hierarchical structure of the question bank and the differences in question quality within the hierarchical system, ensuring that optimization not only focuses on the effectiveness of individual questions but also takes into account the overall balance of the question bank and the coverage of multi-dimensional abilities. Through multi-objective optimization simulation, it achieves efficient exploration of complex improvement spaces, providing diverse candidates for subsequent solution selection.

[0051] After generating multiple optimized schemes, the Monte Carlo simulation method was used to comprehensively test the schemes and evaluate their assessment validity and reliability. Monte Carlo simulation repeatedly simulates the test-takers' answering process by randomly sampling the ability distribution and answering behavior model, reconstructing the assessment results data. For each scheme, thousands to tens of thousands of sample answering runs were simulated to capture the stability and adaptability of question performance under different scenarios. By comparing key indicators such as changes in question discrimination, ability estimation error, and model fit goodness of fit generated by the schemes in the simulation, the improvement in validity and reliability level were quantified. Simultaneously, the robustness of the schemes in dealing with abnormal answering patterns and behavioral fluctuations was evaluated to ensure that the schemes can adapt to the complex changes in the actual assessment environment. This step achieved closed-loop verification between the theoretically optimized schemes and actual assessment performance, ensuring that the optimization strategy is not only theoretically reasonable but also has practical value and stability. A systematic and comprehensive evaluation of the assessment validity and reliability results of each optimized scheme was conducted. The evaluation adopted a multi-dimensional indicator system, covering aspects such as the improvement in discrimination, the accuracy of ability estimation, the balance of answering time, the range of sample adaptation, and the stability of model fit. The comprehensive scoring method uses weighted aggregation technology to balance the weight relationships between different indicators, forming an overall performance score. Subsequently, the solutions are ranked and screened, prioritizing those that demonstrate balanced performance across multiple indicators and significant improvement effects. Considering the real-time and dynamic adjustment requirements of the assessment application, the adjustment costs and implementation complexity of the evaluation solutions are also incorporated into the decision-making factors. The final optimal solution not only possesses the best statistical performance but is also easy to deploy and maintain, becoming the core solution for subsequent question updates and assessment system optimization. This process ensures the scientific validity, feasibility, and practicality of the optimized solution, laying a solid foundation for the high-quality development of the dynamic question bank.

[0052] After selecting the optimal solution, the system enters the practical application phase, generating configuration parameters for real-time assessment question updates based on the solution's content. This step involves translating the optimization strategy into executable update instructions for the question bank management system, such as updating question stem version numbers, adjusting option schemes, correcting question weights, and replacing or adding rules for questions. Simultaneously, adjustment rules for relevant dynamic IRT parameters are incorporated into the parameter configuration, supporting automatic correction and dynamic evolution of parameters during subsequent assessments. To ensure the smooth and secure nature of question bank updates, the update configuration parameters also include version control, change log recording, and rollback mechanisms, enabling full-process tracking and management of the update process. The system distributes the configuration parameters to the online assessment platform via an automated interface, applying them to the testing environment in real time. This process ensures that continuous improvement in question quality can be seamlessly integrated into assessment operations, achieving efficient management and intelligent maintenance of the dynamic question bank, and improving the overall accuracy and adaptability of the assessment.

[0053] Therefore, the embodiments should be considered as exemplary and non-limiting in all respects, and the scope of the invention is defined by the appended claims rather than the foregoing description. Thus, all variations falling within the meaning and scope of the equivalents of the application are intended to be included within the invention.

[0054] The above description is merely a specific embodiment of the present invention, enabling those skilled in the art to understand or implement it. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein are implemented in other embodiments without departing from the spirit or scope of the invention. Therefore, the present invention is not to be limited to the embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novel features of the invention herein.

Claims

1. A dynamic IRT-variance joint detection algorithm for item discrimination in human resource assessment, characterized in that, Includes the following steps: Step S1: Collect the respondent's full-link behavior data based on the assessment terminal, and calculate the initial parameters of the question's IRT; Step S2: Based on the initial parameters of the IRT in the problem, perform evolutionary analysis and time series change rate analysis to obtain the dynamic IRT parameter evolution trajectory and the dynamic drift index of discrimination. Step S3: Based on the evolution trajectory of dynamic IRT parameters, evaluate the applicability of the questions and generate a multi-dimensional question evaluation report; Step S4: Automated quality stratification of the assessment question bank based on the dynamic drift index of discrimination, and construct a dynamic question bank quality stratification system; Step S5: Perform stability identification of question discrimination and analysis of question quality defects on the dynamic IRT parameter evolution trajectory to obtain question quality defect records; Step S6: Based on the question quality defect record and dynamic question bank quality stratification system, conduct multi-objective optimization simulation and comprehensive evaluation of the solution, and extract the optimal solution.

2. The dynamic IRT-variance joint detection method for item discrimination in human resource assessment according to claim 1, characterized in that, The specific steps of step S1 are as follows: Based on the collection of full-link behavior data of test takers by the assessment terminal, the test duration sequence, question modification frequency, pause distribution and mouse heatmap are calculated to generate a multimodal test behavior dataset; Calculate the behavior timestamp of the respondent's full-link behavior data; The multimodal answering behavior dataset is accurately matched and labeled based on the behavior timestamps to generate a time-series behavior matrix; Perform abnormal response pattern recognition on the time-series behavior matrix and label guessing answers and random answer samples; After removing guessing and random answers, a standardized answer behavior matrix is ​​obtained. Ability-level clustering analysis was performed based on a standardized answer behavior matrix to extract the response distribution characteristics of different ability ranges; Based on the response distribution characteristics, the question difficulty coefficient, discrimination parameter, and guessing parameter are calculated, and the initial parameters of the question IRT are obtained by fitting.

3. The dynamic IRT-variance joint detection method for item discrimination in human resource assessment according to claim 1, characterized in that, The specific steps of step S2 are as follows: Based on the initial parameters of the problem's IRT, joint vectorization modeling is performed to construct a high-dimensional parameter representation space; Markov chain Monte Carlo sampling is performed on the high-dimensional parameter representation space to generate posterior distribution estimation coefficients of the parameters. Calculate the confidence interval and peak probability density of the discrimination parameter based on the posterior distribution estimation coefficients of the parameters; Based on the confidence interval and the peak probability density, a reliability assessment of parameter estimation is performed to obtain an estimated reliability value; Based on the estimated reliable values, the initial parameters of the problem's IRT are corrected in real time and the time series change rate is analyzed to obtain the dynamic drift index of discrimination.

4. The dynamic IRT-variance joint detection method for item discrimination in human resource assessment according to claim 3, characterized in that, The specific steps for obtaining the dynamic drift index of discrimination by real-time correction and time series change rate analysis of the initial parameters of the IRT based on the estimated reliable value are as follows: The initial parameters of the problem's IRT are corrected in real time based on the estimated reliable values ​​to obtain the corrected initial parameters of the IRT. Convergence diagnosis is performed on the modified IRT initial parameters to obtain convergence diagnosis results, which include the Gehrman-Rubin statistic and the effective sample size. The convergence diagnosis results are adaptively adjusted by adjusting the sampling step size to generate a dynamic IRT parameter evolution trajectory. The dynamic IRT parameter evolution trajectory is analyzed by time-series change rate analysis to obtain the dynamic drift index of discrimination.

5. The dynamic IRT-variance joint detection method for item discrimination in human resource assessment according to claim 1, characterized in that, Step S3 is as follows: Based on the response distribution characteristics and the evolution trajectory of dynamic IRT parameters, the maximum capability is estimated to generate the subject's capability estimate. Hierarchical variance decomposition was performed on the estimated ability values ​​of the test subjects to generate the total variance, between-group variance and within-group variance components. The hierarchical variance contribution of different questions is calculated based on the overall variance, between-group variance, and within-group variance components, generating the variance contribution of each question. Based on the variance contribution, the applicability of the questions is rated, and a multi-dimensional question evaluation report is generated.

6. The dynamic IRT-variance joint detection method for item discrimination in human resource assessment according to claim 5, characterized in that, The specific steps for rating the suitability of questions based on the variance contribution and generating a multidimensional question evaluation report are as follows: Based on the variance contribution, quality heterogeneity is identified to obtain high-discrimination questions and low-discrimination questions; The synergistic effect analysis of high-discrimination and low-discrimination questions was conducted to obtain the functional complementarity and redundancy characteristics between questions, and a question synergistic effect network was constructed. Calculate the discrimination of high-discrimination questions and low-discrimination questions for different ability levels, and generate local discrimination curves; Extreme point detection and inflection point identification are performed based on local discrimination curves, and the effective scope of discrimination is marked. Question applicability rating is performed based on the effective scope of discrimination and the question synergy network, generating a multi-dimensional question evaluation report.

7. The dynamic IRT-variance joint detection method for item discrimination in human resource assessment according to claim 1, characterized in that, The specific steps of step S4 are as follows: Construct an IRT path discrimination evaluation function based on the discrimination dynamic drift index; Construct a variance path discrimination evaluation function based on a multidimensional question evaluation report; The IRT path discrimination evaluation function is used to perform weighted fusion based on the variance path discrimination evaluation function to obtain the fusion evaluation result; The fusion evaluation results are standardized and mapped to obtain the standardized discrimination score. Based on standardized discrimination scores, question thresholds are defined, and a four-level discrimination judgment standard is constructed. Based on a four-level discrimination criterion, an automated quality stratification system for the assessment question bank is constructed, which is then used to build a dynamic question bank quality stratification system.

8. The dynamic IRT-variance joint detection method for item discrimination in human resource assessment according to claim 1, characterized in that, The specific steps of step S5 are as follows: The temporal entropy value of the problem is obtained by calculating the evolution trajectory of the dynamic IRT parameters; Based on the temporal entropy value of the question, a stable and fluctuating question is identified by discrimination. Perform volatility attribution analysis and abnormal pattern diagnosis on volatility issues, and extract volatility attribution response sample characteristics; Based on the characteristics of the fluctuating attribution response samples, the question quality defects are analyzed to obtain question quality defect records.

9. The dynamic IRT-variance joint detection method for item discrimination in human resource assessment according to claim 1, characterized in that, The specific steps of step S6 are as follows: Based on the record of quality defects in the test questions, we can identify improvement strategies for the test questions. Based on the improvement strategy and the dynamic question bank quality stratification system, multi-objective optimization simulation is conducted to generate multiple sets of optimization schemes; Monte Carlo simulation tests were conducted on multiple optimization schemes to generate evaluation validity and reliability of different schemes; Based on the validity and reliability of the assessment, a comprehensive evaluation of the proposed solutions was conducted, and the optimal solution was selected. Real-time assessment questions are generated based on the optimal solution, and configuration parameters are updated.