Psychological assessment validity dynamic verification algorithm and system based on multivariate variance analysis
The dynamic validity verification algorithm for psychological tests using multivariate analysis of variance solves the problems of dynamic changes and multidimensional factor interactions in traditional psychological test validity verification methods. It enables real-time monitoring and optimization of test results, thereby improving the reliability and comparability of the test results.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-25
- Publication Date
- 2026-04-03
AI Technical Summary
Traditional methods for validating the validity of psychological tests are unable to capture the dynamic changes in an individual's psychological state, leading to unstable test results. Furthermore, they lack systematic research on the interaction of multidimensional factors, affecting the comprehensiveness and accuracy of validity analysis.
A dynamic validity verification algorithm for psychological tests based on multivariate analysis of variance is adopted. By collecting test record information, a test stability curve is generated, multi-dimensional score difference analysis is performed, a dynamic validity response curve is constructed, and validity degradation interval analysis and recalibration are carried out to achieve real-time self-optimization of credibility.
Identify and correct assessment results affected by psychological fatigue, emotional fluctuations, etc., improve the quality of assessment data, realize real-time monitoring and optimization of assessment validity, and ensure the reliability and comparability of assessment results.
Smart Images

Figure CN121789985A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of psychological test validity analysis, and in particular to a dynamic verification algorithm and system for psychological test validity based on multivariate analysis of variance. Background Technology
[0002] With the rapid development of artificial intelligence, big data analytics, and educational assessment technologies, psychological assessment systems have been widely applied in various fields such as educational selection, human resource management, vocational competency assessment, and mental health screening. Psychological assessment results have significant reference value in talent selection and individual decision-making, and their scientific rigor and accuracy directly affect the credibility and fairness of the assessment conclusions. However, against the backdrop of increasingly diverse assessment subjects and more complex assessment environments, traditional methods for validating the validity of psychological assessments are gradually revealing their limitations, making it difficult to meet the needs of dynamic and intelligent assessment scenarios.
[0003] The validity of psychological testing is a key indicator of whether a testing tool can accurately reflect the psychological characteristics of the test taker. Existing validity testing methods typically rely on static statistical models, such as correlation analysis, reliability coefficient testing, and factor analysis. These methods are often based on the assumptions of a fixed sample and constant item parameters, neglecting the dynamic changes in the test taker's psychological state, response behavior, and external environment during the testing process. In practical applications, an individual's response state may be affected by various factors such as fatigue, emotional fluctuations, and environmental interference, leading to fluctuations and instability in the test data, thus causing the test validity to change over time. Traditional static testing methods struggle to capture these dynamic changes in validity, easily resulting in distorted or biased test results.
[0004] Furthermore, most existing validity analyses focus on single-dimensional outcome analysis, lacking systematic research on the interactions between multiple factors. For example, variance differences between different assessment dimensions, different groups, or different assessment batches often reflect the validity fluctuations of assessment tools in different contexts. However, because traditional methods struggle to simultaneously handle multi-factor interactions, assessment systems cannot accurately assess the sources and contribution proportions of validity changes, affecting the comprehensiveness and accuracy of validity analyses. Summary of the Invention
[0005] To address the aforementioned technical problems, this invention proposes a dynamic validity verification algorithm and system for psychological tests based on multivariate analysis of variance, thereby resolving at least one of the aforementioned technical issues.
[0006] To achieve the above objectives, this invention provides a dynamic validity verification algorithm for psychological tests based on multivariate analysis of variance, comprising the following steps: Step S1: Collect the assessment record information of the assessment subjects, conduct a response stability assessment, and generate an assessment state stability curve; Step S2: Extract the group assessment report, perform multi-dimensional question score calculation and score difference analysis, and generate assessment score difference value; Step S3: Based on the stability curve of the assessment state and the difference in assessment scores, conduct a validity contribution response analysis and construct a dynamic validity response curve; Step S4: Perform validity degradation interval analysis on the dynamic validity response curve, and perform dynamic recalibration to extract the validity stability optimization curve; Step S5: Perform multivariate ANOVA and normalization calibration based on the validity stability optimization curve to generate equivalent dynamic calibration; Step S6: Perform dynamic validity self-verification and real-time reliability self-optimization based on equivalent dynamic calibration.
[0007] This specification provides a dynamic validity verification system for psychological tests based on multivariate analysis of variance, used to execute the dynamic validity verification algorithm for psychological tests based on multivariate analysis of variance as described above, including: The response status assessment module is used to collect assessment record information of the assessment subjects, assess the stability of their responses, and generate an assessment status stability curve. The scoring difference module is used to extract group assessment reports, perform multi-dimensional question score calculation and score difference analysis, and generate assessment score difference values. The validity response module is used to perform validity contribution response analysis based on the assessment stability curve and the difference in assessment scores, and to construct a dynamic validity response curve. The recalibration module is used to analyze the validity degradation interval of the dynamic validity response curve, perform dynamic recalibration, and extract the validity stability optimization curve. The normalization calibration module is used to perform multivariate ANOVA and normalization calibration based on the validity stability optimization curve, and generate equivalent dynamic calibration. The validity self-verification module is used for dynamic validity self-verification and real-time reliability self-optimization based on equivalent dynamic calibration.
[0008] The beneficial effects of this invention are as follows: By collecting information such as answering time, question switching frequency, and changes in answering rhythm, it is possible to identify whether the test taker's answering state is stable, thereby avoiding the influence of non-ability factors such as attention fluctuations and fatigue on the assessment results. The generated "assessment state stability curve" can intuitively reflect the trend of the test taker's psychological state changes throughout the assessment process, providing dynamic basic data for subsequent validity analysis. By filtering out abnormal fluctuation intervals through the stability curve, invalid answers or cheating behaviors can be identified in advance, improving the quality of assessment data from the source. By extracting multi-dimensional scores (such as cognitive, emotional, and social dimensions) at the group level, hierarchical analysis of different assessment elements can be achieved. Score difference analysis can reveal systematic differences within groups or between individuals in specific dimensions, helping to identify bias items or validity risks in assessment questions. The generated "assessment score difference value" serves as an important indicator of consistency within the group, providing data support for subsequent validity response analysis. Combining assessment state stability (process factor) with score difference (outcome factor) can more comprehensively characterize assessment validity. By constructing a "dynamic validity response curve," the trend of validity changes over time or during the assessment process can be quantified, enabling real-time monitoring of validity. The curve's range of change reveals which stages of responses or questions contribute most to overall validity, providing a reference for subsequent optimization of the question bank design. Analysis of the degradation range of the validity response curve allows for timely identification of the causes of decreased assessment validity (such as psychological fatigue, emotional fluctuations, or external interference). Dynamic correction of the degradation range using a recalibration algorithm ensures that overall assessment validity remains within a stable range. The extracted "validity stability optimization curve" smooths out abnormal fluctuations, improving the reliability and generalization ability of the validity model. Multivariate analysis of variance (MANOVA) is used to extract validity influence weights from multi-dimensional interaction factors, achieving high-dimensional validity structure modeling. Normalization calibration transforms validity indicators from different dimensions into a unified scale, eliminating dimensional differences between assessment tools. This benchmark can be used for validity alignment across different assessment batches, groups, or environments, ensuring horizontal comparability and vertical traceability of assessment results. Based on real-time data feedback, it can automatically detect validity drift and perform self-correction, thereby maintaining the long-term stability of the assessment model. Through self-optimizing algorithms (such as Bayesian update or recursive least squares), it dynamically optimizes the assessment reliability parameters, achieving intelligent validity maintenance. The final output "real-time reliability" index can serve as an important basis for the credibility of psychological assessment reports, providing immediate reference for applications in education, recruitment, and psychological intervention. Attached Figure Description
[0009] Figure 1 This is a flowchart illustrating the steps of a dynamic validity verification algorithm for psychological assessment based on multivariate analysis of variance according to the present invention. Figure 2 This is a detailed flowchart illustrating the implementation steps of step S1. Figure 3 This is a flowchart illustrating the detailed implementation steps of step S2. Detailed Implementation
[0010] It should be understood that the specific embodiments described herein are for illustrative purposes only and are not intended to limit the scope of the invention.
[0011] This application provides a dynamic validity verification algorithm and system for psychological tests based on multivariate analysis of variance. The executing entities of the dynamic validity verification algorithm and system for psychological tests based on multivariate analysis of variance include, but are not limited to, mechanical equipment, data processing platforms, cloud server nodes, network upload devices, etc., which can be considered as general computing nodes in this application. The data processing platform includes, but is not limited to, at least one of an audio-visual management system, an information management system, and a cloud-based data management system.
[0012] Please see Figures 1 to 3 This invention provides a dynamic validity verification algorithm for psychological tests based on multivariate analysis of variance, comprising the following steps: Step S1: Collect the assessment record information of the assessment subjects, conduct a response stability assessment, and generate an assessment state stability curve; Step S2: Extract the group assessment report, perform multi-dimensional question score calculation and score difference analysis, and generate assessment score difference value; Step S3: Based on the stability curve of the assessment state and the difference in assessment scores, conduct a validity contribution response analysis and construct a dynamic validity response curve; Step S4: Perform validity degradation interval analysis on the dynamic validity response curve, and perform dynamic recalibration to extract the validity stability optimization curve; Step S5: Perform multivariate ANOVA and normalization calibration based on the validity stability optimization curve to generate equivalent dynamic calibration; Step S6: Perform dynamic validity self-verification and real-time reliability self-optimization based on equivalent dynamic calibration.
[0013] In the embodiments of the present invention, see Figure 1 The diagram below illustrates the steps of a dynamic validity verification algorithm for psychological tests based on multivariate analysis of variance (VAV) according to the present invention. In this example, the steps of the dynamic validity verification algorithm for psychological tests based on multivariate analysis of variance include: Step S1: Collect the assessment record information of the assessment subjects, conduct a response stability assessment, and generate an assessment state stability curve; In this embodiment, structured data is collected from the entire response process of the psychological assessment subjects. This includes start and end times, duration of each question, option modification trajectory, submission intervals, and interaction information such as mouse and touch input. The system constructs a response record stream model using a time-series approach, continuously mapping the temporal characteristics of the response process using the response time and response sequence of each question as nodes. Next, the system calculates behavioral parameters such as average response speed, modification rate, and variance of dwell time across different time periods based on a time-series sliding window algorithm. Through multidimensional statistical analysis of these indicators, the system constructs a response stability function to measure the consistency of an individual's behavior and the level of attentional fluctuation during the assessment process. A continuous decrease in the stability function value indicates significant fluctuations in the test subject's response behavior. The system smooths and normalizes the stability values of all time-series nodes, forming a continuous assessment state stability curve. This curve dynamically reflects changes in an individual's psychological state and level of focus during the assessment process, providing a dynamic temporal basis for subsequent validity response analysis.
[0014] Step S2: Extract the group assessment report, perform multi-dimensional question score calculation and score difference analysis, and generate assessment score difference value; In this embodiment, assessment report data is extracted from the group dimension to calculate the multidimensional structural features and dissimilarity indicators of item scores. Assessment results from different test subjects are retrieved from the database, and the raw scores for each item are standardized to eliminate the influence of differences in scoring ranges and item difficulty. Then, the system performs hierarchical aggregation of score data based on the content category, dimensional attributes, and cognitive level requirements of the items. By calculating the mean, standard deviation, and distribution skewness of each item in the group sample, the system obtains the central trend and fluctuation range of the group scores for each item. Based on this, a score dissimilarity index is calculated using the ratio of between-group variance to within-group variance to measure the degree of score dispersion among different groups on a specific assessment dimension. When the dissimilarity index is higher than a set threshold (e.g., 0.15), the system determines that the item has significant group dissimilarity. Subsequently, the system aggregates the score dissimilarity indices of each item by assessment dimension to form a comprehensive assessment score dissimilarity value. This value reflects the overall deviation strength in psychological characteristic responses between different groups or different assessment stages, providing key input parameters for subsequent dynamic correlation analysis of validity response.
[0015] Step S3: Based on the stability curve of the assessment state and the difference in assessment scores, conduct a validity contribution response analysis and construct a dynamic validity response curve; In this embodiment, the stability curve of the assessment state and the score difference value are first synchronized in time and matched in scale to ensure the correspondence of the two types of data on the time axis. Then, the system uses multiple regression and variance decomposition methods to model the collaborative change trend of stability and score difference, extracting three types of validity contribution coefficients: content validity, construct validity, and response validity. Content validity reflects the strength of the association between the item content and the target psychological construct; construct validity measures the logical consistency between assessment dimensions; and response validity characterizes the degree of matching between individual behavioral responses and psychological states. The system uses a dynamic weight adjustment strategy to update the changes of each validity coefficient over time in real time, generating a multidimensional validity response sequence. Subsequently, a smoothing fitting function is used to continuously process the validity coefficient curve, forming a dynamic validity response curve. This curve not only shows the evolution trend of assessment validity throughout the process but also reveals the fluctuation pattern of validity contribution in different time periods, providing a continuous response basis for subsequent degradation analysis and validity optimization.
[0016] Step S4: Perform validity degradation interval analysis on the dynamic validity response curve, and perform dynamic recalibration to extract the validity stability optimization curve; In this embodiment, based on the dynamic validity response curve, potential validity decay regions are identified and corrected. First and second derivatives are calculated on the validity response curve to extract local slope and curvature information of validity changes. When the local slope continuously decreases and the curvature sign changes, the system identifies this time period as a potential validity degradation interval. Subsequently, the system calculates a validity degradation index based on the length, amplitude, and impact dimension of the degradation interval to quantify the severity of validity decline. For the detected degradation intervals, the system employs a dynamic recalibration mechanism, comparing the validity mean and standard deviation of the non-degradation intervals to readjust the baseline value and fluctuation amplitude of the degradation segment. This recalibration process combines weighted smoothing and local regression fitting methods to ensure the continuity and authenticity of the validity recovery curve. The recalibrated validity curve is then renormalized to generate a validity stability optimization curve. This curve reflects the adjusted dynamic characteristics of validity, eliminating biases caused by unstable behavior or short-term fluctuations.
[0017] Step S5: Perform multivariate ANOVA and normalization calibration based on the validity stability optimization curve to generate equivalent dynamic calibration; In this embodiment, multivariate analysis of variance (MANOVA) is used to systematically compare the validity stability optimization curves of different groups or individuals, extracting the main effects and interaction effects of the validity structure, thereby establishing a cross-group consistent equivalence calibration model. A multivariate analysis matrix is constructed using the validity stability optimization curve data with time nodes as observed variables and validity type as the dependent variable. Subsequently, ANOVA is performed on different groups (or different assessment conditions) to calculate key indicators such as Wilks' Lambda, Pillai's Trace, and F-statistic to identify significant group effects and interaction effects. When the significance level is less than 0.05, the system determines that there are differences in validity structure. To achieve validity consistency, the system uses a standardized equivalence transformation method to linearly normalize the validity curves of different groups, ensuring that their variance ratio is controlled within a set range (e.g., 1 ± 0.05). Furthermore, the system introduces polynomial fitting and Sigmoid function mapping to dynamically correct nonlinear biases, ensuring the comparability of the calibrated validity response across the entire time domain.
[0018] Step S6: Perform dynamic validity self-verification and real-time reliability self-optimization based on equivalent dynamic calibration.
[0019] In this embodiment, the system continuously monitors response behavior characteristics, item responses, and validity trends during the assessment process, and compares them in real time with the baseline curve of the equivalent dynamic calibration model. When the system detects that the validity deviation rate exceeds a preset threshold (e.g., 0.08) or the credibility drops to a critical value (e.g., 0.75), it immediately triggers a self-verification process. The self-verification module re-estimates the validity parameters of the current stage through variance reanalysis and weighted regression algorithms, and locally corrects the item presentation strategy, time interval, and weight distribution. Simultaneously, the credibility self-optimization module adjusts the assessment content according to the direction of validity deviation, such as inserting high-focus items when attention declines, or adding stability measurement items when detecting emotional fluctuations. The system also uses dynamic Kalman filtering to smooth and predict real-time validity data, ensuring the continuity and robustness of feedback. Finally, the system forms a dynamic self-verification loop based on equivalent calibration, realizing real-time correction of validity and adaptive improvement of credibility, thereby constructing a high-precision and highly reliable dynamic validity verification system for psychological assessment.
[0020] In this embodiment, see Figure 2 The diagram below illustrates the detailed implementation steps of step S1. In this embodiment, the detailed implementation steps of step S1 include: Collect assessment record information from the assessment subjects; Based on the assessment record information, the response time, option modification trajectory, and question dwell time distribution are calculated to obtain multi-dimensional assessment response data; The multi-dimensional evaluation response data is processed by time-series segmentation to extract evaluation response data from multiple time periods; Based on the evaluation response data, the answer behavior characteristics are analyzed for each time period to generate the answer behavior pattern for each time period; Behavioral fluctuation analysis is performed on the aforementioned response behavior pattern, and response stability is assessed to generate an evaluation state stability curve.
[0021] In this embodiment, raw assessment record information generated by the test taker during the response process is collected from the psychological assessment platform. This information includes, but is not limited to: test taker identification, assessment questionnaire ID, question number, response timestamp, option submission record, option modification log, page dwell time, mouse and keyboard interaction trajectory, etc. Data collection is achieved through an embedded front-end event listening module, which automatically records every operation behavior of the test taker during the response. To ensure data accuracy, the sampling frequency is set to 50 Hz, that is, the interaction state is recorded once every 20 milliseconds.
[0022] In terms of data structure, the raw behavior flow is encapsulated in JSON format, with each record containing fields such as {question_id, timestamp, selected_option, modified_flag, dwell_time, cursor_position}. Data is cached locally and then uploaded to the server for centralized processing. To prevent interference from outliers, the data collection module incorporates an automatic detection mechanism that automatically marks records with overdue answers (exceeding the question's time limit by three times) or missing behavior trajectories as "invalid records" and removes them from subsequent statistical analysis.
[0023] In the experimental phase, 300 university students were selected as the participants, each completing standardized psychological assessment scales (such as the EPQ and SCL-90). Each questionnaire contained 90 questions, with an average response time of approximately 15 minutes. The recorded data amounted to approximately 2.7 million behavioral log entries, providing a sufficient basic sample for subsequent response data calculation and behavioral pattern analysis. The key objective of this step was to obtain fine-grained, high-fidelity behavioral data to support subsequent dynamic psychological state modeling and validity assessment.
[0024] The collected raw assessment records are characterized to generate multi-dimensional assessment response data for analysis. First, the response time (RT) for each question is calculated based on the timestamp difference, using the following formula: RT(i) = T_submit(i) - T_start(i).
[0025] Then, the trajectory of participants' option modifications on the same question was statistically analyzed, including the number of modifications made in a single response, the direction of modification (from extreme options to neutral options or the opposite direction), and the time interval between modifications. This part was achieved through time series analysis of the modified_flag field.
[0026] Next, the dwell time distribution for each item was calculated, which is the relative proportion of time spent on different items in the questionnaire, to characterize the participants' attention distribution pattern. Quantile normalization was used to standardize the dwell time for each item to the [0,1] interval to facilitate cross-individual comparisons.
[0027] After calculation, the three types of data (RT, modified trajectory, and dwell time) are combined into a multidimensional response vector R = [RT, MT, DT]. In the experimental environment, to reduce the influence of noise, a sliding window smoothing process (window width w = 3 questions) is used to eliminate extreme values. The final generated data dimension is N × 3 (N is the number of questions), forming the dynamic behavioral response matrix of the test takers throughout the assessment process. This matrix provides the basic input for subsequent validity verification based on multivariate analysis of variance (MANOVA), enabling the algorithm to simultaneously capture the changing trends of the test takers' states in both the time and behavioral dimensions. The sliding variance σ(t) of the response time series is calculated, and when the variance exceeds 1.5 times the global mean, it is determined as a state change point. In this way, several key moments are automatically identified, and the entire assessment process is divided into 3–5 time periods, corresponding to psychological state stages such as the "initial adaptation period," "stable response period," and "final fatigue period."
[0028] To ensure the statistical validity of each time segment, a minimum duration threshold of 60 seconds was set, and each segment contained at least 15 questions. After segmentation, statistical features such as average response time, modification frequency, and attention concentration were extracted for each segment. The experiment observed that individuals generally experienced a 20% increase in response time and a 15% increase in option modification rate in later segments, demonstrating a typical fatigue effect.
[0029] The core advantage of this segmented processing lies in its ability to perform localized analysis of assessment data based on the time dimension. This allows the algorithm to move beyond relying solely on global average indicators and dynamically reflect the psychological fluctuations of the test-takers during the assessment process. The segmented data will then be fed into the next step, the behavioral feature analysis module, to further generate time-series-based response behavior patterns.
[0030] For each time period, a set of response characteristic indicators are calculated, including: average response time, modification rate, standardized attention index, option consistency index (i.e., the average value of option variation), reaction time coefficient of variation, etc.
[0031] Then, Principal Component Analysis (PCA) was used to reduce the dimensionality of these features to extract the principal components that best represent the differences in participants' response behaviors. By setting the retention rate to 85% of the cumulative variance, 2–3 principal dimensions were typically obtained, such as "response speed factor," "prudence factor," and "stability factor." Next, K-means clustering (K=3, determined by the elbow method) was used to cluster the dimensionality-reduced feature space to classify the participants' response behavior patterns over the given time period. Each pattern represents a psychological characteristic state, such as: quick decision-making, cautious revision, and hesitant repetition.
[0032] In the experiment, statistical analysis of the behavioral patterns of 300 participants revealed that approximately 60% of individuals maintained a stable response pattern during the mid-stage, while 35% shifted to a hesitant behavioral pattern in the later stage, indicating the prevalence of assessment fatigue. The behavioral pattern vector BM(t) for each time period was ultimately output as a time-series representation of psychological and behavioral characteristics. This process not only revealed the participants' cognitive processing methods at different stages but also provided quantifiable input features for subsequent stability analysis.
[0033] The Euclidean distance D(t,t+1) between behavioral pattern vectors in adjacent time periods is calculated to quantify the magnitude of pattern variation. Then, the overall stability of the subjects' response behavior is measured by calculating the Fluctuation Coefficient (FC), which is defined as the square root of the mean of the squared differences between each segment.
[0034] Next, multivariate analysis of variance (MANOVA) was used to test the significance of the mean vectors of behavioral characteristics at different time periods to determine whether there were statistically significant differences in response behavior between different stages. When Wilks' Lambda value was less than 0.05, it indicated that there were significant stage-specific changes in the participants' behavior, suggesting that the validity of the responses might be affected by state fluctuations. The fluctuation coefficient was further processed exponentially (α=0.3) and plotted as a stability curve. This curve, with time on the horizontal axis and stability value on the vertical axis, reflects the dynamic evolution of the participants' psychological state throughout the assessment process.
[0035] In the experimental results, the curves of stable individuals tended to be smooth with fluctuations less than 0.15, while the fluctuation coefficients of individuals with emotional or attentional fluctuations exceeded 0.25. This analysis allows for the dynamic assessment of the stability of the test validity, providing a basis for determining the authenticity of psychological test data and detecting anomalies. The final output stability curve can be directly used in psychological state monitoring or dynamic validity calibration modules, enabling real-time quantitative assessment of psychological test validity at the behavioral level.
[0036] In this embodiment, see Figure 3The diagram below illustrates the detailed implementation steps of step S2. In this embodiment, the detailed implementation steps of step S2 include: Collect group evaluation reports; Based on the group assessment report, the assessment subjects are identified, and multi-dimensional question scores are calculated to extract the assessment scores of different subjects in multiple dimensions; the multi-dimensional assessment scores include the original score matrix, reverse question score transformation records, and missing value imputation trajectory; Based on the group assessment report, the question difficulty parameter, discrimination parameter, and guessing coefficient are calculated to obtain the question complexity. Based on the complexity of the assessment questions, the assessment scores of the multiple dimensions are grouped and the score difference is calculated to generate an assessment score difference value.
[0037] In this embodiment, a front-end behavior capture module and a back-end data acquisition module work together to comprehensively record the test subjects' response behavior during the psychological assessment process. The front-end module embeds an event listener program into the test subject's interface to capture information such as mouse clicks, option switching, page dwell time, keyboard input, and scrolling operations. Each operation is automatically timestamped and cached in real time, forming high-precision time-series data. A sampling interval of 20 milliseconds is set to ensure the temporal integrity of the interactive data and the continuity of behavioral detail capture. All collected data includes fields such as question number, start time, submission time, option status, modification identifier, dwell time, and behavior coordinates, and is uploaded to a server for centralized storage via a secure encrypted channel. To ensure data quality, an anomaly detection mechanism is configured during the acquisition process to identify and label missing records, timed responses, and data interruptions in real time, and invalid records are automatically removed during the data preprocessing stage. Finally, the raw behavioral data is saved in a structured format (such as JSON or a tabular database) to form a complete raw assessment log file, providing basic data support for subsequent response calculation, temporal modeling, and validity analysis. The response time (RT) for each question is calculated based on its start and submission times, reflecting the participant's reaction speed at the single-question level. Next, option modification trajectory data (MT) is generated by analyzing option modification identifiers and time series information, including the number of modifications, modification direction, and interval between modifications, reflecting the participant's hesitation and reflection characteristics. Subsequently, the page dwell time (DT) for each question is calculated and normalized to the [0,1] interval, thus obtaining the participant's attention distribution characteristics throughout the entire assessment. To eliminate the influence of occasional anomalies, a sliding window smoothing method (window width can be set to 3 questions) is used to smooth and correct the reaction time and dwell time series of consecutive questions. Finally, a response matrix R = [RT, MT, DT] is formed, consisting of three types of features: RT, MT, and DT, representing the participant's multidimensional behavioral response characteristics throughout the assessment process in matrix form. This matrix structure provides a calculable input data foundation for subsequent time-series segmentation, behavioral modeling, and multivariate variance analysis.
[0038] The multidimensional response time series of the participants is dynamically segmented to identify changes in response behavior characteristics at different stages of the assessment process. First, the moving variance series of response duration and dwell time is calculated to quantify the temporal variation in the stability of the participants' responses. When the local variance value exceeds a threshold set by the global mean (e.g., 1.5 times), that moment is determined as a state inflection point. By continuously detecting inflection points and combining them with a minimum time limit (e.g., 60 seconds or 15 questions), the entire assessment process is adaptively divided into several consecutive time periods. The response data within each time period is extracted separately to form an independent data subset for subsequent behavioral analysis. To avoid over-segmentation, a smoothing judgment mechanism is introduced, automatically merging consecutive fluctuation intervals into a single interval when the interval is too short (e.g., less than 30 seconds). Through this adaptive segmentation strategy, the assessment behavior can be decomposed into representative stage segments based on the time dimension, reflecting the state change characteristics of the participants from the initial response to the later response. The structure of each data segment remains consistent with the original response matrix, providing staged input for subsequent behavioral feature analysis. Within each time period, statistics such as mean response duration, variance, modification rate, option consistency, coefficient of variation of response, and concentration of dwell time are calculated to construct a time-period-level feature vector. Subsequently, to reduce feature redundancy and noise interference, principal component analysis (PCA) is used for dimensionality reduction, retaining principal components that explain more than 85% of the total variance. The dimensionality-reduced feature space reflects the main response characteristic dimensions of the participants, such as speed factors, prudence factors, and stability factors. Next, clustering algorithms (such as K-means) are used to classify the time-period features, dividing similar response patterns into several behavioral categories. Each category represents a typical response pattern, such as quick decision-making, cautious revision, or hesitant and iterative. Corresponding pattern labels are assigned to each time period, and time-series-based behavioral pattern trajectories are generated for subsequent dynamic change analysis. Through this process, a structured mapping from raw response characteristics to behavioral patterns is achieved, enabling a quantitative description of the evolution of participants' response characteristics at different stages of the assessment process.
[0039] The Euclidean distance or cosine similarity of behavioral pattern vectors between adjacent time periods is calculated to measure the magnitude of pattern changes between time periods. Subsequently, the overall stability of responses is quantified by calculating the Fluctuation Coefficient (FC), defined as the root mean square of the squared differences between each segment. To determine whether behavioral differences are statistically significant, a multivariate analysis of variance (MANOVA) model is used to test the significance of the mean vectors of behavioral characteristics for each time period, and the Wilks' Lambda statistic is calculated. When this value is below a preset threshold (e.g., 0.05), significant stage-specific differences in response behavior are considered, indicating potential fluctuations in response validity. Finally, a stability curve of the assessment state is plotted based on the time-series fluctuation index, with time on the horizontal axis and stability value on the vertical axis. The data is processed using an exponential smoothing algorithm (smoothing coefficient α=0.3). The continuity and fluctuation amplitude of the curve can intuitively reflect the psychological stability of the subjects and the reliability of the assessment process, forming a visualized result of dynamic verification of assessment validity, achieving quantitative analysis and dynamic evaluation of assessment validity from a behavioral perspective.
[0040] In this embodiment, step S3 includes the following steps: Extract the identity information of the evaluation subjects; match the evaluation state stability curves based on the evaluation subject identity information, and extract the state stability curve for each subject; Based on the state stability curve and the difference in evaluation scores, feature alignment is performed to generate a state temperature-score difference collaborative matrix. Define the dependent variable matrix, including item factors, state factors, and time factors; Based on the dependent variable matrix, the state temperature-score difference co-factor matrix is decomposed into factor interaction terms to obtain different validity contribution rates, including content validity, construct validity, and response validity. Dynamic validity response curves are constructed based on different validity contribution rates.
[0041] In this embodiment, the participant's identity information is extracted from the assessment database, including but not limited to individual ID, assessment time, scale type, gender, age, and assessment scene identifier. All identity information is hashed and encrypted to protect privacy. Subsequently, an index table is built based on the individual ID, and the identity information is matched with the state stability curve generated in the previous stage. The stability curve is stored in time series form, containing the stability value S(t) at each time point, along with an assessment stage label. After matching, the complete set of stability curves corresponding to each participant is extracted to form an individualized dynamic stability dataset. To avoid matching errors, a dual verification mechanism is adopted, that is, the curve attribution is verified by both the assessment timestamp and the scale ID, thereby ensuring data consistency and accuracy. Through this step, a mapping relationship between individual identity and dynamic stability characteristics is established, providing basic data support for subsequent cross-object comparisons and validity factor decomposition, and realizing individualized association management of the assessment data layer. The state stability curve of each participant is normalized, and the stability value S(t) is linearly converted into a state temperature T(t) to more intuitively reflect the trend of psychological state changes. The conversion formula uses an interval mapping method, mapping low stability values to a high "temperature" state, representing increased psychological fluctuations. Subsequently, the score difference ΔS between adjacent assessment phases or different assessment dimensions is calculated as a structured indicator reflecting changes in assessment results. To ensure comparability across different scales, the score difference values are Z-score standardized.
[0042] Based on this, a collaboration matrix M_c is constructed to represent the state temperature and score difference, where rows represent time series nodes and columns represent items or assessment dimensions. Each element M_c(i,j) in the matrix represents the collaboration strength between the state temperature and score difference of the j-th assessment dimension at time node i. To enhance the stability of the matrix, a moving weighted average method (weight decay coefficient α=0.7) is introduced during the generation process to reduce the interference of sudden fluctuations on the overall feature alignment. The final generated M_c matrix is the core feature expression of the state-score difference interaction, providing an input basis for subsequent factor decomposition based on multivariate analysis of variance, and realizing the synchronous characterization of psychological dynamic changes and assessment result differences.
[0043] Based on the assessment content and response characteristics, the key dimensions affecting assessment validity are divided into three categories of factors: item factors, state factors, and time factors. Item factors reflect the attributes of the assessment content, including question type, measurement dimensions, difficulty level, and number of options; state factors correspond to the degree of fluctuation in an individual's psychological state during the response process, and their values are derived from the state temperature sequence T(t); time factors represent the dynamic evolution characteristics of assessment behavior over time, which can be represented by time period division labels or assessment stage numbers.
[0044] The three types of factors are encoded into a dependent variable matrix Y in matrix form, where each row represents a response event or item response sample, and the column vectors correspond to the values of the three types of factors. To achieve numericalization and unified analysis, one-hot encoding or standardized numerical transformation methods are used to map discrete variables into a continuous feature space. Simultaneously, factor independence tests (e.g., variance inflation factor VIF < 5) ensure that there is no significant multicollinearity among the three types of factors, thus satisfying the statistical assumptions of multivariate analysis of variance. Through the definition and construction of this dependent variable matrix, a comprehensive input structure capable of representing response characteristics, psychological state, and temporal dynamics is formed, providing a solid data foundation for the subsequent accurate decomposition of validity contribution rates. The dependent variable matrix Y constructed in the previous step is interactively modeled with the state-temperature-score difference co-operation matrix M_c, and the statistical decomposition of factor interaction terms is achieved through multivariate analysis of variance (MANOVA). First, the co-operation strength values in M_c are used as independent variable inputs, and the three types of factors in Y—items, state, and time—are used as explanatory variables to construct a multivariate variance model. The model measures the degree of significance of each factor on the synergistic feature by calculating Wilks' Lambda, Pillai's Trace, and other statistics for each factor and its interaction terms (item × state, state × time, item × time).
[0045] Subsequently, the explained variance was decomposed into different validity contribution rates according to factor type: item-related items corresponded to content validity, state-related items to response validity, and structure-related items to structure validity. To enhance the accuracy of the decomposition, a weighted covariance matrix adjustment mechanism was introduced to balance the weights of segments with uneven sample size and time distribution, ensuring the comparability of the contribution rates of each factor. The final numerical results of the three types of validity contribution rates were obtained, representing the combined contribution of the assessment content, structure, and response dimensions to overall validity under dynamic psychological state changes. Through this process, the constituent sources of assessment validity can be quantified, enabling dynamic analysis and visual modeling of validity factors.
[0046] The contribution rates of content validity, construct validity, and response validity are expanded along a time series dimension, and a time alignment algorithm is used to ensure the correspondence between different validity curves at the same time points. Subsequently, the validity series is made continuous through smooth interpolation (using cubic spline functions) to make the curve changes more stable and observable. Each curve, with time as the horizontal axis and validity contribution rate as the vertical axis, represents the dynamic changes of the three types of validity during the assessment process.
[0047] Further calculations of the slope change rate and crossover point positions of each validity curve were performed to identify the dynamic shifts in validity structure during the assessment process. For example, when the response validity curve rises rapidly while the content validity curve declines, it indicates that individuals exhibited a state-driven response tendency in later assessment stages. Finally, the three curves were normalized and superimposed to generate a comprehensive dynamic validity response curve, representing the overall trend of assessment validity changes over time. This curve not only reflects the overall stability of assessment validity but also reveals the dynamic patterns of validity composition changes with the psychological state of the test takers, providing a basis for real-time validity monitoring and algorithm adaptive calibration, and realizing visualization of dynamic validity response from the statistical level to the time series level.
[0048] In this embodiment, step S4 includes the following steps: The dynamic validity response curve is used to identify and eliminate dishonest responses, and the standardized validity response curve is extracted. The time drift rate is obtained by calculating the slope of the local temporal variation of the standardized validity response curve. The variance imbalance is obtained by calculating the structural skewness of the contribution rates of each validity function based on the dynamic validity response curve. Regional validity degradation interval analysis was conducted based on time drift rate and variance imbalance, and potential degradation intervals were marked. Dynamically recalibrate the potential degradation interval and extract the core parameters to recalibrate the weights. The validity stability of the dynamic validity response curve is optimized based on the recalibration weights of the core parameters, and the validity stability optimization curve is extracted.
[0049] In this embodiment, a complete scan of the time-series characteristics of the dynamic validity response curve is performed to extract its key statistical indicators, including the rate of change, the number of abrupt change points, the amplitude of continuous fluctuations, and the distribution of local abnormal peaks. By comprehensively judging these indicators, anomaly detection algorithms (such as the three-standard-deviation outlier detection method) are used to identify segments in the curve that exhibit significant jumps, abnormal periodic reversals, or short-period violent oscillations. These characteristics typically correspond to dishonest behaviors such as inattentiveness, random selection, or emotional interference during the response process. To avoid false rejection, each abnormal interval is reconfirmed by cross-referencing the overall trend of the curve, the fluctuation cycle, and the score stability index. When an abnormal feature simultaneously meets multiple threshold conditions (such as three consecutive segments with local fluctuation amplitudes exceeding twice the overall mean square deviation), the interval is marked as a dishonest response interval. Subsequently, time-series repair and interpolation reconstruction methods are used to repair the data of the rejected abnormal segments, and the overall curve is standardized to stabilize the validity value distribution within a unified range (such as [0,1]). The final standardized validity response curve serves as the basic input for subsequent time-series analysis and validity optimization. The standardized curve is divided into equally spaced segments according to the time series (e.g., each segment contains a fixed number of time nodes), and the slope of local linear change is calculated within each segment. This slope reflects the rate of increase or decrease in validity response within that time interval and is an important indicator for measuring the sensitivity to changes in psychological state or response validity. To ensure computational stability, a weighted least squares (WLS) method is used to linearly fit each curve segment, assigning higher weights to the most recent time nodes to improve the accuracy of capturing immediate changes. Subsequently, the average slope change rate of all segments is calculated and smoothed to obtain a continuous time drift rate curve. Positive drift rates indicate an increasing trend in validity, while negative values represent a decreasing trend. To avoid misjudgments caused by instantaneous fluctuations, a dynamic threshold range is set (e.g., the range of ±0.05 to ±0.15 is considered a normal fluctuation range). When the time drift rate continuously exceeds the threshold range, it is considered that there may be a shift in psychological state or a decline in response validity during that period.
[0050] The contribution rates of the three types of validity at each time point are vectorized to form a three-dimensional validity distribution vector V(t) = [C(t), S(t), R(t)], corresponding to content, structure, and response validity, respectively. Next, the covariance matrix Σ(t) of V(t) is calculated to characterize the synergistic changes among the dimensions. Large differences in the eigenvalues of the covariance matrix indicate a significant imbalance in the distribution of validity. The variance imbalance degree D_v is defined as the normalized result of the ratio of the largest eigenvalue to the smallest eigenvalue minus one, i.e., D_v = (λ_max - λ_min) / λ_max, used to represent the strength of the structure skewness. To further enhance the temporal continuity analysis, the variation curve of the variance imbalance degree over time, D_v(t), is calculated, and instantaneous disturbances are reduced through smoothing filtering. If D_v(t) consistently exceeds a set threshold (e.g., 0.3), a significant validity structure imbalance is considered to exist in that interval, potentially indicating that a certain type of validity factor has an excessively strong dominance over the overall assessment, affecting the stability of the comprehensive validity of the assessment. The time drift rate curve and the variance imbalance curve are synchronized along the time dimension to ensure the comparability of the two indicators at the same time point. Then, the joint change function F(t) = |dT(t) / dt| × D_v(t) is calculated, which comprehensively reflects the coupling strength between the rate of validity change and the degree of structural skewness. When F(t) exceeds a preset threshold (e.g., 0.05), that time point is determined to belong to a potential validity degradation interval. To enhance the continuity of identification, a sliding window analysis method is used to merge adjacent abnormal nodes into continuous degradation regions, and short-term fluctuations are further filtered by the duration and amplitude of the regions. Finally, a set of degradation intervals {Z_k} is output, with each interval containing start time, end time, and degradation intensity level information. The results are used to dynamically monitor validity decay during the assessment process and provide target area input for subsequent parameter recalibration. Through this analysis, the time period in which validity decline occurs can be effectively located, and local abnormal areas leading to a decrease in overall validity can be identified, thereby achieving a temporal diagnosis of the validity of psychological tests.
[0051] A dynamic recalibration process is performed on identified potential degradation intervals to restore the validity balance and stability of the curve. First, the local statistical characteristics of the main validity parameters are recalculated within each degradation interval, including the mean of content validity, variance of construct validity, and coefficient of variation of response validity. Then, a weight update model W(t) is constructed, adjusting the weights of each validity dimension inversely based on the fluctuation range within the degradation interval. Specifically, the greater the fluctuation of a validity factor, the larger the weight adjustment, with the correction coefficient calculated proportionally (e.g., adjusting weight Δw_i = β × σ_i / Σσ), where β is the adjustment coefficient used to control the overall balance rate. An iterative optimization method is used to ensure that the updated weight set satisfies the validity balance constraint (e.g., Σw_i = 1) while maintaining the continuity of the validity curve at the boundaries of degradation intervals. The final core parameter recalibration weight vector W* records the dynamic correction ratio of each validity dimension at different time periods. This weighting parameter serves as a key input for subsequent validity optimization, correcting the skewed structure and time drift effects in the curve, achieving quantitative compensation and dynamic rebalancing of the degenerate interval. The overall validity response curve is stabilized through dynamic adjustment of the core parameter weights, generating an optimized validity curve. The recalibrated weights W are applied to each validity dimension in the original validity response curve, performing a weighted reconstruction operation. The reconstruction function is defined as E_opt(t) = Σ[w_i(t) × E_i(t)], where E_i(t) is the contribution rate of different validity dimensions at time t. This function calculates the dynamic balance of validity components over time, reducing high-frequency oscillations and structural skewness in the curve. To further optimize curve continuity, an adaptive smoothing algorithm (such as exponential smoothing or Savitzky-Golay filtering) is used for post-processing of the reconstructed curve, making the validity change trend more consistent with the natural laws of psychological state evolution. Subsequently, the stability improvement rate and the decrease in fluctuation amplitude of the curve before and after optimization are calculated to verify the validity correction effect. When the average variance of the optimized curve is less than 80% of that of the original curve, stability optimization is considered achieved. The final output validity stability optimization curve reflects the dynamic trend of validity under the actual psychological state of the subjects, and can be used as a validity reliability indicator of psychological assessment, providing a quantifiable basis for dynamic verification and validity maintenance.
[0052] In this embodiment, the specific steps for identifying and eliminating dishonest responses from the dynamic validity response curve and extracting the standardized validity response curve are as follows: During the evaluation process, the evaluation subjects undergo continuous facial image scanning to extract facial monitoring images; The background light intensity of the face monitoring image is calculated, and the image brightness is harmonized to obtain a brightness-optimized monitoring image. Dynamic eye movement tracking is performed based on brightness-optimized monitoring images to extract eye movement trajectories; Based on the eye movement trajectory, the trajectory distribution and movement distribution degree are obtained to obtain the gaze point thermal distribution and saccade path complexity. The pupil diameter change rate and periodic blink frequency were calculated based on the brightness-optimized monitoring images. The evaluation concentration is calculated based on the pupil diameter change rate, periodic blink frequency, fixation point thermal distribution, and saccade path complexity, and the authenticity of the answers is analyzed to generate a dynamic profile of the answer integrity. Based on the dynamic profile of the integrity of the responses, an integrity threshold is determined, and dishonest respondents are marked. Based on the identification and elimination of dishonest responses from respondents, the standardized validity response curve is extracted.
[0053] In this embodiment, a high frame rate camera unit (frame rate set to the range of 30fps to 60fps) is embedded in the evaluation terminal device, and a real-time image acquisition module runs in the background. The module adopts a multi-threaded asynchronous capture mechanism to ensure that image acquisition and evaluation interface operation do not interfere with each other. Face detection algorithms (such as face localization models based on MTCNN or Haar features) are used to quickly detect and crop each frame of video image, retaining only image data containing complete facial regions. Each monitoring image is accompanied by timestamp information for subsequent time-series synchronization with the answer behavior data. To ensure the continuity and reliability of monitoring, a dynamic frame loss detection mechanism is set up. When consecutive frame detection failures exceed a set threshold (such as 3 frames), the re-identification and tracking process is automatically started. At the same time, the image sequence is lightweight compressed and noise suppressed to reduce storage pressure and improve subsequent analysis efficiency. Finally, a set of face images labeled with time series is formed, providing high-quality visual data input for subsequent brightness optimization, eye-tracking analysis, and answer integrity recognition. The average background light intensity is calculated for each monitored image. A region segmentation algorithm is used to separate the face region from the background region. The ambient light level is determined by calculating the mean and variance of the grayscale values of the background pixels. When the light intensity deviates from the standard brightness value (set to 120 grayscale units) by more than ±30, the brightness harmonization module is activated. This module uses an adaptive histogram equalization algorithm (CLAHE) to locally enhance or suppress image brightness, preventing loss of detail caused by overall brightness adjustment. Simultaneously, gamma correction (γ value between 0.8 and 1.2) is combined to achieve global brightness balance and skin tone restoration. To maintain the continuity of illumination between time-series images, an inter-frame brightness smoothing algorithm is introduced to dynamically compensate for the brightness difference between adjacent frames, ensuring smooth brightness changes in the video sequence. The processed face images still maintain clear eye areas and facial contours under different lighting conditions, providing stable and reliable input image data for subsequent eye movement tracking and pupil parameter extraction.
[0054] The positions of both eyes are determined using an eye region localization algorithm, and the orbital region of interest (ROI) is extracted in each frame. Subsequently, a pupil center localization algorithm based on Hough circle transform and edge detection is used to calculate the geometric center coordinates of the pupil, generating a time-series of coordinate points. To improve tracking accuracy, optical flow is employed for inter-frame pupil position prediction and dynamic correction, enabling continuous tracking of minute eye movements. Each time point's eye position is accompanied by a time label to ensure the temporal integrity of the trajectory. A moving average filtering algorithm is used to smooth the original trajectory data, eliminating noise interference caused by blinking or changes in lighting. When eye occlusion is detected in consecutive frames or the localization confidence score falls below a threshold (e.g., 0.85), trajectory updates are automatically paused, and trajectory compensation is performed after visibility conditions recover. The final generated eye movement trajectory includes a gaze point sequence, gaze duration, and saccade speed information, providing fundamental data for gaze distribution and visual attention feature calculation. The screen answering interface is divided into grid regions (e.g., 20×20 cells), and the number of gazes and gaze duration in each cell are calculated to form a gaze heatmap matrix. After normalization, the matrix is mapped to a fixation heatmap, reflecting the individual's visual concentration area and attentional diffusion range during the response process. Subsequently, the spatial distribution and path complexity of the eye trajectory are calculated. The distribution is obtained by calculating the variance and mean square distance of the fixation point coordinates to quantify the dispersion of fixation; the saccade path complexity is calculated based on the ratio of the total path length of the trajectory polygon to the straight-line distance, with a higher ratio indicating more complex visual search behavior. Furthermore, the average saccade speed and fixation pause ratio are calculated in conjunction with the time dimension to comprehensively characterize the level of visual focus and cognitive engagement. The final generated fixation heatmap and saccade complexity parameters will be used in conjunction with pupil change characteristics to assess concentration and analyze the authenticity of responses.
[0055] In the brightness-optimized image, the gray-level gradient method was used to identify the pupil edge, and the pupil diameter was estimated using an ellipse fitting algorithm. To reduce the impact of illumination changes on the measurement, a brightness normalization coefficient was used to correct the diameter value. The pupil diameter change rate was calculated by dividing the diameter difference between adjacent time frames by the initial value, and continuous calculations formed a pupil change rate sequence to reflect the instantaneous fluctuations in psychological load. When the change rate exceeded a set threshold (e.g., ±5%), it was recorded as a stress response event. Blink frequency was obtained by detecting changes in the eye closure ratio between frames. The convolutional feature matching method was used to determine the eye opening and closing state, and the number of blinks per minute was counted, and its periodic stability was calculated. Both the pupil change rate and blink frequency were processed using moving averages to remove noise. These two types of indicators together constitute a physiological response feature set, which is used for subsequent assessment of concentration and the authenticity of responses, helping to determine whether the subject is in a state of distraction or emotional interference. The pupil change rate, blink frequency, fixation distribution concentration, and saccade complexity were normalized, mapping each feature to the same numerical range (0 to 1). Subsequently, the overall assessment concentration C(t) was calculated using a weighted fusion model, with weights allocated based on feature stability (e.g., pupil change rate 0.4, blink frequency 0.2, fixation distribution 0.3, saccade complexity 0.1). Higher concentration indicates more sufficient visual and cognitive engagement. Further correlation analysis was performed between the concentration time series and assessment response behavior data, using multiple regression to determine the consistency between concentration and response. When the concentration consistently fell below 0.7 times the overall average, it was marked as a low-focus interval. Combining behavioral characteristics such as abnormal response speed and repeated option modification, the answer authenticity index R(t) was calculated, and a dynamic profile of answer integrity was created. A visual curve of answer integrity change was formed with time on the horizontal axis and R(t) on the vertical axis, providing a basis for subsequent non-honesty identification.
[0056] The algorithm calculates the mean and standard deviation of the dynamic profile of an individual's honesty throughout the test, and sets a threshold range for honesty judgment (e.g., below 0.8 times the mean is considered suspicious). When an individual's honesty curve falls below this threshold for multiple consecutive time periods, it is recorded as a state of non-honesty risk. To avoid misjudgment, the algorithm also detects the number of abnormal pupil changes, sudden increases in saccade path complexity, and abnormal fluctuations in reaction time. When at least two of the three indicators overlap with the low honesty range, the individual is judged as a non-honest respondent. While marking non-honest individuals, the corresponding time period information and behavioral characteristics are stored for subsequent validity elimination and correction. This judgment mechanism can automatically distinguish between individuals who answer with normal focus and those who answer without focus, randomly, or mechanically, providing reliable data support for the validity control of psychological assessments. The marked non-honest individuals or time periods are time-stamped and aligned with the dynamic validity response curve; data in the corresponding intervals are marked as abnormal segments. An abnormal segment elimination strategy is adopted, setting the validity value of the abnormal segment as a missing value, and using time series interpolation methods (such as cubic spline interpolation) to smoothly reconstruct the missing areas. To ensure the reconstructed curve aligns with the overall trend, the local slope difference between the curves before and after removal is calculated, and the interpolation results are corrected to ensure continuity at the boundaries and smooth derivatives. Finally, the reconstructed curve is standardized, mapping the validity value to the range [0,1] to eliminate scale differences between individuals. This standardized validity response curve serves as the core input for subsequent dynamic validity analysis and stability optimization, ensuring that the input data of the validity model reflects only genuine and honest response behavior characteristics, thereby improving the accuracy and reliability of the dynamic validity verification algorithm.
[0057] In this embodiment, the specific steps of step S5 are as follows: Multivariate analysis of variance was conducted based on the validity stability optimization curve to extract the main effect and interaction effect of different groups; Based on the subject-matter effect and interaction effect, a validity response heterogeneity analysis was performed to generate a validity difference matrix. Principal component dimensionality reduction and cluster identification were performed on the validity difference matrix to extract the core common validity of the group; Based on the core of common validity of the group, equivalent transformation and normalization calibration are performed to generate equivalent dynamic calibration.
[0058] In this embodiment, the validity stability optimization curves of all individuals are time-series aligned and normalized to ensure consistency in time scales across different groups. Then, the samples are grouped according to group categorical variables (such as gender, education level, occupation type, or assessment background). A multivariate ANOVA model is constructed, using the stability values of the validity curves at multiple time points as the multivariate dependent variable, and group categories and their interaction terms as independent variables. By calculating the Wilks' Lambda statistic, Pillai's Trace value, and F-statistic corresponding to each independent variable, group factors with significant main effects or interaction effects over time are automatically identified. When the main effect is significant (significance level p < 0.05), it indicates a difference in the overall validity curves between groups; when the interaction effect is significant, it indicates that the validity of different groups changes differently over time. Subsequently, the time periods and validity dimensions corresponding to significant effects are extracted to form a group main effect and interaction effect matrix, recording the direction, intensity, and temporal sequence of differences between groups in matrix form. This matrix provides structured input for subsequent validity heterogeneity analysis. The main effects matrix and interaction effects matrix were standardized to transform the effect strength values at different time periods and validity dimensions into comparable numerical ranges (e.g., between 0 and 1). Then, based on the time-series characteristics of the validity stability optimization curve, the validity difference vectors between groups at key time points were calculated. The difference values were obtained by calculating the ratio of the difference between the group means to the within-group variance to characterize the relative dispersion of the validity response. To reveal the heterogeneity in the structure of validity responses between groups, the dynamic correlation coefficient matrix between the validity change curves of each group was further calculated and element-wise weighted and superimposed with the main effects matrix to generate the final validity difference matrix. This matrix uses the group as the row and column index, and the unit value represents the intensity of the difference in dynamic validity changes between groups. Higher values indicate greater inconsistency in validity characteristics between groups. Threshold segmentation was also used to identify regions in the difference matrix that significantly deviate from the overall mean, marked as "high heterogeneity regions." This matrix result provides a multi-dimensional input data structure for subsequent principal component analysis and group cluster identification, used to mine common validity characteristics between groups.
[0059] Principal component analysis (PCA) was performed on the validity variance matrix to extract the most representative feature dimensions. By calculating the covariance matrix and extracting eigenvalues and eigenvectors, principal components that explain more than 85% of the total variance were retained, thus preserving key validity variance information while reducing dimensionality. The dimensionality reduction result forms a principal component feature space, where the coordinates of each group represent its validity response feature distribution. Subsequently, K-means clustering or hierarchical clustering was used to classify the principal component feature vectors, grouping groups with similar validity patterns into the same class. Euclidean distance was used as the similarity metric during clustering, and the principle of minimizing the within-group sum of squares was adopted to achieve optimal partitioning. Finally, the common validity core of the groups was extracted by calculating the principal component mean vector of each group. This validity core reflects the common stable characteristics and key validity structures of different groups in the psychological testing process, and is an important parameter describing the consistency of psychological testing at the group level, providing a standard reference for subsequent equivalence dynamic calibration. The validity stability optimization curve of each group was aligned and compared with the common validity core of the groups, and the deviation vector at each time point was calculated. To eliminate the gender bias caused by differences in assessment characteristics among different groups, an equivalent transformation function is introduced to perform linear and nonlinear dual mapping correction on the validity curves of each group. The linear transformation uses least squares regression to shift the group curves to the central trend line of the common core; the nonlinear transformation uses polynomial fitting or sigmoid function transformation to dynamically compensate for nonlinear bias. Subsequently, the corrected curves are normalized to map the validity responses of each group to a unified interval (e.g., [0,1]), ensuring comparability between different groups on the same scale. To verify the calibration effect, the ratio of inter-group variance and the validity similarity index before and after equivalent calibration are calculated. Calibration is considered effective when the inter-group variance decrease rate exceeds 20%. The final equivalent dynamic calibration model maintains the consistency of validity curves across multiple groups and time dimensions, providing a dynamic, cross-group unified validity benchmark for psychological assessment and achieving adaptive group-based optimization of the dynamic validity verification algorithm.
[0060] In this embodiment, step S6 includes the following steps: Based on equivalent dynamic calibration, intelligent optimization of the weights and presentation order of assessment items is carried out to build a dynamic scheduling engine for assessment. Dynamic validity self-verification and real-time reliability self-optimization are performed based on the dynamic scheduling engine of the assessment.
[0061] In this embodiment, the output data of the equivalent dynamic calibration model is used to extract the validity sensitivity intervals and response difference characteristics of each group at different time periods. Based on these difference characteristics, an item weight adjustment matrix is constructed, using the contribution rate of each item in different validity dimensions (content validity, construct validity, and response validity) as the basis for weighting. By calculating the matching degree index between an individual's current validity status and item validity sensitivity, the priority order of items is determined. When an individual's validity stability is in a low range, the weight of items related to emotion regulation or attention dimensions is automatically increased to enhance the relevance and effectiveness of the assessment. A dynamic weighted strategy update mechanism in reinforcement learning algorithms is adopted to adjust the weight parameters in real time based on individual response behavior feedback and changes in validity stability, maintaining the adaptability of the scheduling process. At the same time, a sequence optimization model is introduced at the item presentation end, using Markov chain transition probabilities to predict and sort the order of items, ensuring that high-weight items appear first and avoiding continuous accumulation of cognitive load. Finally, a dynamic scheduling engine for the assessment is formed, realizing multi-dimensional real-time control of item weights and order, providing basic support for the accurate and intelligent execution of the assessment. The assessment process implements dynamic validity self-verification and real-time credibility optimization at the implementation level, enabling the assessment to possess self-awareness, self-correction, and self-adaptive capabilities. During the assessment, it receives input information from the dynamic scheduling engine in real time, including changes in item weights, adjustments to presentation order, and the flow of response behavior characteristics. By comparing the output with the equivalent dynamic calibration model in real time, it calculates the validity shift rate and credibility index of the current assessment state. When the validity shift rate exceeds a set threshold (e.g., 0.1) or the credibility index falls below a standard value (e.g., 0.8), a self-verification procedure is immediately initiated to make local adjustments to the assessment path. The self-verification process includes two parts: first, dynamic validity feedback correction, which recalculates the validity distribution of the current stage and corrects the item presentation strategy through multivariate analysis of variance; second, a credibility self-optimization mechanism, which improves the authenticity of assessment behavior data by adding auxiliary highly diagnostic items or inserting attention-checking items. After each item is completed, the validity assessment parameters are updated in real time, and the Kalman filter algorithm is used to smooth the validity change curve over continuous time periods, ensuring the stability of the response. Ultimately, a closed-loop dynamic validity self-verification process is formed throughout the entire assessment cycle. Through continuous self-optimization, the psychological assessment can automatically maintain a high level of credibility and efficiency during operation, thereby realizing the intelligent closed-loop operation of the dynamic validity verification algorithm for psychological assessment based on multivariate analysis of variance.
[0062] In this embodiment, a dynamic validity verification system for psychological tests based on multivariate analysis of variance is provided, used to execute the dynamic validity verification algorithm for psychological tests based on multivariate analysis of variance as described above, including: The response status assessment module is used to collect assessment record information of the assessment subjects, assess the stability of their responses, and generate an assessment status stability curve. The scoring difference module is used to extract group assessment reports, perform multi-dimensional question score calculation and score difference analysis, and generate assessment score difference values. The validity response module is used to perform validity contribution response analysis based on the assessment stability curve and the difference in assessment scores, and to construct a dynamic validity response curve. The recalibration module is used to analyze the validity degradation interval of the dynamic validity response curve, perform dynamic recalibration, and extract the validity stability optimization curve. The normalization calibration module is used to perform multivariate ANOVA and normalization calibration based on the validity stability optimization curve, and generate equivalent dynamic calibration. The validity self-verification module is used for dynamic validity self-verification and real-time reliability self-optimization based on equivalent dynamic calibration.
[0063] Therefore, the embodiments should be considered as exemplary and non-limiting in all respects, and the scope of the invention is defined by the appended claims rather than the foregoing description. Thus, all variations falling within the meaning and scope of the equivalents of the application are intended to be included within the invention.
[0064] The above description is merely a specific embodiment of the present invention, enabling those skilled in the art to understand or implement it. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein are implemented in other embodiments without departing from the spirit or scope of the invention. Therefore, the present invention is not to be limited to the embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novel features of the invention herein.
Claims
1. A dynamic validity verification algorithm for psychological tests based on multivariate analysis of variance, characterized in that, Includes the following steps: Step S1: Collect the assessment record information of the assessment subjects, conduct a response stability assessment, and generate an assessment state stability curve; Step S2: Extract the group assessment report, perform multi-dimensional question score calculation and score difference analysis, and generate assessment score difference value; Step S3: Based on the stability curve of the assessment state and the difference in assessment scores, conduct a validity contribution response analysis and construct a dynamic validity response curve; Step S4: Perform validity degradation interval analysis on the dynamic validity response curve, and perform dynamic recalibration to extract the validity stability optimization curve; Step S5: Perform multivariate ANOVA and normalization calibration based on the validity stability optimization curve to generate equivalent dynamic calibration; Step S6: Perform dynamic validity self-verification and real-time reliability self-optimization based on equivalent dynamic calibration.
2. The dynamic validity verification algorithm for psychological tests based on multivariate analysis of variance as described in claim 1, characterized in that, The specific steps of step S1 are as follows: Collect assessment record information from the assessment subjects; Based on the assessment record information, the response time, option modification trajectory, and question dwell time distribution are calculated to obtain multi-dimensional assessment response data; The multi-dimensional evaluation response data is processed by time-series segmentation to extract evaluation response data from multiple time periods; Based on the evaluation response data, the answer behavior characteristics are analyzed for each time period to generate the answer behavior pattern for each time period; Behavioral fluctuation analysis is performed on the aforementioned response behavior pattern, and response stability is assessed to generate an evaluation state stability curve.
3. The dynamic validity verification algorithm for psychological tests based on multivariate analysis of variance as described in claim 1, characterized in that, The specific steps of step S2 are as follows: Collect group evaluation reports; Based on the group assessment report, the assessment subjects are identified, and multi-dimensional question scores are calculated to extract the assessment scores of different subjects in multiple dimensions; the multi-dimensional assessment scores include the original score matrix, reverse question score transformation records, and missing value imputation trajectory; Based on the group assessment report, the question difficulty parameter, discrimination parameter, and guessing coefficient are calculated to obtain the question complexity. Based on the complexity of the assessment questions, the assessment scores of the multiple dimensions are grouped and the score difference is calculated to generate an assessment score difference value.
4. The dynamic validity verification algorithm for psychological tests based on multivariate analysis of variance as described in claim 1, characterized in that, Step S3 is as follows: Extract the identity information of the evaluation subjects; match the evaluation state stability curves based on the evaluation subject identity information, and extract the state stability curve for each subject; Based on the state stability curve and the difference in evaluation scores, feature alignment is performed to generate a state temperature-score difference collaborative matrix. Define the dependent variable matrix, including item factors, state factors, and time factors; Based on the dependent variable matrix, the state temperature-score difference co-factor matrix is decomposed into factor interaction terms to obtain different validity contribution rates, including content validity, construct validity, and response validity. Dynamic validity response curves are constructed based on different validity contribution rates.
5. The dynamic validity verification algorithm for psychological tests based on multivariate analysis of variance as described in claim 1, characterized in that, The specific steps of step S4 are as follows: The dynamic validity response curve is used to identify and eliminate dishonest responses, and the standardized validity response curve is extracted. The time drift rate is obtained by calculating the slope of the local temporal variation of the standardized validity response curve. The variance imbalance is obtained by calculating the structural skewness of the contribution rates of each validity function based on the dynamic validity response curve. Regional validity degradation interval analysis was conducted based on time drift rate and variance imbalance, and potential degradation intervals were marked. Dynamically recalibrate the potential degradation interval and extract the core parameters to recalibrate the weights. The validity stability of the dynamic validity response curve is optimized based on the recalibration weights of the core parameters, and the validity stability optimization curve is extracted.
6. The dynamic validity verification algorithm for psychological tests based on multivariate analysis of variance as described in claim 5, characterized in that, The specific steps for identifying and eliminating dishonest responses from the dynamic validity response curve and extracting the standardized validity response curve are as follows: During the evaluation process, the evaluation subjects undergo continuous facial image scanning to extract facial monitoring images; The background light intensity of the face monitoring image is calculated, and the image brightness is harmonized to obtain a brightness-optimized monitoring image. Dynamic eye movement tracking is performed based on brightness-optimized monitoring images to extract eye movement trajectories; Based on the eye movement trajectory, the trajectory distribution and movement distribution degree are obtained to obtain the gaze point thermal distribution and saccade path complexity. The pupil diameter change rate and periodic blink frequency were calculated based on the brightness-optimized monitoring images. The evaluation concentration is calculated based on the pupil diameter change rate, periodic blink frequency, fixation point thermal distribution, and saccade path complexity, and the authenticity of the answers is analyzed to generate a dynamic profile of the answer integrity. Based on the dynamic profile of the integrity of the responses, an integrity threshold is determined, and dishonest respondents are marked. Based on the identification and elimination of dishonest responses from respondents, the standardized validity response curve is extracted.
7. The dynamic validity verification algorithm for psychological tests based on multivariate analysis of variance as described in claim 1, characterized in that, The specific steps of step S5 are as follows: Multivariate analysis of variance was conducted based on the validity stability optimization curve to extract the main effect and interaction effect of different groups; Based on the subject-matter effect and interaction effect, a validity response heterogeneity analysis was performed to generate a validity difference matrix. Principal component dimensionality reduction and cluster identification were performed on the validity difference matrix to extract the core common validity of the group; Based on the core of common validity of the group, equivalent transformation and normalization calibration are performed to generate equivalent dynamic calibration.
8. The dynamic validity verification algorithm for psychological tests based on multivariate analysis of variance as described in claim 1, characterized in that, The specific steps of step S6 are as follows: Based on equivalent dynamic calibration, intelligent optimization of the weights and presentation order of assessment items is carried out to build a dynamic scheduling engine for assessment. Dynamic validity self-verification and real-time reliability self-optimization are performed based on the dynamic scheduling engine of the assessment.
9. A dynamic validity verification system for psychological assessments based on multivariate analysis of variance, characterized in that, The algorithm for performing dynamic validity verification of psychological tests based on multivariate analysis of variance as described in claim 1 includes: The response status assessment module is used to collect assessment record information of the assessment subjects, assess the stability of their responses, and generate an assessment status stability curve. The scoring difference module is used to extract group assessment reports, perform multi-dimensional question score calculation and score difference analysis, and generate assessment score difference values. The validity response module is used to perform validity contribution response analysis based on the assessment stability curve and the difference in assessment scores, and to construct a dynamic validity response curve. The recalibration module is used to analyze the validity degradation interval of the dynamic validity response curve, perform dynamic recalibration, and extract the validity stability optimization curve. The normalization calibration module is used to perform multivariate ANOVA and normalization calibration based on the validity stability optimization curve, and generate equivalent dynamic calibration. The validity self-verification module is used for dynamic validity self-verification and real-time reliability self-optimization based on equivalent dynamic calibration.