Intelligent evaluation methods, systems, and readable storage media based on observable indicators

By constructing observable indicator units and consistency rules with psychological theories, and combining AI large models and retrieval-enhanced knowledge bases, the problems of subjective dependence and theoretical disconnect in existing psychological assessment technologies are solved, and objective, interpretable and stable psychological assessments are achieved.

CN122074989APending Publication Date: 2026-05-26浙江连信科技有限公司
View PDF 4 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
浙江连信科技有限公司
Filing Date
2026-04-21
Publication Date
2026-05-26

AI Technical Summary

Technical Problem

Existing psychological assessment techniques suffer from problems such as reliance on the subjective cooperation of subjects, lack of psychological theoretical support for assessment results, insufficient model generalization ability, and poor system scalability. They are difficult to adapt to high-frequency, real-time, or large-scale screening scenarios, and the assessment results are difficult to verify.

Method used

By constructing observable indicator units and consistency rules with psychological theories, and combining them with AI large models and retrieval-enhanced knowledge bases, the mapping and reasoning between behavioral characteristics and psychological constructs are realized, generating traceable psychological assessment results.

Benefits of technology

It achieves objective behavioral analysis, eliminates social desirability bias, improves the interpretability and professional credibility of the evaluation results, enhances the comprehensiveness and stability of the evaluation, and strengthens the generalization ability and test-retest consistency of the evaluation system.

✦ Generated by Eureka AI based on patent content.
Patent Text Reader

Abstract

This invention discloses an intelligent assessment method, system, and readable storage medium based on observable indicators, belonging to the field of information data processing technology. The method includes: acquiring the definition of a target psychological construct from a standardized psychological scale or psychological theoretical system; converting it into one or more observable indicator units according to observability rules and consistency rules with psychological theories; constructing an enhanced knowledge base for observable indicator retrieval; structurally binding the observable indicator units with evidence from psychological literature; acquiring multimodal behavioral data and extracting features; aligning the data with the observable indicator units to generate a set of behavioral evidence; and performing inference assessment based on a collaborative mechanism of rules, knowledge base, and machine learning models. This invention overcomes the limitations of existing technologies that rely on subjective self-assessment or black-box model prediction by constructing a computational representation layer from psychological constructs to observable indicators, achieving objective psychological assessment that requires no subject self-assessment, has a theoretical basis, and is traceable.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of information and data processing technology, and specifically relates to an intelligent evaluation method, system, and readable storage medium based on observable indicators. Background Technology

[0002] With the continuous development of artificial intelligence, multimodal signal processing, and psychological research, technologies for automating the assessment of individual psychological states, personality traits, or psychological risks using computer systems have been gradually applied in fields such as mental health screening, educational assessment, talent selection, judicial auxiliary analysis, and public safety. Currently, psychological assessment technologies mainly fall into three categories: first, traditional assessment methods based on standardized psychological scales, where subjects actively complete questionnaires or self-assessment scales to obtain psychological trait scores; second, prediction methods based on machine learning or deep learning, which utilize single or multimodal data such as text, voice, or video to directly output predicted results of psychological states or personality traits through data-driven models; and third, analytical methods based on rules or expert systems, which analyze and infer specific behavioral patterns through manually formulated expert experience rules.

[0003] However, the aforementioned existing technical solutions all have significant limitations in practical applications. Assessment methods based on psychological scales heavily rely on the subject's subjective cooperation and self-awareness, making them susceptible to social desirability bias, self-perception bias, and the subject's emotional state. The assessment process is time-consuming, making it difficult to adapt to high-frequency, real-time, or large-scale screening scenarios, and even more unsuitable for non-cooperative or passive assessment scenarios. While prediction methods based on artificial intelligence models automate the assessment process, they are essentially black-box mappings based on statistical correlation, lacking clear psychological theoretical support. The reasoning process is uninterpretable, and the results cannot be traced back to verifiable theoretical basis. The model's generalization ability and test-retest consistency heavily depend on the distribution characteristics of the training data. Although rule-based or expert system-based methods offer some interpretability, rule construction is costly and maintenance is difficult. They also struggle to effectively handle complex multimodal, high-dimensional behavioral data, resulting in poor system scalability and environmental adaptability. More fundamentally, existing technologies generally suffer from a disconnect between psychological theories and computational processes. Psychological theories are used only as reference information for post-hoc interpretation of results, rather than as constraints in the computational process. There is a lack of a systematic intermediate representation layer between behavioral characteristics and abstract psychological constructs, making it difficult to verify evaluation results and preventing the formation of effective evidence to support the reasoning process. Summary of the Invention

[0004] To address the problems existing in the prior art, this invention provides an intelligent evaluation method, system, and readable storage medium based on observable indicators. By transforming abstract psychological concepts into concrete behavioral indicators and combining them with psychological evidence for evaluation, the accuracy is improved.

[0005] The technical solution adopted in this invention is as follows:

[0006] In a first aspect, the present invention provides an intelligent evaluation method based on observable indicators, comprising the following steps:

[0007] First, target psychological constructs derived from standardized psychological scales and / or psychological theory systems are selected. Then, through observability rules and consistency rules of psychological theories, the target psychological constructs are transformed into observable indicator units in compliance with regulations. Finally, a retrieval-enhanced knowledge base is constructed that structurally binds observable indicator units and target psychological constructs and supports the mapping relationship between the two.

[0008] Based on the non-self-evaluation multimodal behavioral data of the object to be evaluated, the AI ​​big model is used to match behavioral features with observable indicator units through the theoretical boundary constraints of the retrieval enhancement knowledge base. Combined with the preset collaborative reasoning mechanism, the psychological construct assessment is completed, and a traceable psychological assessment result with a complete reasoning path and theoretical basis is generated.

[0009] In conjunction with the first aspect, the present invention provides a first implementation of the first aspect, wherein the mental theory consistency rule is a rigid mapping constraint rule that runs through the entire process of construct transformation, knowledge base construction, and reasoning evaluation. The specific implementation is as follows:

[0010] In the construct transformation stage, the mapping relationship between observable indicator units and target psychological constructs is verified. The exclusive definition boundaries and causal relationship requirements of the target psychological constructs in their respective psychological theories are strictly matched, and behavioral representations that have no theoretical causal relationship with the target psychological constructs and only have statistical correlation are eliminated.

[0011] In the knowledge base construction phase, the theoretical definition boundaries of the target mental construct, the mutual exclusion relationship with other mental constructs, and the sub-dimension attribution information are structurally bound and stored with observable indicator units to form a reasoning constraint ontology.

[0012] During the reasoning and evaluation phase, the feature matching results and intermediate reasoning conclusions output by the AI ​​large model are verified in real time, and association matching and reasoning conclusions that exceed the boundaries defined by the target mental construct theory are filtered out.

[0013] In conjunction with the first aspect, the present invention provides a second implementation of the first aspect, wherein the observability rule is an admission verification rule for observable index units, specifically as follows:

[0014] Verify whether the candidate observable indicator units correspond to explicit behavioral representations that can be directly captured by audio and video acquisition devices and text acquisition interfaces. Eliminate introspective descriptions that can only be obtained through subjective self-evaluation by the object to be evaluated. The observable indicator units are unambiguously and repeatedly collected and verified by third-party devices.

[0015] In conjunction with the first aspect or the second implementation of the first aspect, the present invention provides a third implementation of the first aspect, wherein the retrieval enhancement knowledge base adopts a hybrid architecture of vector database and graph database, and stores structured data in RDF triple format;

[0016] The psychological literature evidence includes research DOI, sample size, construct-indicator association effect size, confidence interval, and theoretical attribution metadata.

[0017] In conjunction with the third embodiment of the first aspect, the present invention provides a fourth embodiment of the first aspect, which specifically includes the following steps:

[0018] S1. Obtain the standardized definition text of the target psychological construct, wherein the target psychological construct is derived from a standardized psychological scale or a mature psychological theory system that has been validated by peers.

[0019] S2. The admission verification of candidate observable indicator units is completed through observability rules, and the mapping compliance verification between candidate observable indicator units and target mental constructs is completed through psychological theory consistency rules. The observable indicator units that pass the dual verification are established with the target mental constructs.

[0020] S3. Construct a retrieval-enhanced knowledge base, which structurally binds and stores observable indicator units, theoretical boundary constraints on the ontology of target psychological constructs, and psychological literature evidence supporting the mapping relationship between the two.

[0021] S4. Obtain non-self-evaluation multimodal behavioral data of the object to be evaluated, extract corresponding multi-dimensional behavioral features, and based on the theoretical boundary constraints in the retrieval enhancement knowledge base, complete the matching mapping between behavioral features and observable indicator units to generate a set of structured behavioral evidence.

[0022] S5, based on the collaborative reasoning mechanism of preset expert rules, retrieval-enhanced knowledge base, and AI big model, quantifies and evaluates the behavioral evidence set, generating traceable psychological assessment results.

[0023] In conjunction with the fourth implementation of the first aspect, the present invention provides a fifth implementation of the first aspect, wherein the dual verification further includes document support constraints and multimodal identifiability constraints;

[0024] The literature support constraints are: verify the mapping relationship between candidate observable indicator units and target psychological constructs, have published empirical psychological research data to support them, and retain only candidate indicator units whose associated effect size reaches a preset threshold.

[0025] The multimodal identifiability constraint is that the candidate observable indicator unit can be identified and verified by behavioral data from at least two of the three modalities: text, audio, and video.

[0026] In conjunction with the fourth implementation of the first aspect, the present invention provides a sixth implementation of the first aspect. In the collaborative reasoning mechanism of step S5, the preset expert rules are used to limit the boundary conditions and mutual exclusion rules of the reasoning logic, the retrieval enhancement knowledge base is used to provide theoretical boundary constraints and empirical support for the entire reasoning process, and the AI ​​big model is only used for pattern recognition of behavioral features and preliminary matching of features and observable indicator units, and does not directly output the final evaluation conclusion of the mental construct.

[0027] In a second aspect, the present invention provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements an intelligent evaluation method based on observable indicators as described in any of the preceding claims.

[0028] Thirdly, the present invention provides a psychological assessment system that is applied to the intelligent assessment method based on observable indicators described in any of the above claims to acquire behavioral data of the subject to be assessed, extract behavioral features, and perform reasoning analysis on psychological constructs to generate psychological assessment results.

[0029] The beneficial effects of this invention are as follows:

[0030] (1) By constructing a conversion mechanism from psychological constructs to observable indicator units, this invention transforms abstract psychological concepts into intermediate representations that can be recognized and verified by computer systems, realizing the transformation of psychological assessment from subjective self-evaluation to objective behavioral analysis, and effectively eliminating the influence of social desirability bias and self-cognition bias on the assessment results.

[0031] (2) By introducing a retrieval-enhanced knowledge base into the psychological assessment process, the present invention structurally binds observable indicator units with psychological literature evidence, so that the model reasoning process is constrained by clear theoretical evidence, ensuring that the assessment conclusions can be traced back to specific psychological theoretical basis, and significantly improving the interpretability and professional credibility of the assessment results.

[0032] (3) By establishing an alignment mechanism between multimodal behavioral features and observable indicator units, this invention realizes the fusion analysis of multi-source heterogeneous data such as text, audio, and video under a unified semantic framework, avoids the information limitations of single-modal data, and improves the comprehensiveness and stability of psychological assessment.

[0033] (4) By adopting a reasoning framework that combines rules, knowledge and models, this invention overcomes the uncontrollability of pure data-driven model reasoning, so that the evaluation process can maintain the high efficiency of automated processing and meet the theoretical consistency requirements of professional psychological evaluation, thereby enhancing the generalization ability and retest consistency of the evaluation system in different application scenarios. Detailed Implementation

[0034] The present invention will be further illustrated below with reference to specific embodiments. It should be understood that the following embodiments are only used to illustrate the technical concept of the present invention, and are not an exhaustive list.

[0035] Example 1:

[0036] As one implementation method, this embodiment discloses an intelligent assessment method based on observable indicators, applicable to passive psychological characteristic analysis scenarios where subjects do not need to actively participate in self-assessment. This method transforms abstract psychological concepts into a calculable, verifiable, and traceable objective assessment system by constructing a three-layer computational representation architecture of psychological constructs, observable indicators, and multimodal behavioral evidence.

[0037] It should be understood that the "observable indicator unit" described in this invention is not a simple behavioral label, but an intermediate representation layer between abstract mental constructs and raw behavioral data.

[0038] From a psychometric perspective, psychological constructs (such as neuroticism and social anxiety) are latent variables that cannot be directly observed. Traditional self-rating scales map latent variables to explicit indicators through the introspection of the subjects, but are limited by social desirability bias and cognitive blind spots.

[0039] This invention introduces observability rules and consistency rules with psychological theories to construct computer-recognizable operational definitions of latent variables, enabling AI systems to complete compliant reasoning from multimodal behavioral data to latent variable evaluation under clear theoretical boundary constraints.

[0040] In one embodiment, the method includes the following core modules, the functional configuration of which corresponds to the technical features of the present invention:

[0041] Module 1: Compliance Conversion of Psychological Constructs. This module is configured to acquire target psychological constructs derived from standardized psychological scales or mature psychological theoretical systems that have been peer-validated, and to complete the compliance conversion of target psychological constructs into observable indicator units through observability rules and psychological theory consistency rules.

[0042] Module 2: Retrieval Enhancement Knowledge Base, configured to store observable indicator units, theoretical boundary constraint ontology of target psychological constructs, and psychological literature evidence supporting the mapping relationship between the two, forming a structured and bound retrieval enhancement knowledge base.

[0043] Module 3: Multimodal Behavior Analysis and AI Matching Module, is configured to acquire non-self-evaluation multimodal behavior data of the object to be evaluated, and use the AI ​​big model to complete the matching of behavioral features with observable indicator units through the theoretical boundary constraints of the retrieval enhancement knowledge base.

[0044] Module 4: Collaborative Reasoning Assessment Module, configured to combine a preset collaborative reasoning mechanism to complete the assessment of mental constructs and generate traceable mental assessment results with complete reasoning paths and theoretical basis.

[0045] The following section details the specific implementation parameters, algorithm logic, and verification data for each module, using a scenario of stress susceptibility assessment in corporate recruitment as an example.

[0046] Step S1: Acquiring and Standardizing the Target Mental Construct

[0047] Obtain the target psychological construct derived from the NEO-PI-R personality scale and the Costa & McCrae five-factor personality theory, namely the standardized definition text of neuroticism.

[0048] The theoretical definition of this construct is clearly defined as follows:

[0049] Individuals exhibit emotional instability tendencies when facing stressful situations, with core sub-dimensions including anxiety, anger and hostility, depression, self-awareness, impulsivity, and vulnerability. At the same time, it is important to clarify the mutually exclusive relationship between this construct and other personality constructs such as extraversion and conscientiousness to ensure that subsequent reasoning does not result in conceptual drift.

[0050] Step S2: Compliance conversion and dual verification of observable indicator units

[0051] By employing both observability rules and consistency rules with mental theory, a compliant transformation from neurotic constructs to observable indicator units is achieved.

[0052] The specific implementation of the observability rule involves conducting an admission check on several initially generated candidate behavioral representations.

[0053] Verify whether the candidate indicators correspond to explicit behavioral representations that can be directly captured by audio / video capture devices and text capture interfaces. Remove introspective descriptions such as "I often feel anxious" that can only be obtained through subjective self-assessment, and retain only a number of candidate indicators that can be unambiguously and repeatedly captured and verified by third-party devices, including: co-occurrence frequency of facial action units, amplitude of speech fundamental frequency micro-vibration, and frequency of head avoidance posture.

[0054] The specific implementation of the theory-of-mind consistency rule involves: verifying the mapping relationship between several candidate indicators and the neuroticism construct, strictly matching the specific definition boundaries and causal relationship requirements of neuroticism in the Big Five personality theory; eliminating behavioral representations that have no theoretical causal relationship with neuroticism and only have statistical correlations; and finally retaining several candidate indicators that pass the theory consistency verification.

[0055] The specific implementation of the literature support constraint is as follows: For the mapping relationship between candidate observable indicator units and target psychological constructs, psychological literature databases are searched, and metadata from published empirical studies is extracted. Only candidate indicator units with peer-reviewed journal publication records, sample size N≥100, association effect size reaching a preset threshold, and confidence intervals not containing zero are retained. For example, when verifying the association between the fundamental frequency variation coefficient of speech and the neurotic construct, the literature evidence includes the study DOI (10.1037 / pas0000287), sample size (N=2145), and association effect size (r=0.42, 95%CI[0.38,0.46]), eliminating candidate indicators based solely on small samples (N<50) or with insufficient effect sizes (r<0.3).

[0056] The specific implementation of the multimodal identifiability constraint is as follows: verify whether the candidate observable indicator unit can be independently identified and verified through at least two of the three modalities: text, audio, and video. For example, the indicator of the proportion of negative facial expression duration must pass the cross-validation of both video and audio modalities to ensure that the single-modal false recognition rate is less than 15% and the dual-modal joint recognition accuracy is not less than 85% before it can pass the admission verification.

[0057] The following core observable units were ultimately identified:

[0058] Percentage of negative facial expressions (video modality, based on FACS coding system).

[0059] Speech fundamental frequency variation coefficient (audio modality, based on PRAAT acoustic analysis);

[0060] Frequency of pauses in stressed scenarios (audio + text modality);

[0061] Frequency of absolute negative expressions in text (text modality, based on LIWC dictionary);

[0062] Frequency of head avoidance posture (video modality, based on 3D head pose estimation algorithm).

[0063] Duration of limb contraction posture (video modality, based on OpenPose skeleton keypoint tracking);

[0064] The magnitude of the decrease in semantic coherence (text modality, calculated based on BERT semantic similarity).

[0065] Establish a specific mapping relationship between the above 7 observable indicator units and neurotic constructs to form construct-indicator mapping pairs.

[0066] Step S3: Construction of Enhanced Knowledge Base

[0067] A retrieval-enhanced knowledge base is constructed using a hybrid architecture combining vector and graph databases. Structured data is stored in RDF triple format, and the specific implementation includes:

[0068] Knowledge graph ontology construction: The core triplet is neurotic construct-theoretical definition boundary-observable indicator unit, which binds and stores 7 observable indicator units and standardized definition text of neurotic construct.

[0069] Construct a reasoning constraint ontology to structurally bind and store the mutual exclusion relationship between neuroticism and other personality constructs, sub-dimension attribution information, and corresponding observable indicator units.

[0070] Structured binding of documentary evidence: This involves binding corresponding psychological documentary evidence to each observable indicator unit, forming a traceable chain of empirical support. Taking the fundamental frequency variation coefficient of speech as an example, the bound documentary metadata includes:

[0071] The study included the DOI (10.1037 / pas0000287), sample size (N=2145), construct-index association effect size (r=0.42, 95%CI[0.38,0.46]), and theoretical attribution. The association structure of neurotic constructs, fundamental frequency variation coefficients of speech, and literature evidence was stored in RDF triple format to ensure that statistical evidence from the literature could be used in real-time as a basis for confidence weighting during subsequent inference.

[0072] Hybrid storage architecture implementation: Milvus vector database is used to store the semantic embedding vectors of observable indicator units, supporting fast retrieval based on semantic similarity; Neo4j graph database is used to store the relationships between constructs, indicators, and documents, supporting graph traversal queries with complex reasoning paths. A unified API interface enables collaborative calling of vector retrieval and graph queries.

[0073] Step S4: Multimodal behavioral data acquisition and AI matching processing

[0074] Non-self-report multimodal behavioral data acquisition: A semi-structured stress interview task was designed to acquire non-self-report multimodal behavioral data of the subjects to be evaluated. Data acquisition was completed passively, without requiring subjects to fill out any self-report scales.

[0075] Video data: Acquired using a binocular camera with a resolution of 1080P and a frame rate of 30fps. The entire process captures the candidate's facial expressions, head posture, and body movements. The face resolution is no less than 300×300 pixels to ensure the accuracy of facial motion unit recognition.

[0076] Audio data: Acquired using an omnidirectional microphone, with a sampling rate of 44.1kHz, 16-bit quantization, and synchronous acquisition of candidate speech data throughout the process, with a signal-to-noise ratio of no less than 60dB, ensuring the integrity of acoustic features such as fundamental frequency micro-vibration;

[0077] Text data: Based on ASR speech-to-text technology, the entire interview conversation is transcribed into text data with a word error rate of less than 3%. The conversation timestamp is preserved simultaneously to ensure time alignment accuracy with the audio and video channels of ±100ms.

[0078] Multi-dimensional behavioral feature extraction and AI matching: Utilizing a large AI model and based on the theoretical boundary constraints of the aforementioned retrieval enhancement knowledge base, the matching and mapping of behavioral features extracted from multimodal behavioral data with observable indicator units is completed.

[0079] Regarding the use of large AI models: This embodiment adopts a technical approach of calling existing pre-trained models and performing domain adaptation, rather than developing new model architectures. Specifically:

[0080] An adaptive pre-trained language model in the field of psychology was selected as the base model. This model is based on the open-source Llama2-7B architecture and fine-tuned on corpora in the field of psychology.

[0081] The fine-tuned dataset includes the publicly available Chinese psycholinguistics dataset CLP and the Chinese depression corpus CCMD, totaling 1.2 million annotated text sentences. The training hyperparameters are a learning rate of 2e-5, a batch size of 16, and 8 training epochs.

[0082] The fine-tuning objective is to enable the model to have the ability to recognize behavioral features and to initially match features with observable indicator units, while prohibiting the training model from directly outputting the final evaluation conclusion of mental constructs.

[0083] During the inference phase, the AI ​​model performs feature recognition and preliminary matching only within the seven observable indicator units limited by the knowledge base. This includes features such as eye contact duration and body posture changes in videos, speech rate and pauses in audio, and self-negating statements in text. No feature associations are made beyond the theoretical boundaries. The entire model training and inference process is constrained by the consistency rule of mental theory. When the model outputs associations that exceed the boundaries defined by the target mental construct theory, the system automatically filters and removes them.

[0084] Step S5: Collaborative Reasoning Mechanism and Quantitative Evaluation

[0085] Based on a collaborative reasoning mechanism utilizing pre-defined expert rules, a retrieval-enhanced knowledge base, and a large AI model, the behavioral evidence set is quantitatively evaluated.

[0086] Collaborative reasoning division of labor mechanism:

[0087] Pre-defined expert rule layer: This layer defines the boundary conditions and mutual exclusion rules for reasoning logic. For example, a rule could be set such that if a candidate exhibits a high frequency of proactive speaking, the neuroticism score weight should be reduced even if nervous behavior occurs, to prevent construct confusion.

[0088] The knowledge base layer, enhanced by retrieval, provides theoretical boundary constraints and empirical support throughout the reasoning process. It retrieves the knowledge base in real-time during reasoning to obtain the construct-indicator correlation effect size and confidence interval corresponding to the current behavioral feature, which serves as the basis for scoring weights.

[0089] AI large model layer: only used for pattern recognition of behavioral features, preliminary matching of features and observable indicator units, outputting intermediate matching results, such as detecting high fundamental frequency variation - matching indicator 2: speech fundamental frequency variation coefficient, and does not directly output the final evaluation conclusion of neurotic constructs.

[0090] Full-process rule verification: Based on the consistency rule of psychological theory, the feature matching results and intermediate inference conclusions output by the AI ​​large model are verified in real time.

[0091] All association matching and inference conclusions that exceed the boundaries defined by the neuroticism construct theory will be filtered out. For example, if the AI ​​model incorrectly matches the extraversion indicator of frequent smiling as a neuroticism indicator, the system will automatically filter out the match based on the mutual exclusion relation ontology in the knowledge base.

[0092] Multi-dimensional quantitative scoring: Based on the knowledge base relevance effect size corresponding to each piece of behavioral evidence in the behavioral evidence set, a reliability weight is assigned to each piece of behavioral evidence. Quantitative scoring is completed according to three fixed dimensions: strength, importance, and relevance.

[0093] Intensity dimension: Calculates the degree of deviation of behavioral characteristics from the norm sample of the adult population;

[0094] Importance dimension: The predictive validity (AUC value under the ROC curve) of neuroticism constructs is weighted based on behavioral characteristics;

[0095] Relevance dimension: Assess the theoretical correlation between behavioral characteristics and neurotic constructs, based on theoretical attribution metadata in the knowledge base.

[0096] By integrating the three-dimensional scores and confidence weights, the final quantitative T-score of the neuroticism construct is generated, and the stress susceptibility risk level is generated simultaneously.

[0097] Traceable results generation: The generated psychological assessment results include:

[0098] Standardized T-scores and population percentile rankings of neuroticism constructs;

[0099] The list of behavioral evidence corresponding to the seven observable indicators supporting the score includes specific behavioral fragments, timestamps, and modal sources;

[0100] The complete reasoning path from behavioral characteristics to observable indicator units and then to construct scoring, such as: Video frame #1200AU4 activation - indicator 1: negative facial expression - neuroticism score increased by 0.5 points;

[0101] The psychological literature corresponding to each conclusion is based on the DOI list.

[0102] Validation results: This embodiment was validated in the enterprise IT job recruitment scenario. The correlation coefficient r with the gold standard NEO-PI-R scale score was 0.74, the consistency with the clinical expert interview assessment results reached 82%, the overall assessment accuracy was 73.74%, and the test-retest reliability coefficient reached 0.90, which is better than the direct prediction of the general large model with an accuracy of 40%.

[0103] Test-retest reliability and criterion-related validity verification: To verify the stability and accuracy of the assessment results, the STOR-RAG model was validated using the repeated measures method.

[0104] The test was run five times using the same dataset. The results showed an overall test-retest reliability level of ICC(2,k)=0.923 and Cronbach's α=0.926, meeting the excellent reliability standard. Regarding criterion-related validity, the correlation coefficient r with the gold standard NEO-PI-R scale score was 0.74, and the consistency with the clinical expert interview assessment results reached 82%. The objective accuracy was 0.74, and the subjective accuracy reached 0.9025, demonstrating the high professional credibility of the model evaluation results.

[0105] Cross-scenario and cross-cultural stability verification: In 15 cross-scenario cases, experts scored the reports from 0 to 5 based on six dimensions: report fluency, depth and insight, logic, information accuracy, credibility, and interpretability. The results showed M=4.627 and SD=0.220, proving that the model maintains stable overall performance in different application scenarios.

[0106] In terms of cross-cultural validation, the expert scores (M=4.56, SD=0.598) of the samples from European and American white populations and Middle Eastern populations were not significantly different from those of the domestic samples, proving that the model has cross-cultural generalization ability.

[0107] Ablation experiment verification of rule-knowledge-model collaborative mechanism: In order to prove the technical advantages and necessity of collaborative reasoning architecture, an ablation experiment was designed to compare the evaluation accuracy of different architectures.

[0108] The baseline of a single AI large model directly predicts the Big Five personality dimensions using only a general multimodal large model, without accessing the retrieval enhancement knowledge base and expert rules. The average accuracy of the evaluation results is 0.41 (Agreeableness 0.27, Extraversion 0.67, Conscientiousness 0.20, Openness 0.67, Neuroticism 0.27).

[0109] Using the collaborative architecture described in this invention, the average accuracy of the evaluation results is 0.74 (agreeableness 0.70, extraversion 0.84, conscientiousness 0.72, openness 0.66, neuroticism 0.78).

[0110] Experimental results show that the collaborative architecture improves the average accuracy by 80.5% compared to the single-model baseline, with significant improvements in due diligence and human-centeredness, demonstrating the key role of knowledge base constraints and expert rules in improving the accuracy of the assessment.

[0111] As an implementation method, this embodiment addresses the potential black-box risks that may arise during the inference process of large AI models, especially small open-source models such as Llama2-7B, namely, the inference path within the model is not visible, intermediate conclusions that exceed theoretical boundaries are generated, and the model cannot self-correct. This embodiment adopts a layered progressive verification and hard constraint backoff mechanism to ensure that the entire inference process is transparent and controllable and to avoid wasting computational resources.

[0112] (1) Setting up inference chain slices and verification points

[0113] The single inference process of a large AI model is forcibly split into three logical slices, and a knowledge base verification node is inserted after each slice:

[0114] Slice1: Feature semantic description generation. The AI ​​model receives multimodal feature vectors and outputs candidate behavior representations in natural language.

[0115] Verification Point 1: The rule engine verifies whether the description belongs to the explicit behavioral representation based on the observability rule. Only after introspective descriptions are removed can the next slice be moved on.

[0116] Slice2: Candidate matching of observable indicator units. Based on the knowledge base retrieval results, the AI ​​model establishes candidate associations between the behavioral representations output by Slice1 and specific observable indicator units, and outputs confidence scores and knowledge base reference numbers.

[0117] Verification point 2: The knowledge base interface verifies whether the reference number exists and belongs to the current target mental construct. If the verification fails, the reasoning path is immediately blocked and Slice3 is no longer entered.

[0118] Slice3: Evidence strength assessment. The AI ​​model evaluates the evidence weight of the matching result based on the effect size of the literature evidence. For example, a high weight is given because this indicator is associated with the neuroticism effect size r=0.42.

[0119] Verification point 3: The rule engine compares the effect size metadata stored in the knowledge base to verify the rationality of the weights output by the AI.

[0120] (2) Hard constraints on knowledge base and anti-hallucination mechanism

[0121] To avoid the illusions created by small-scale models due to parameter limitations, such as fabricating non-existent knowledge base entries, a structured output template is implemented:

[0122] The system uses hard regular expression validation to ensure that the kb_reference_id field is not empty and exists in the index of the retrieval enhancement knowledge base; if the field is empty or the ID does not exist, the result is marked as unauthorized reasoning and discarded directly, regardless of how high the confidence value is, and will not enter the subsequent scoring process.

[0123] (3) Conflict detection and progressive rollback

[0124] When verification point 2 or verification point 3 finds a conflict between the AI ​​output and the knowledge base rules, instead of requiring the AI ​​to rethink, a gradual rollback mechanism is initiated:

[0125] First conflict: Reduce the confidence weight of the candidate indicator by 50%, mark it as questionable evidence, allow it to enter the collaborative reasoning stage but reduce its influence;

[0126] The second conflict: a direct fallback to the rule engine as a fallback. That is, the specific behavioral characteristic is no longer processed by the AI ​​model, but instead is determined by the expert rule engine based on the mutual exclusion relation ontology in the knowledge base;

[0127] Anti-dead-loop counter: Set a maximum retry threshold for each evaluation task. If the AI ​​model outputs a result rejected by the checkpoint twice in a single evaluation for the same behavioral feature, the system automatically freezes the AI ​​model's further reasoning for that feature and hands it over to the rule engine for processing, completely avoiding the waste of computing power.

[0128] (4) Small model lightweight monitoring adaptation

[0129] Given the limited inference capabilities of the open-source, small-scale model used in this embodiment, instead of employing a complex approach of monitoring a small model with a large model, a lightweight monitoring architecture using a rule engine and a BERT-based classifier is adopted:

[0130] BERT-based Consistency Classifier: A lightweight BERT model is deployed independently, with only 110M parameters, far fewer than the main model. Its sole function is to compare the theoretical_basis text output by the AI ​​model with the standard theoretical description of the metric in the knowledge base to calculate semantic similarity. If the similarity is <0.75, a fallback mechanism is triggered.

[0131] Prompt word injection and solidification: The whitelist of observable indicator units of the current target mental construct is explicitly embedded in the system prompt words, and hard interception is performed before the model output through the keyword filtering layer.

[0132] (5) Process data traceability is achieved

[0133] To achieve transparency in the black box, the system generates inference logs at each verification point in Slices 1-3:

[0134] Record the raw output of the AI ​​model.

[0135] Record the verification results (Pass / Blocked / Fallback);

[0136] Record the basis for knowledge base verification, such as if IDKB-NEU-001 is verified and the theoretical boundary is matched.

[0137] This log, as part of the traceable psychological assessment results, is output along with the final score for reviewers or clinical experts to examine the compliance of the AI ​​reasoning.

[0138] Through the above mechanism, even for small-scale open-source AI models, the reasoning process is constrained within the scope of the knowledge base authorization: the output space is limited in advance by the Prompt whitelist; during the process, illusions are intercepted by slice verification and hard constraints of regular expressions; and after the process, infinite loops are avoided by progressive rollback, ensuring that the system can still generate reliable conclusions by relying on the rule engine when the AI ​​model makes a mistake, thus realizing a rigid security architecture that is AI-assisted but does not make decisions.

[0139] It should be understood that the above-described STOR project implementation scenario is only one specific embodiment of the present invention. In other embodiments, the technical modules may adopt the following parallel alternatives, and all of these alternatives fall within the protection scope defined by the present invention:

[0140] In some embodiments, the target constructs are derived from the DSM-5 clinical diagnostic system, such as the diagnostic criteria for social anxiety disorder. Based on the observability rule and the theory-of-mind consistency rule, a compliant conversion of social anxiety constructs into observable indicator units is completed: The observability rule filters out overt behavioral representations that can be directly captured by third-party devices, such as avoidance of eye contact in interpersonal interactions, unstable speech rhythm, and avoidant body postures in social scenarios, while eliminating subjective self-assessments; the theory-of-mind consistency rule verifies the theoretical fit between the filtered behavioral representations and the social anxiety constructs, ultimately forming three core observable indicator units.

[0141] In other embodiments, the target psychological constructs are derived from the standardized definitions of the ICD-11 International Classification of Diseases or the MMPI Minnesota Multiphasic Personality Inventory, and corresponding stress susceptibility indicator sets or psychopathological indicator sets are constructed respectively.

[0142] In some embodiments, the observability rule not only verifies whether candidate indicators can be captured by audio / video devices, but also further verifies their cross-context stability. It requires that candidate observable indicator units exhibit a stable association with the target mental construct in at least three different social scenarios, eliminating behavioral representations that are overly context-specific.

[0143] In other embodiments, the observability rule employs a third-party verifiability grading mechanism to classify candidate metrics into:

[0144] Level A (can be unambiguously identified by a regular camera)

[0145] Level B (requires specialized sensors such as eye trackers for recognition)

[0146] Level C (requires multimodal fusion algorithm for identification)

[0147] Only A-level and B-level indicators are retained for subsequent processes to ensure the universality of the evaluation system.

[0148] In some embodiments, the retrieval enhancement knowledge base adopts a pure vector database architecture (such as Chroma or Weaviate), and realizes semantic matching between behavioral features and observable indicator units through high-dimensional vector similarity calculation, which is suitable for lightweight deployment scenarios.

[0149] In other embodiments, the retrieval-enhanced knowledge base adopts a pure graph database architecture (such as Amazon Neptune) and uses ontology reasoning and rule engines to perform complex consistency checks of psychological theories, making it suitable for clinical diagnostic scenarios that require deep logical reasoning.

[0150] In a further embodiment, the psychological literature evidence includes not only research DOI, sample size, and effect size, but also links to the original experimental data, cultural background labels of the participants, and data collection date labels. During the inference phase, literature evidence with similar cultural backgrounds and time periods to the subjects being evaluated is prioritized to improve the accuracy of cross-cultural assessment.

[0151] In some embodiments, the collaborative reasoning mechanism employs a hierarchical adjudication model:

[0152] The first layer: The AI ​​model proposes preliminary assessment hypotheses, such as the candidate having a high level of neuroticism;

[0153] The second layer: The expert rule engine performs logical verification based on the mutual exclusion relationships in the knowledge base. For example, if the candidate also exhibits high extraversion, which is theoretically negatively correlated with the neurotic construct, it needs to be reviewed.

[0154] The third layer involves retrieving relevant literature evidence from the knowledge base and calculating the empirical support for the hypothesis.

[0155] Final decision: The final evaluation conclusion will only be output if the hypothesis passes the expert rule verification and the empirical support is >0.7.

[0156] In other embodiments, the collaborative reasoning mechanism employs a human-in-the-loop model. After the AI ​​model completes behavioral feature matching, it submits the intermediate reasoning path to a licensed psychological counselor for manual review; the manual review feedback is then fed back to the system for iterative optimization of the AI ​​model's matching threshold.

[0157] In other embodiments, the collaborative reasoning mechanism employs a multi-model ensemble approach. Three independently trained AI models, such as a BERT-based text model, a Wav2Vec2-based audio model, and a VideoSwinTransformer-based video model, are used to output behavioral feature matching results. The outputs of these multiple models are then fused through a voting mechanism or a weighted average mechanism, and finally, after verification by an expert rule layer, a final evaluation conclusion is generated.

[0158] In some embodiments, the dual verification also includes a timeliness constraint: verifying whether the mapping relationship between candidate observable indicator units and target mental constructs is based on research literature published within the last 10 years, and eliminating indicator-construct associations that are no longer valid due to the evolution of psychological theories.

[0159] In other embodiments, the dual validation also includes a population-specific constraint: validating the applicability of the candidate observable indicator unit in the target assessment population, requiring the indicator to have a predictive validity of not less than 0.75 in that specific population.

[0160] In some embodiments, the AI ​​large model is a multimodal large model (such as GPT-4V, Gemini, Qwen-VL), and the theoretical boundary constraints of the knowledge base are enhanced by prompting engineering injection retrieval.

[0161] The prompt template is fixed as follows: You can only complete behavioral feature matching based on the provided psychology knowledge base, and you are prohibited from outputting association conclusions that exceed the theoretical boundaries of the knowledge base. The knowledge base is limited to: a list of specific observable indicator units of the current target psychological construct, and behavioral feature identification and matching are only based on the above indicators.

[0162] In other embodiments, the large AI model employs a combination of cloud API calls and local lightweight deployment. For high-concurrency scenarios, the cloud-based large model service is called via a RESTful API; for data privacy-sensitive scenarios, a locally deployed lightweight model (such as quantizedLlama2-7B) is used for inference, with model parameters and knowledge base stored on a local server.

[0163] It should be understood that, in addition to the specific embodiments described above, those skilled in the art can make equivalent substitutions or rearrange the steps of the embodiments without departing from the principles of the present invention.

[0164] In some embodiments, the execution order of steps S1-S5 can be adjusted according to specific application scenarios.

[0165] For example, the general retrieval enhancement knowledge base construction in step S3 can be completed in advance to form a general knowledge base covering mainstream psychological constructs such as the Big Five personality traits, social anxiety, and depressive tendencies. Then, for different application scenarios such as corporate recruitment, community screening, and clinical diagnosis, step S1 can be executed to access specific target psychological construct definitions without having to repeatedly build the knowledge base, which greatly improves the efficiency of scenario adaptation.

[0166] In this implementation, the general knowledge base serves as a reusable infrastructure that supports rapid adaptation to new assessment scenarios. By simply updating the definition of the target mental construct and the corresponding mapping of observable indicator units, the deployment of the psychological assessment system for the new scenario can be completed within 24 hours.

[0167] In other embodiments, the AI ​​large model is not limited to adaptive pre-trained language models in the field of psychology; it can also be a general large model implemented by retrieving and enhancing the knowledge base in real time. In this implementation, there is no need to fine-tune the model; relevant psychological theoretical evidence is directly injected into the model in context through knowledge base retrieval, guiding the model to output matching results that conform to the theoretical boundaries.

[0168] As one implementation method, this embodiment also provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the intelligent evaluation method based on observable indicators as described in any of the preceding claims.

[0169] In some embodiments, the computer-readable storage medium includes, but is not limited to, ROM, RAM, magnetic disk, optical disk, USB flash drive, portable hard drive, cloud storage space, etc.

[0170] In other embodiments, the present invention also provides a psychological assessment system applied to the aforementioned methods, used to acquire behavioral data of the subject to be assessed, extract behavioral features, and perform reasoning analysis on psychological constructs to generate psychological assessment results. The system includes:

[0171] Data acquisition layer: includes a high-definition camera array, an omnidirectional microphone array, and a text input interface, used to acquire multimodal behavioral data of the object to be evaluated;

[0172] Feature extraction layer: includes a facial action unit recognition module based on FACS, a pose estimation module based on OpenPose, an acoustic feature extraction module based on PRAAT, and a semantic analysis module based on BERT;

[0173] Reasoning and evaluation layer: includes a retrieval-enhanced knowledge base interface, an AI large-model reasoning engine (configured to call external pre-trained models or locally fine-tuned models), an expert rule engine, and a collaborative adjudication module;

[0174] Results output layer: includes an evaluation report generator, a visualization interface, and an API interface (supporting RESTful calls and returning evaluation results in JSON format).

[0175] It should be understood that the above division of system modules is only a logical functional division, and there may be overlap or merging in actual implementation. For example, the feature extraction layer and the inference evaluation layer can be integrated into the same edge computing device to achieve low-latency real-time evaluation; or deployed on a cloud server to support large-scale parallel processing.

[0176] This invention is not limited to the optional embodiments described above, and anyone can derive other various forms of products based on the inspiration of this invention. The specific embodiments described above should not be construed as limiting the scope of protection of this invention; the scope of protection of this invention should be determined by the claims, and the specification can be used to interpret the claims.

Claims

1. An intelligent evaluation method based on observable indicators, characterized in that, Includes the following steps: First, target psychological constructs derived from standardized psychological scales and / or psychological theory systems are selected. Then, through observability rules and consistency rules of psychological theories, the target psychological constructs are transformed into observable indicator units in compliance with regulations. Finally, a retrieval-enhanced knowledge base is constructed that structurally binds observable indicator units and target psychological constructs and supports the mapping relationship between the two. Based on the non-self-evaluation multimodal behavioral data of the object to be evaluated, the AI ​​big model is used to match behavioral features with observable indicator units through the theoretical boundary constraints of the retrieval enhancement knowledge base. Combined with the preset collaborative reasoning mechanism, the psychological construct assessment is completed, and a traceable psychological assessment result with a complete reasoning path and theoretical basis is generated.

2. The intelligent evaluation method based on observable indicators according to claim 1, characterized in that, The theory-of-mind consistency rule is a rigid mapping constraint rule that runs through the entire process of construct transformation, knowledge base construction, and reasoning evaluation. Its specific implementation method is as follows: In the construct transformation stage, the mapping relationship between observable indicator units and target psychological constructs is verified. The exclusive definition boundaries and causal relationship requirements of the target psychological constructs in their respective psychological theories are strictly matched, and behavioral representations that have no theoretical causal relationship with the target psychological constructs and only have statistical correlation are eliminated. In the knowledge base construction phase, the theoretical definition boundaries of the target mental construct, the mutual exclusion relationship with other mental constructs, and the sub-dimension attribution information are structurally bound and stored with observable indicator units to form a reasoning constraint ontology. During the reasoning and evaluation phase, the feature matching results and intermediate reasoning conclusions output by the AI ​​large model are verified in real time, and association matching and reasoning conclusions that exceed the boundaries defined by the target mental construct theory are filtered out.

3. The intelligent evaluation method based on observable indicators according to claim 1 or 2, characterized in that, The observability rule is the admission verification rule for observable indicator units, specifically: Verify whether the candidate observable indicator units correspond to explicit behavioral representations that can be directly captured by audio and video acquisition devices and text acquisition interfaces. Eliminate introspective descriptions that can only be obtained through subjective self-evaluation by the object to be evaluated. The observable indicator units are unambiguously and repeatedly collected and verified by third-party devices.

4. The intelligent evaluation method based on observable indicators according to claim 1 or 2, characterized in that, The enhanced retrieval knowledge base adopts a hybrid architecture of vector database and graph database, and stores structured data in RDF triple format; The psychological literature evidence includes research DOI, sample size, construct-indicator association effect size, confidence interval, and theoretical attribution metadata.

5. The intelligent evaluation method based on observable indicators according to claim 4, characterized in that, Specifically, the steps include the following: S1. Obtain the standardized definition text of the target psychological construct, wherein the target psychological construct is derived from a standardized psychological scale or a mature psychological theory system that has been validated by peers. S2. The admission verification of candidate observable indicator units is completed through observability rules, and the mapping compliance verification between candidate observable indicator units and target mental constructs is completed through psychological theory consistency rules. The observable indicator units that pass the dual verification are established with the target mental constructs. S3. Construct a retrieval-enhanced knowledge base, which structurally binds and stores observable indicator units, theoretical boundary constraints on the ontology of target psychological constructs, and psychological literature evidence supporting the mapping relationship between the two. S4. Obtain non-self-evaluation multimodal behavioral data of the object to be evaluated, extract corresponding multi-dimensional behavioral features, and based on the theoretical boundary constraints in the retrieval enhancement knowledge base, complete the matching mapping between behavioral features and observable indicator units to generate a set of structured behavioral evidence. S5, based on the collaborative reasoning mechanism of preset expert rules, retrieval-enhanced knowledge base, and AI big model, quantifies and evaluates the behavioral evidence set, generating traceable psychological assessment results.

6. The intelligent evaluation method based on observable indicators according to claim 5, characterized in that, In step S2, the dual verification also includes document support constraints and multimodal identifiability constraints; The literature support constraints are: verify the mapping relationship between candidate observable indicator units and target psychological constructs, have published empirical psychological research data to support them, and retain only candidate indicator units whose associated effect size reaches a preset threshold. The multimodal identifiability constraint is that the candidate observable indicator unit can be identified and verified by behavioral data from at least two of the three modalities: text, audio, and video.

7. The intelligent evaluation method based on observable indicators according to claim 5, characterized in that, In the collaborative reasoning mechanism of step S5, the preset expert rules are used to limit the boundary conditions and mutual exclusion rules of the reasoning logic, the retrieval enhancement knowledge base is used to provide theoretical boundary constraints and empirical support throughout the reasoning process, and the AI ​​big model is only used for pattern recognition of behavioral features and preliminary matching of features and observable indicator units, and does not directly output the final evaluation conclusion of the mental construct.

8. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements an intelligent evaluation method based on observable indicators as described in any one of claims 1-7.

9. A psychological assessment system, characterized in that, The method is applied to any one of the above claims 1-7 to obtain behavioral data of the object to be evaluated, extract behavioral features, and perform reasoning analysis on psychological constructs to generate psychological evaluation results.