Cognitive state evaluation and intervention method and system based on semantic dynamic weighting, electronic device and storage medium
By establishing personalized baselines and aligning multimodal time sequences, dynamically adjusting behavioral feature weights, and combining semantic labels and a policy library, the problems of misjudgment and inaccurate intervention in cognitive state assessment in existing technologies are solved, achieving accurate assessment and effective intervention of cognitive state.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- SHANGHAI HAOYI INFORMATION SCI & TECH CO LTD
- Filing Date
- 2026-03-09
- Publication Date
- 2026-05-08
AI Technical Summary
Existing technologies cannot distinguish different internal states of the same behavioral manifestation in cognitive state assessment. They suffer from temporal mismatch of multimodal data and lack of personalized adaptability, leading to misjudgment and inaccurate intervention.
By collecting initial user behavior data to establish a personalized baseline, using a circular buffer for multimodal temporal backtracking alignment, dynamically adjusting behavioral feature weights, and combining semantic tags and a preset strategy library for precise intervention.
It enables accurate attribution and effective intervention of users' cognitive states, improves the temporal accuracy and personalized adaptability of the analysis, reduces the false alarm rate, and enhances the precision and interpretability of intervention strategies.
Smart Images

Figure CN121809487B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the interdisciplinary fields of artificial intelligence, affective computing and human-computer interaction, and in particular to a method, system, electronic device and storage medium for cognitive state assessment and intervention based on semantic dynamic weighting. Background Technology
[0002] In fields such as intelligent training and human-computer dialogue, real-time and accurate assessment of users' cognitive states is crucial. Existing cognitive state recognition technologies based on multimodal data (such as voice and video) typically determine a user's psychological state by analyzing behavioral characteristics (such as posture, facial expressions, and voice).
[0003] However, these technologies generally suffer from the following drawbacks:
[0004] First, the mapping relationship between features and states is rigid. Existing technologies typically establish static mapping models from "behavioral features to cognitive states." For example, they simply map "body trembling" to "tension." However, in complex interactive situations, the same behavioral manifestation may stem from completely different internal states. For instance, in price negotiations, "body trembling" might indicate stress and tension; but when recalling complex technical details, the same "body trembling" is more likely to indicate deep thinking under high cognitive load. Existing technologies cannot distinguish between such "identical but different meanings," leading to frequent misjudgments of user states.
[0005] Secondly, there is a temporal mismatch in multimodal data. In multimodal analysis, the processing speeds of different modalities vary significantly. For example, capturing visual features such as facial micro-expressions takes milliseconds, while speech recognition and natural language understanding (semantic analysis) typically involve delays of several seconds. Existing technologies often simply match the semantic analysis results of the current moment with the behavioral features of the current moment, leading to incorrect associations between actions from several seconds ago and semantics from several seconds later, resulting in "causal inversion" and severely impacting the accuracy of the analysis.
[0006] Secondly, there is a lack of personalized adaptation capabilities. Different users have their own inherent physiological and behavioral patterns, such as basic speech rate and habitual body swaying amplitude. Most existing systems use globally uniform thresholds for judgment, which can lead to continuous false alarms for certain users with specific behavioral habits, resulting in a low signal-to-noise ratio and poor reliability.
[0007] Therefore, existing technologies still have significant shortcomings in terms of the accuracy, adaptability, and precision of intervention in cognitive state assessment. Summary of the Invention
[0008] The purpose of this application is to provide a method, system, electronic device and storage medium for cognitive state assessment and intervention based on semantic dynamic weighting, which aims to solve the problems existing in the prior art, such as cognitive state misjudgment due to the inability to distinguish between "same meanings and different meanings", temporal mismatch due to data processing delays, poor adaptability due to the lack of personalized models, and difficulty in identifying emotional masquerading, so as to achieve more accurate attribution of user cognitive state and effective intervention.
[0009] To achieve the above objectives, this application provides a cognitive state assessment and intervention method based on semantic dynamic weighting, comprising the following steps: S1: In the initial stage of the interaction process, initial behavioral data of the user is collected, and personalized baseline data is established based on the initial behavioral data; the initial behavioral data includes a first initial modality data stream and a second initial modality data stream, and the personalized baseline data includes a first baseline feature value calculated and stored based on the first initial modality data stream, and a second baseline feature value calculated and stored based on the second initial modality data stream;
[0010] S2: After the initial stage of the interaction process, user behavior data is continuously collected, and the behavior data and the timestamps corresponding to the behavior data are stored in a circular buffer in real time to form a behavior data stream; the behavior data includes a first modal data stream and a second modal data stream, and the circular buffer stores the first modal data stream and the timestamps of the first modal data stream, the second modal data stream and the timestamps of the second modal data stream;
[0011] S3: When a semantic tag is identified in the interaction process, the second modal data stream is parsed to obtain the semantic tag and the time interval corresponding to the semantic tag; based on the time interval, a first modal data segment and a second modal data segment aligned with the time of the semantic tag are retrieved from the circular buffer; features are extracted from the first modal data segment to obtain a first modal feature sequence, and the first modal feature sequence is compared with the first baseline feature value and normalized to obtain a first modal behavioral feature; features are extracted from the second modal data segment to obtain a second modal feature sequence, and the second modal feature sequence is compared with the second baseline feature value and normalized to obtain a second modal behavioral feature;
[0012] S4: Based on the semantic tags, obtain the first feature weight corresponding to the first modal behavior feature and the second feature weight corresponding to the second modal behavior feature from the preset semantic-feature weight mapping table; use the first feature weight to perform a multiplicative weighted calculation on the first modal behavior feature to obtain a first dynamic weighted feature index; use the second feature weight to perform a multiplicative weighted calculation on the second modal behavior feature to obtain a second dynamic weighted feature index.
[0013] S5: Based on the preset cognitive state classification model, the classification result of the cognitive state is obtained according to the first dynamic weighted feature index and the second dynamic weighted feature index. The contribution of each dynamic weighted feature index to the classification result is calculated to determine the main attribution features.
[0014] S6: Based on the semantic tags and the main attribution features, match and execute the corresponding intervention strategies from the preset strategy library.
[0015] Optionally, the method further includes the step: S7: evaluating changes in the user's cognitive state as a feedback signal, and optimizing and updating the feature weights in the semantic label-feature weight correspondence mapping table and / or the intervention strategies in the strategy library using online incremental learning based on the feedback signal; the optimization and update includes adjusting the feature weights corresponding to the main attribution features under the semantic labels based on the difference between the main attribution features and the actual cognitive state of the external feedback, and / or adjusting the priority or parameters of the corresponding intervention strategies in the strategy library based on the immediate or delayed feedback effect obtained by the executed intervention strategies. Wherein, the actual cognitive state of the external feedback includes manually labeled cognitive states or cognitive states reported by the user themselves.
[0016] Optionally, both the first initial modal data stream and the second initial modal data stream are video streams, and both are audio streams. Furthermore, the first modal feature sequence includes a head posture angle sequence, and the first dynamically weighted feature index includes a dynamically weighted posture micro-tension index. The dynamically weighted posture micro-tension index is obtained by: performing a 4Hz to 12Hz bandpass filter on the head posture angle sequence, extracting high-frequency micro-vibration components, calculating the current energy value of the high-frequency micro-vibration components, comparing and normalizing it with the baseline energy value obtained from the first initial modal data stream to obtain a relative rate of change, and then performing the weighted calculation on the relative rate of change; or, the first modal feature sequence includes a facial optical flow change sequence, and the second modal feature stream includes a head posture angle sequence, and the second initial modal data ... The feature sequence includes a harmonic-to-noise ratio (NNR) sequence of speech, and the second dynamically weighted feature index includes a multidimensional micro-expression-acoustic synchronicity index. The multidimensional micro-expression-acoustic synchronicity index is obtained by: calculating the decorrelation measure in time between the facial optical flow change sequence and the normalized NNR sequence, normalizing the decorrelation measure, and performing the weighted calculation on the normalization result; or, the second modal feature sequence includes real-time speech rate, the second baseline feature value includes a speech rate baseline, and the second dynamically weighted feature index includes a semantically weighted speech rate change index. The semantically weighted speech rate change index is obtained by: comparing and normalizing the real-time speech rate with the speech rate baseline to obtain a speech rate change rate, and performing the weighted calculation on the speech rate change rate.
[0017] Optionally, after step S4, the method includes: inputting the semantic label, the primary attribution feature, and the cognitive dissonance score calculated based on the dynamic weighted feature index into a pre-trained decision tree model to classify the user's cognitive state; wherein the cognitive dissonance score is obtained by weighted summation of the dynamic weighted feature indexes, or by inputting the dynamic weighted feature indexes into a preset regression model; or, in the step of executing the intervention strategy, the strategy library is a dual-indexed intervention matrix with the semantic label and the primary attribution feature as dual indices.
[0018] Optionally, when the strategy library is a dual-index intervention matrix, after the step of executing the intervention strategy, the method further includes: evaluating the user's cognitive state change as a feedback signal, and optimizing and updating the feature weights in the semantic label-feature weight correspondence mapping table and / or the intervention strategies in the corresponding units of the dual-index intervention matrix in an online incremental learning manner based on the feedback signal; the optimization and update includes adjusting the feature weights corresponding to the main attribution features under the semantic labels based on the difference between the main attribution features and the actual cognitive state of the external feedback, and / or adjusting the priority or parameters of the corresponding intervention strategies in the dual-index intervention matrix based on the immediate feedback effect or delayed feedback effect obtained by the executed intervention strategy.
[0019] To achieve the above objectives, this application also provides a cognitive state assessment and intervention system based on semantic dynamic weighting, comprising: a baseline calibration module, used to collect initial behavioral data of the user at the initial stage of the interaction process, and establish personalized baseline data based on the initial behavioral data; a data update module, used to continuously collect behavioral data of the user after the initial stage of the interaction process, and store the behavioral data and the timestamps corresponding to the behavioral data in a circular buffer in real time to form a behavioral data stream; a time sequence alignment module, used to identify semantic tags in the interaction process, and use the circular buffer to backtrack and extract behavioral data segments that are time-aligned within the time interval corresponding to the semantic tags; and a dynamic weighting module, used to extract features from the behavioral data segments extracted by the time sequence alignment module to obtain a feature sequence. The system compares and normalizes the feature sequence with the personalized baseline data to obtain at least one behavioral feature; based on the semantic label, it searches for the feature weight corresponding to the behavioral feature under the semantic label from a preset semantic-feature weight mapping table; it uses the feature weight to perform a multiplicative weighted calculation on the behavioral feature to obtain at least one dynamically weighted feature index; the strategy execution module receives the dynamically weighted feature index output by the dynamic weighting module, obtains the classification result of the cognitive state based on the dynamically weighted feature index according to a preset cognitive state classification model, calculates the contribution of each of the dynamically weighted feature indices to the classification result to determine the main attribution feature, and matches and executes the corresponding intervention strategy from a preset strategy library based on the semantic label and the main attribution feature.
[0020] Optionally, the system further includes: an optimization and update module, used to evaluate changes in the user's cognitive state as feedback signals, and to optimize and update the feature weights in the semantic label-feature weight correspondence mapping table and / or the intervention strategies in the strategy library in an online incremental learning manner based on the feedback signals; the optimization and update includes adjusting the feature weights corresponding to the main attribution features under the semantic labels based on the difference between the main attribution features and the actual cognitive state of the external feedback, and / or adjusting the priority or parameters of the corresponding intervention strategies in the strategy library based on the immediate or delayed feedback effects obtained by the executed intervention strategies.
[0021] To achieve the above objectives, this application also provides an electronic device, including a processor and a memory, wherein the memory stores a computer program, and the processor is configured to execute the computer program to implement any of the methods described above.
[0022] To achieve the above objectives, this application also provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements any of the methods described above.
[0023] The technical solution provided in this application has the following beneficial effects: By introducing a semantic dynamic weighting mechanism, the weights of behavioral features can be dynamically adjusted according to the real-time interaction context, enabling the system to accurately distinguish the true internal state of the same behavioral manifestation under different semantics, solving the problem of rigid feature-state mapping, and achieving accurate attribution; By using a circular buffer for multimodal temporal backtracking alignment, the slower-processing semantic results are forced to be precisely matched with the high-speed behavioral features generated during the semantic period, solving the "causal inversion" problem caused by processing delays and ensuring the temporal accuracy of the analysis; By establishing a personalized relative baseline system, all analyses are based on the user's own changes, effectively eliminating the interference caused by individual physiological and behavioral pattern differences, reducing false positives and improving the personalized adaptability of the system; By quantifying the inconsistency of cross-modal features, a calculable technical path is provided for effectively identifying users' deliberate emotional masquerading; By adopting a "semantic-attribution" dual-index strategy library, intervention strategies are doubly bound to triggering scenarios and internal causes, realizing a complete closed loop from identifying the state to understanding the cause and then to precise intervention, enhancing the precision, effectiveness, and interpretability of the intervention strategy. Attached Figure Description
[0024] To more clearly illustrate the technical solutions of the embodiments of this application, the drawings used in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0025] Figure 1 This is a flowchart illustrating the cognitive state assessment and intervention method according to an embodiment of this application.
[0026] Figure 2 This is a schematic diagram of the architecture of a cognitive state assessment and intervention system according to an embodiment of this application.
[0027] Figure 3 This is a schematic diagram of the multimodal data timing backtracking alignment process according to an embodiment of this application.
[0028] Explanation of reference numerals in the attached figures:
[0029] 10: Perception Unit; 20: Temporal Alignment and Semantic Unit; 21: Circular Buffer; 30: Dynamic Weighted Calculation Unit; 31: Weight Table; 40: Cognitive Attribution and Diagnosis Unit; 50: Policy Mapping and Execution Unit; 51: Dual-Index Intervention Matrix. Detailed Implementation
[0030] To make the objectives, technical solutions, and advantages of this application clearer, the application will be further described in detail below with reference to the accompanying drawings. The described embodiments should not be regarded as limitations on this application. All other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0031] Unless otherwise defined, all technical and scientific terms used in the embodiments of this application have the same meaning as commonly understood by one of ordinary skill in the art. The terminology used in the embodiments of this application is for the purpose of describing the embodiments of this application only and is not intended to limit this application.
[0032] Before providing a further detailed description of the embodiments of this application, some of the nouns and terms involved in the embodiments of this application will be explained. The nouns and terms involved in the embodiments of this application are subject to the following interpretations.
[0033] (1) Ring Buffer: This refers to a logically connected First-In-First-Out (FIFO) data structure used to efficiently store real-time generated data streams. It is typically implemented using a buffer (a fixed-length array) and two pointers (a head pointer and a tail pointer, or a read pointer and a write pointer). When the buffer is full, new input data will automatically overwrite the oldest data starting from the beginning. In this application, it is used to temporarily store fixed-length behavioral data with precise timestamps (e.g., multimodal behavioral data) to support fast and efficient backtracking queries for any past time interval. It is a key data structure for solving the problem of time-series mismatch caused by data processing delays.
[0034] (2) Semantic Label: This refers to keywords or classification labels extracted by analyzing the user's speech or text content during the interaction process using Natural Language Processing (NLP) technology. These labels can summarize the current dialogue context, topic, or intent. Examples include "S_Pressure_PriceObjection" and "S_Cognitive_TechnicalDetail". Semantic labels provide contextual basis for subsequent dynamic weighting.
[0035] (3) Primary Attributing Feature: This refers to at least one feature that has the greatest impact on the classification result of the current cognitive state, calculated by contribution analysis methods (such as SHAP value, LIME, etc.) after inputting multiple dynamically weighted feature indicators into the cognitive state classification model. It reveals the root cause of the user's current cognitive state at the physiological or behavioral level and is a key input for achieving precise intervention. Among them, SHAP (SHapley Additive explanations) is a model interpretation framework based on Shapley values (SHAP) used to interpret the prediction results of machine learning models; LIME (Local Interpretable Model-agnostic Explanations) is a local interpretability method used to interpret the prediction results of arbitrary machine learning models, using a "simple model" to locally mimic a "complex model".
[0036] (4) Strategy Library: This refers to a pre-defined or learned dataset used for storing and retrieving intervention strategies. In a preferred embodiment of this application, the strategy library is constructed as an efficient dual-index intervention matrix. The dual-index intervention matrix is a pre-defined strategy library data structure, using a two-dimensional matrix or hash table, with "semantic labels" and "primary attribution features" as the two-dimensional indexes (keys). The matrix cells store precise intervention strategies for specific scenarios (defined by semantic labels) and specific causes (defined by primary attribution features). This structure allows the system to retrieve highly contextualized and cause-oriented precise intervention actions in a deterministic and near-instantaneous manner, avoiding complex real-time decision calculations at the moment of intervention.
[0037] (5) Dynamic Posture Micro-tension Index (d-PMI): This is a composite index used to quantify involuntary physiological responses induced by factors such as intrinsic stress or high cognitive load. The calculation of the dynamic weighted posture micro-tension index includes:
[0038] Step 1, High-frequency micro-tremor component extraction; By performing bandpass filtering on the head posture angle sequence aligned with the semantic label S in a specific frequency band (preferably 4 Hz to 12 Hz, which has been shown to be strongly correlated with stress levels in physiological and psychological studies), the high-frequency micro-tremor component P_high(t) characterizing involuntary micro-tremors is extracted.
[0039] Step 2, Current Energy Value Calculation; The current energy value E_curr is obtained by calculating the root mean square value of the high-frequency micro-vibration component P_high(t) within the analysis window. The calculation formula is as follows:
[0040] ;
[0041] Step 3, normalize the personalized baseline data; calculate the relative rate of change of the current energy value E_curr relative to the baseline energy value E_base, which is the normalized relative change of micro-tension ΔE: ΔE=E_curr / E_base;
[0042] The baseline energy value E_base is obtained as follows: In the initial stage of interaction, the user's first initial modal data stream (e.g., video stream) is collected, the initial head posture angle sequence is extracted from it, a 4 Hz to 12 Hz bandpass filter is performed on the initial head posture angle sequence, the initial high frequency micro-vibration component is extracted, and the root mean square value of the initial high frequency micro-vibration component is calculated, which is the user's personalized baseline energy value E_base.
[0043] Step 4, dynamic weighting; by multiplying the normalized relative change in micro-tension ΔE by the attitude micro-tension weight w_pmi(S) corresponding to the current semantic label in the semantic-feature weight mapping table, the dynamic weighted attitude micro-tension index d-PMI is obtained, that is: d-PMI = w_pmi(S) × ΔE.
[0044] (6) Multidimensional Micro-expression and Acoustic Synchrony Index (m-MAS): This is an index used to quantify cross-modal feature inconsistency, often used to identify higher-order cognitive states such as emotional masquerading. Specifically, it quantifies the inconsistency between facial expression and vocal features by calculating the decorrelation between the facial optical flow change sequence and the speech harmonic-to-noise ratio change sequence, multiplying it by the feature weight corresponding to the current semantic label, and selectively incorporating acoustic tension for correction, thus obtaining the multidimensional micro-expression and acoustic synchrony index.
[0045] (7) Cognitive Dissonance Score (CDS): This is a numerical value that provides a refined and quantifiable assessment of a user's cognitive state. Specifically, it can be calculated based on a weighted sigmoid function or other regression models, using at least one dynamically weighted feature index, to characterize the degree to which a user's cognitive state deviates from its normal state.
[0046] Please see Figure 1 The illustration shows a flowchart of a semantically dynamically weighted cognitive state assessment and intervention method (hereinafter referred to as the "method") provided in an embodiment of this application. This method aims to address problems in existing technologies such as cognitive state misjudgment due to the inability to distinguish between "synonyms," temporal mismatch due to data processing delays, and poor adaptability due to the lack of personalized models. In a basic embodiment, this method can be applied to a human-computer interaction system, such as an intelligent training system or a human-computer dialogue system.
[0047] like Figure 1As shown, the method first executes step S1, which, in the initial stage of the interaction process, collects the user's initial behavioral data and establishes personalized baseline data based on the initial behavioral data. The purpose of this step is to establish a unique, neutral behavioral pattern reference system for each user, thereby eliminating interference from individual differences. Specifically, the system can guide the user to engage in a neutral dialogue or activity unrelated to the core interaction task, such as reading a neutral text or giving a simple self-introduction. In a specific embodiment, during this period, the system collects the user's initial behavioral data through sensing devices (such as a camera or microphone), such as a first initial modal data stream containing facial expressions and head posture, and a second initial modal data stream containing voice tone and speech rate. Subsequently, the system extracts features from these initial behavioral data, calculates and stores a series of baseline feature values to form personalized baseline data. For example, it calculates the user's average head sway energy, average speech rate, and fundamental frequency range in a relaxed state.
[0048] Next, step S2 is executed. After the initial stage of the interaction process, the formal interaction phase begins. The system continuously collects user behavior data and stores the behavior data and its corresponding timestamps in a circular buffer in real time, forming a behavior data stream. Please refer to [reference needed]. Figure 2 and Figure 3 The circular buffer 21 is a fixed-size data structure, for example, capable of storing data from the past 60 seconds. When new behavioral data (such as a video frame or an audio clip) and its corresponding timestamp (e.g., accurate to milliseconds) are generated, they are written to the end of the circular buffer 21. If the circular buffer 21 is full, the oldest data will be overwritten by the new data. This "first-in, first-out" mechanism ensures that the system always retains a complete record of behavior within the most recent period, providing a data foundation for subsequent backtracking analysis. In the specific embodiment described above, the behavioral data may, for example, be a first modal data stream of a video stream and a second modal data stream of an audio stream. Based on this, the circular buffer stores the first modal data stream and its timestamp, and the second modal data segment and its timestamp.
[0049] Then, step S3 is executed: when the system recognizes a semantic tag in the interaction process, it extracts the behavioral data segment within the time interval corresponding to the semantic tag from the circular buffer. For example... Figure 3As shown, this step is crucial for resolving the "temporal mismatch" problem in multimodal data. For example, when a user's voice stream is sent to the NLP module for analysis, this process may involve a delay of hundreds of milliseconds or even several seconds Δt. When the NLP module completes processing at time T and outputs a semantic label (e.g., "price objection") and its corresponding time interval [T_start, T_end], this time interval actually corresponds to a past voice event. At this point, the system does not use the behavioral data at time T, but instead uses this time interval [T_start, T_end] as a query index to initiate a "backtracking query" to the circular buffer 21, precisely extracting the behavioral data segment recorded within that time period. Subsequently, the system performs feature extraction on this behavioral data segment to obtain a feature sequence, and compares and normalizes the feature sequence with the personalized baseline data established in step S1 to obtain at least one behavioral feature. For example, calculating the rate of change of the current head shaking energy relative to the baseline energy.
[0050] Subsequently, step S4 is executed. Based on the semantic label, the feature weight corresponding to the behavioral feature under the semantic label is found from the preset semantic-feature weight mapping table (i.e., weight table 31 in the attached figure). The feature weight is then used to perform a multiplicative weighted calculation on the behavioral feature to obtain at least one dynamically weighted feature index. For example, for the behavioral feature "head shaking," if the current semantic label is "thinking about complex problems," its corresponding weight in weight table 31 may be low (e.g., 0.5), because the shaking at this time is likely to represent cognitive load; while if the semantic label is "encountering strong rebuttal," its corresponding weight may be high (e.g., 1.8), because the shaking at this time is more likely to represent stress and tension. In this way, the system can dynamically adjust the interpretation of the same behavioral manifestation according to the interaction context, solving the problem of "same meaning, different interpretations."
[0051] Next, step S5 is executed. Based on a preset cognitive state classification model, the classification result of the cognitive state is obtained according to the dynamically weighted feature indicators, and the contribution of each dynamically weighted feature indicator to the classification result is calculated to determine the primary attribution feature. The classification model can be a support vector machine, neural network, or logistic regression, etc. For example, after receiving multiple dynamically weighted feature indicators, the classification model outputs the classification result as "stress". Then, the system uses an interpretability analysis technique (such as SHAP) to analyze whether the "dynamically weighted posture indicator" or the "dynamically weighted acoustic indicator" contributes the most in this determination of the "stress" state. The behavioral feature corresponding to the dynamically weighted feature indicator with the largest contribution is determined as the primary attribution feature.
[0052] Finally, step S6 is executed, whereby, based on the semantic tags and the primary attribution features, a corresponding intervention strategy is matched and executed from a preset strategy library. The system no longer simply intervenes in a state of "tension," but rather in a state of "tension caused by 'postural instability' in a 'price negotiation' scenario." For example, the strategy library can be a dual-indexed intervention matrix 51, where the system uses semantic tags and primary attribution features as dual indices to find the most matching intervention strategy, such as "stress testing by having a virtual human display a more assertive posture" or "guiding the user to relax through gentle language."
[0053] The above embodiments address the problem of poor adaptability by establishing a personalized baseline, the problem of temporal mismatch by using a circular buffer and backtracking query, the problem of homonymy by using semantic dynamic weighting, and the problem of accurate and interpretable feedback loop by using semantic and attribution-based intervention, thereby improving the accuracy of cognitive state assessment and the effectiveness of intervention.
[0054] Furthermore, in a preferred embodiment, the multimodal data in the above embodiments will be described in detail. Please refer to [reference needed]. Figure 3 The initial behavioral data and the behavioral data may include data from multiple modalities. For example, the first initial modal data stream and the second initial modal data stream are both video streams, collected by devices such as cameras; the second initial modal data stream and the second modal data stream are both audio streams, collected by devices such as microphones. Accordingly, in step S1, the personalized baseline data includes a first baseline feature value (such as a head pose energy baseline) calculated based on the video stream and a second baseline feature value (such as a speech rate baseline and a harmonic-to-noise ratio baseline) calculated based on the audio stream.
[0055] In this embodiment, the circular buffer 21 in step S2 simultaneously stores both timestamped video and audio data streams. Step S3 is more detailed: the system parses the second modality data stream (e.g., the audio stream), for example, by performing Automatic Speech Recognition (ASR) and Natural Language Processing (NLP) to obtain semantic tags and their corresponding time intervals. The accuracy of this time interval is crucial; preferably, its timestamp resolution is accurate to 1 millisecond. Subsequently, based on this time interval, the system backtracks through the circular buffer 21 to extract the first modality data segment (video segment) and the second modality data segment (audio segment) that are strictly aligned in time with the semantic tags. Time alignment means that the absolute value of the difference between the timestamps of the extracted first modality data stream and the second modality data stream is extremely small, for example, less than or equal to 1 millisecond, thus ensuring the causal correctness of the analysis.
[0056] After extracting the aligned data segments, the system performs feature extraction. For example, the first modality data segment (video segment) is processed to obtain a first modality feature sequence (such as a head pose angle sequence), which is then compared with and normalized to obtain the first modality behavioral feature. Simultaneously, the second modality data segment (audio segment) is processed to obtain a second modality feature sequence (such as a harmonic-to-noise ratio sequence), which is then compared with and normalized to obtain the second modality behavioral feature. In step S4, based on the semantic labels, the system retrieves the first feature weight corresponding to the first modality behavioral feature and the second feature weight corresponding to the second modality behavioral feature from the semantic-feature weight mapping table (weight table 31). The first feature weight is used to weight the first modality behavioral feature to obtain a first dynamic weighted feature index; the second feature weight is used to weight the second modality behavioral feature to obtain a second dynamic weighted feature index.
[0057] In one specific implementation, the dynamically weighted feature index can be a composite index with clear physical or psychological significance. For example, the first modal feature sequence may include a head posture angle sequence, and the first dynamically weighted feature index may specifically be a dynamically weighted posture micro-tension index (d-PMI). Its calculation method is as follows: First, a bandpass filter of 4 Hz to 12 Hz is applied to the head posture angle sequence. Signals in this frequency band are usually associated with involuntary physiological tremors and are reliable physiological representations of states such as tension and stress. Then, high-frequency tremor components are extracted, and the current energy value of the filtered high-frequency tremor components is calculated. The current energy value is compared with and normalized to the baseline energy value of the first initial modal data stream (neutral state video) obtained in step S1 to obtain a relative rate of change. Finally, the relative rate of change is weighted, that is, the relative rate of change is multiplied by the posture micro-tension weight corresponding to the current semantic label found in weight table 31 to obtain the final d-PMI value. Its calculation formula can be expressed as:
[0058] d-PMI = w_pmi(S) × (E_curr / E_base)
[0059] Wherein, d-PMI represents the dynamically weighted posture micro-tension index; w_pmi(S) represents the weight corresponding to the posture micro-tension feature under the semantic label S; E_curr represents the energy of the high-frequency micro-vibration component within the current time window; and E_base represents the energy of the high-frequency micro-vibration component at the baseline state. This dynamically weighted posture micro-tension index can sensitively capture the physiological reactions that users cannot control voluntarily due to internal stress, and through semantic weighting, it effectively distinguishes tension under stress from body swaying caused by other reasons (such as fatigue).
[0060] In another specific implementation, to identify potential emotional charades by users, the first modal feature sequence may include a facial optical flow change sequence, the second modal feature sequence may include a harmonic-to-noise ratio (HNR) sequence, and the second dynamically weighted feature index may specifically be a multidimensional micro-expression-acoustic synchronicity index (m-MAS). The m-MAS is calculated as follows: First, the temporal decorrelation measure between the facial optical flow change sequence (representing minute facial muscle movements) and the normalized HNR sequence (representing sound stability) is calculated. For example, the absolute value of the Pearson correlation coefficient is calculated, and then this absolute value is subtracted from 1. If a user expresses a positive emotion, their facial expression is relaxed but their voice is tense (HNR decreases), the correlation between the two decreases, and the decorrelation measure increases. Then, the decorrelation measure is normalized, and the normalization result is weighted, i.e., multiplied by the synchronicity weight corresponding to the current semantic label. The calculation formula can be expressed as:
[0061] m-MAS = (1 - |ρ|_avg) × w_mas(S) × (1 + λ × |ΔHNR_std / HNR_base_std|)
[0062] Wherein, m-MAS represents the multidimensional micro-expression-acoustic synchronicity index; |ρ|_avg represents the absolute value of the average correlation coefficient between the facial optical flow change sequence and the harmonic-to-noise ratio sequence over time, therefore (1 -|ρ|_avg) represents the average decorrelation; w_mas(S) represents the second feature weight corresponding to the cross-modal synchronicity feature under the semantic label S; ΔHNR_std / HNR_base_std represents the relative acoustic tension change, used to correct the sensitivity of the index; λ is a weighting coefficient used to adjust the influence of the acoustic tension correction term. This multidimensional micro-expression-acoustic synchronicity index provides an effective technical means to identify higher-order cognitive states such as "saying one thing but meaning another" and "pretending to be calm" by quantifying the inconsistency between visual and auditory channel information, thus improving the depth and accuracy of the assessment.
[0063] In another specific implementation, the second modal feature sequence may include real-time speech rate, the second baseline feature value may include a speech rate baseline, and the second dynamic weighted feature index may specifically be a semantically weighted speech rate change index. The semantically weighted speech rate change index is calculated by comparing and normalizing the real-time calculated speech rate with the speech rate baseline established in step S1 to obtain the speech rate change rate, and then performing a weighted calculation on this rate change. For example, under the semantics of "technical Q&A," a slower speech rate may indicate deep thinking, resulting in a lower weight for the second feature; while under the semantics of "casual conversation," a sudden slowdown in speech rate may indicate hesitation or uncertainty, resulting in a higher weight for the second feature. This approach makes the interpretation of speech rate changes more flexible and accurate.
[0064] Furthermore, the quantification of cognitive states and the implementation methods of intervention strategies are refined. In a preferred embodiment, step S5 may specifically include calculating the cognitive dissonance score (CDS). For example, multiple dynamically weighted feature indicators, such as the aforementioned dynamically weighted postural micro-tension index (d-PMI) and multidimensional micro-expression-acoustic synchronicity index (m-MAS), are input into a preset regression model (such as a weighted sigmoid activation function) to obtain a continuous value, i.e., the cognitive dissonance score. The formula for calculating the cognitive dissonance score can be expressed as:
[0065] CDS = σ(α·d-PMI + β·m-MAS + γ·F_other + b)
[0066] In this system, CDS represents the cognitive dissonance score; σ is the sigmoid activation function, mapping the output value to between 0 and 1; d-PMI and m-MAS are the aforementioned dynamically weighted feature indicators; F_other represents other weighted feature indicators; α, β, and γ are the trainable weight parameters for each indicator, reflecting the degree of contribution of different features to cognitive dissonance; and b is the bias term. The system can classify the user's cognitive state based on whether the cognitive dissonance score exceeds a preset threshold (e.g., 0.8), obtaining a classification result (e.g., "cognitive dissonance" or "normal state"). This quantitative approach makes state assessment more refined, facilitating degree judgment and trend analysis.
[0067] Accordingly, in step S6, the strategy library can be specifically implemented as a dual-indexed intervention matrix 51 with semantic labels and primary attribution features as dual indices. When the system determines that the current semantic label is "price objection" and the primary attribution feature is "d-PMI" (i.e., excessive micro-tension of posture), it will locate the cell ["price objection", "d-PMI"] in the dual-indexed intervention matrix and execute the preset intervention strategy therein, such as "applying adversarial pressure to conduct desensitization training". The technical effect of this dual-indexing mechanism is that it deeply binds the intervention strategy with the "triggering scenario (semantics)" and the "intrinsic cause (attribution)", realizing a complete closed loop from "identifying the state" to "understanding the cause" to "precise intervention", enhancing the precision, effectiveness and interpretability of the intervention strategy.
[0068] In an optional implementation, after step S4, the semantic label, the primary attribution features, and the cognitive dissonance score calculated based on the dynamically weighted feature index are all input features to a pre-trained decision tree model to determine and execute the corresponding intervention strategy. The leaf nodes of this decision tree model directly correspond to specific intervention strategies. Therefore, through a single model inference, the input features can be directly mapped to the final intervention strategy to be executed, thus achieving end-to-end decision-making from state assessment to strategy execution. This approach simplifies the system process and is particularly suitable for scenarios where there is a complex nonlinear relationship between the intervention strategy and the input features.
[0069] In a preferred embodiment, to enable the system to adaptively optimize, this application also provides an online incremental learning mechanism. After step S6, the system can evaluate the user's cognitive state change as a feedback signal, and optimize and update the feature weights in the semantic-feature weight mapping table (weight table 31) and / or the intervention strategies in the strategy library (such as the dual-index intervention matrix 51) through online incremental learning based on the feedback signal. The feedback signal can be the degree of improvement in the cognitive state after intervention automatically evaluated by the system, or it can be feedback from external input, such as manually labeled cognitive state or user self-reported cognitive state (e.g., a human coach's rating of the intervention effect).
[0070] The specific optimization and update process may include adjusting the feature weights corresponding to the primary attribution features under the semantic label based on the difference between the primary attribution features and the actual cognitive state of external feedback. For example, if the system determines that a user is tense based on a high d-PMI, but the coach reports that the user is actually excited, the system will reduce the weight of d-PMI under the current semantics. Simultaneously, the system can also adjust the priority or parameters of corresponding intervention strategies in the strategy library based on the immediate or delayed feedback effects obtained from the implemented intervention strategies. For example, if the user's cognitive dissonance score significantly decreases after multiple executions of strategy A, the priority of intervention strategy A in the corresponding cell of the dual-index intervention matrix 51 will be increased. This online incremental learning mechanism enables the system to continuously learn from interactions and iteratively optimize its model, thereby making the assessment of specific users or user groups increasingly accurate and the intervention increasingly effective, achieving true personalized adaptation.
[0071] Please see Figure 2 This application also provides a cognitive state assessment and intervention system based on semantic dynamic weighting. This system can be a hardware platform for implementing the above-described methods. The system includes a baseline calibration module, a data update module, a temporal alignment module, a dynamic weighting module, and a policy execution module. The functions of these modules correspond one-to-one with the aforementioned method steps and can be implemented through software, hardware, or a combination of both. Specifically:
[0072] The baseline calibration module is used to collect initial behavioral data of the user at the initial stage of the interaction process and to establish personalized baseline data based on the initial behavioral data.
[0073] The data update module is used to continuously collect user behavior data after the initial stage of the interaction process, and store the behavior data and the timestamps corresponding to the behavior data in a circular buffer in real time to form a behavior data stream.
[0074] The time alignment module is used to identify semantic tags in the interaction process and use the circular buffer to backtrack and extract time-aligned behavioral data segments within the time interval corresponding to the semantic tags;
[0075] The dynamic weighting module is used to extract features from the behavioral data segments extracted by the time-series alignment module to obtain a feature sequence, compare the feature sequence with the personalized baseline data and normalize it to obtain at least one behavioral feature; according to the semantic label, find the feature weight corresponding to the behavioral feature under the semantic label from a preset semantic-feature weight mapping table; use the feature weight to perform a multiplicative weighted calculation on the behavioral feature to obtain at least one dynamically weighted feature index;
[0076] The strategy execution module is used to receive the dynamically weighted feature indexes output by the dynamic weighting module, obtain the classification result of the cognitive state based on the dynamically weighted feature indexes according to the preset cognitive state classification model, calculate the contribution of each of the dynamically weighted feature indexes to the classification result to determine the main attribution features, and match and execute the corresponding intervention strategy from the preset strategy library based on the semantic label and the main attribution features.
[0077] Specifically, in an electronic device (such as a smart interactive screen, server, or personal computer), the above system can be implemented as follows: the electronic device includes at least one processor and a memory, in which a computer program is stored. When the processor executes the computer program, it implements the functions of the aforementioned modules and the aforementioned method. Alternatively, a computer-readable storage medium stores a computer program thereon, which, when executed by the processor, implements the functions of the aforementioned modules and the aforementioned method. The sensing unit 10 may include a built-in or external camera and microphone for collecting audio and video data from the user. The baseline calibration module is configured to, during the initial stage of interaction, control the sensing unit 10 to collect initial behavioral data and process this initial behavioral data to establish personalized baseline data, which is then stored in the memory.
[0078] The data update module is configured to continuously acquire behavioral data streams from the sensing unit 10 during the interaction and write them, along with a high-precision timestamp, into a circular buffer 21 in memory. The timing alignment module (corresponding to...) Figure 2 The temporal alignment and semantic unit 20 is configured to receive semantic tags and time intervals parsed by the NLP engine (which can be built-in or called from the cloud), and perform a backtracking query on the circular buffer 21 based on the time interval to extract the time-aligned behavioral data segment.
[0079] Dynamic weighting module (corresponding to) Figure 2 The dynamic weighted calculation unit 30 is configured to extract features from the extracted behavioral data segments, compare them with the baseline, normalize them, and then look up weights from the weight table 31 loaded in memory based on semantic labels, perform weighted calculations, and output dynamic weighted feature indicators. The strategy execution module is configured to receive these indicators, call a preset classification model to perform state classification, perform contribution analysis to determine the main attribution features, and finally match and execute intervention strategies from the dual-index intervention matrix 51 loaded in memory based on semantic labels and main attribution features. The execution of the intervention strategy can be through displaying specific content on the screen, playing specific voice through a speaker, or controlling the behavior of a virtual human. The function of this strategy execution module integrates... Figure 2 The functions of the cognitive attribution and diagnosis unit 40 and the strategy mapping and execution unit 50 are shown.
[0080] Furthermore, the system may also include an optimization and update module for evaluating changes in the user's cognitive state as feedback signals, and optimizing and updating the feature weights in the semantic label-feature weight correspondence mapping table and / or the intervention strategies in the strategy library through online incremental learning based on the feedback signals. The optimization and update includes adjusting the feature weights corresponding to the main attribution features under the semantic labels based on the difference between the main attribution features and the actual cognitive state of the external feedback, and / or adjusting the priority or parameters of the corresponding intervention strategies in the strategy library based on the immediate or delayed feedback effects obtained by the executed intervention strategies. This module is configured to receive external feedback signals (such as user ratings input via a touchscreen) and, based on these signals, use online incremental learning algorithms (such as reinforcement learning or gradient descent) to adjust the parameters of the weight table 31 and the dual-index intervention matrix 51 stored in memory. This makes the entire system a self-evolving intelligent system.
[0081] To better understand this application, a specific embodiment will be used below to illustrate all the technical features. Assume this application is applied to a business negotiation training system for sales personnel, which uses a large smart screen with a camera and microphone as the interactive terminal, displaying an AI virtual negotiation opponent.
[0082] First, when student Zhang San uses the system for the first time, the system executes S1, guiding Zhang San to engage in two minutes of neutral activities such as looking at the scenery and chatting. The baseline calibration module collects audio and video data of this process through the sensing unit 10, calculates and stores the head posture energy baseline E_base=0.15, as well as the speech rate baseline, harmonic-to-noise ratio baseline, etc.
[0083] Next, the training began, moving into the price negotiation phase. The data update module started working, continuously writing Zhang San's real-time video frames and audio clips, along with millisecond-level timestamps, into the circular buffer 21.
[0084] The AI opponent says, "Your price is far too high compared to our market research results; we cannot accept it at all." This audio is sent to the NLP engine. 1.5 seconds later, the timing alignment module receives the semantic tag S = "S_Pressure_PriceObjection" and the corresponding time interval [T1, T2]. The timing alignment module immediately performs a backtracking query on the circular buffer 21 to extract Zhang San's video frames and audio segments within the time interval [T1, T2].
[0085] The dynamic weighting module begins processing these data segments. It extracts the head pose angle sequence from the video, bandpass filters it, and calculates the current micro-vibration energy E_curr = 0.45. Simultaneously, it extracts the harmonic-to-noise ratio sequence from the audio and compares it with facial optical flow. Looking up the weights in table 31, under the semantics of "S_Pressure_PriceObjection," the weight w_pmi for pose micro-tension is 1.9, while the feature weight w_mas for cross-modal synchronicity is 0.8. It calculates d-PMI = 1.9 × (0.45 / 0.15) = 5.7. Simultaneously, it calculates a low m-MAS value, for example, 0.3, indicating that Zhang San is not deliberately faking it and that it is a genuine physiological response.
[0086] The strategy execution module received indicators such as d-PMI=5.7 and m-MAS=0.3. It input these indicators into a pre-defined regression model, calculating a cognitive dissonance score (CDS) of 0.92, far exceeding the threshold of 0.8, thus classifying it as "highly stressed." Subsequent contribution analysis showed that d-PMI was the primary attribution characteristic for this judgment.
[0087] Subsequently, the policy execution module used {semantic="S_Pressure_PriceObjection", attribution="d-PMI"} as a dual index to query the dual-index intervention matrix 51. The found policy was "Non-verbal adversarial training under high-pressure situations". The system then prompted the AI opponent on the screen to lean forward slightly, sharpen its gaze, and say in a steady but authoritative tone: "Is that so? Then please explain in detail what cost components your quote is based on?"
[0088] After the training session, the live coach gave the intervention a score of 4 out of 5 for effectiveness. The optimization and update module received this positive feedback. It fine-tuned the internal model, for example, slightly increasing the weight of d-PMI under the semantics of "S_Pressure_PriceObjection" and reinforcing the priority of the "nonverbal adversarial training under high-pressure situations" strategy in this scenario. Through multiple training sessions and feedback from multiple participants, the system's weight table 31 and dual-index intervention matrix 51 will become increasingly optimized, achieving precise control over the trainees' states and efficient training.
[0089] The technical solutions provided in this application are not limited to the embodiments described above, and can be widely applied in various scenarios requiring real-time assessment and intervention of human cognitive states. For example, in high-risk positions such as aerospace and financial transactions, this system can be used to monitor the cognitive load and mental stress of operators in real time. When cognitive dissonance caused by excessive tension or fatigue is detected, timely warnings or suggestions for rest are issued to ensure operational safety.
[0090] In intelligent education scenarios, this system can be used to analyze students' focus and comprehension during the learning process. When an AI teacher explains a complex concept, if the system detects and attributes a student's d-PMI (representing micro-tremors under cognitive load) to a sustained increase, the AI teacher can determine that the student may be experiencing difficulty in understanding. It can then proactively switch to a simpler explanation method or provide a concrete example to achieve personalized teaching and guidance.
[0091] In summary, this application constructs a closed loop for cognitive state assessment and intervention through a series of interconnected technical means, enabling personalized, context-aware, time-precise, attributable, explainable, and adaptively optimized cognitive state assessment. This improves the accuracy and reliability of cognitive state assessment and broadens its application scenarios.
[0092] Those skilled in the art will understand that the systems and methods of the above embodiments can be implemented by a computer program, which can be stored in a computer-readable storage medium, such as a hard disk, optical disk, or flash memory. When the processor executes the program, it implements the steps of any of the above method embodiments or the functions of each module in any of the above system embodiments. Similarly, the electronic device provided in this application is also driven by a processor executing a program in memory to implement the technical solution of this application.
Claims
1. A method for assessing and intervening in cognitive states based on semantically dynamic weighting, characterized in that, Includes the following steps: S1: In the initial stage of the interaction process, the user's initial behavior data is collected, and personalized baseline data is established based on the initial behavior data; the initial behavior data includes a first initial modality data stream and a second initial modality data stream, and the personalized baseline data includes a first baseline feature value calculated and stored based on the first initial modality data stream, and a second baseline feature value calculated and stored based on the second initial modality data stream; S2: After the initial stage of the interaction process, continuously collect user behavior data, and store the behavior data and the timestamps corresponding to the behavior data in a circular buffer in real time to form a behavior data stream; The behavioral data includes a first modal data stream and a second modal data stream. The circular buffer stores the first modal data stream and its timestamp, as well as the second modal data stream and its timestamp. S3: When a semantic tag is identified in the interaction process, the second modal data stream is parsed to obtain the semantic tag and the time interval corresponding to the semantic tag; based on the time interval, a first modal data segment and a second modal data segment aligned with the time of the semantic tag are extracted from the circular buffer; features are extracted from the first modal data segment to obtain a first modal feature sequence; the first modal feature sequence is compared with the first baseline feature value and normalized to obtain the first modal behavior feature; Feature extraction is performed on the second modality data segment to obtain the second modality feature sequence. The second modality feature sequence is compared with the second baseline feature value and normalized to obtain the second modality behavior feature. S4: Based on the semantic tags, obtain the first feature weight corresponding to the first modal behavior feature and the second feature weight corresponding to the second modal behavior feature from the preset semantic-feature weight mapping table; use the first feature weight to perform a multiplicative weighted calculation on the first modal behavior feature to obtain a first dynamic weighted feature index; use the second feature weight to perform a multiplicative weighted calculation on the second modal behavior feature to obtain a second dynamic weighted feature index. S5: Based on the preset cognitive state classification model, the classification result of the cognitive state is obtained according to the dynamic weighted feature index, and the contribution of each of the dynamic weighted feature indexes to the classification result is calculated to determine the main attribution features. S6: Based on the semantic tags and the main attribution features, match and execute the corresponding intervention strategy from the preset strategy library.
2. The method according to claim 1, characterized in that, The method also includes the following steps: S7: Evaluate changes in the user's cognitive state as a feedback signal, and optimize and update the feature weights in the semantic label-feature weight correspondence mapping table and / or the intervention strategies in the strategy library using online incremental learning based on the feedback signal; the optimization and update include, Based on the difference between the primary attribution features and the actual cognitive state of external feedback, the feature weights corresponding to the primary attribution features under the semantic tags are adjusted, and / or, based on the immediate or delayed feedback effect obtained by the implemented intervention strategy, the priority or parameters of the corresponding intervention strategies in the strategy library are adjusted.
3. The method according to claim 1, characterized in that, Both the first initial modal data stream and the second modal data stream are video streams, and both the second initial modal data stream and the second modal data stream are audio streams. The first modal feature sequence includes a head posture angle sequence, and the first dynamic weighted feature index includes a dynamic weighted posture micro-tension index. The dynamic weighted posture micro-tension index is obtained by performing a bandpass filter of 4 Hz to 12 Hz on the head posture angle sequence to extract high-frequency micro-vibration components, calculating the current energy value of the high-frequency micro-vibration components, comparing it with the baseline energy value obtained from the first initial modal data stream and normalizing it to obtain a relative rate of change, and then performing the weighted calculation on the relative rate of change. or, The first modal feature sequence includes a facial optical flow change sequence, the second modal feature sequence includes a harmonic noise ratio (NNR) sequence of speech, and the second dynamic weighted feature index includes a multidimensional micro-expression-acoustic synchronicity index; the multidimensional micro-expression-acoustic synchronicity index is obtained by: calculating the decorrelation measure in time between the facial optical flow change sequence and the normalized NNR sequence, normalizing the decorrelation measure, and performing the weighted calculation on the normalization result; or, The second modal feature sequence includes real-time speech rate, the second baseline feature value includes a speech rate baseline, and the second dynamic weighted feature index includes a semantically weighted speech rate change index; the semantically weighted speech rate change index is obtained by comparing the real-time speech rate with the speech rate baseline and normalizing them to obtain a speech rate change rate, and then performing the weighted calculation on the speech rate change rate.
4. The method according to claim 1, characterized in that, Step S4 is followed by: inputting the semantic labels, the primary attribution features, and the cognitive dissonance score calculated based on the dynamically weighted feature indicators into a pre-trained decision tree model to determine and execute corresponding intervention strategies. The cognitive dissonance score is obtained by weighted summation of the dynamically weighted feature indicators, or by inputting the dynamically weighted feature indicators into a preset regression model; or... In step S6, the strategy library is a dual-indexed intervention matrix with the semantic tags and the main attribution features as dual indices.
5. The method according to claim 4, characterized in that, When the strategy library is a dual-index intervention matrix, after step S6, the method further includes: evaluating changes in the user's cognitive state as a feedback signal, and optimizing and updating the feature weights in the semantic label-feature weight correspondence mapping table and / or the intervention strategies in the corresponding units of the dual-index intervention matrix in an online incremental learning manner based on the feedback signal; the optimization and update includes adjusting the feature weights corresponding to the main attribution features under the semantic labels based on the difference between the main attribution features and the actual cognitive state of the external feedback, and / or adjusting the priority or parameters of the corresponding intervention strategies in the dual-index intervention matrix based on the immediate or delayed feedback effect obtained by the executed intervention strategies.
6. A cognitive state assessment and intervention system based on semantic dynamic weighting, characterized in that, include: The baseline calibration module is used to collect initial behavioral data of the user at the initial stage of the interaction process and to establish personalized baseline data based on the initial behavioral data. The data update module is used to continuously collect user behavior data after the initial stage of the interaction process, and store the behavior data and the timestamps corresponding to the behavior data in a circular buffer in real time to form a behavior data stream. The time alignment module is used to identify semantic tags in the interaction process and use the circular buffer to backtrack and extract time-aligned behavioral data segments within the time interval corresponding to the semantic tags; The dynamic weighting module is used to extract features from the behavioral data segments extracted by the time alignment module to obtain a feature sequence, compare the feature sequence with the personalized baseline data and normalize it to obtain at least one behavioral feature. Based on the semantic label, the feature weight corresponding to the behavioral feature under the semantic label is found from the preset semantic-feature weight mapping table; the feature weight is used to perform a multiplicative weighted calculation on the behavioral feature to obtain at least one dynamically weighted feature index. The strategy execution module is used to receive the dynamically weighted feature indexes output by the dynamic weighting module, obtain the classification result of the cognitive state based on the dynamically weighted feature indexes according to the preset cognitive state classification model, calculate the contribution of each of the dynamically weighted feature indexes to the classification result to determine the main attribution features, and match and execute the corresponding intervention strategy from the preset strategy library based on the semantic label and the main attribution features.
7. The system according to claim 6, characterized in that, Also includes: The optimization and update module is used to evaluate changes in the user's cognitive state as feedback signals, and to optimize and update the feature weights in the semantic label-feature weight correspondence mapping table and / or the intervention strategies in the strategy library through online incremental learning based on the feedback signals. The optimization and update includes adjusting the feature weights corresponding to the main attribution features under the semantic labels based on the difference between the main attribution features and the actual cognitive state of the external feedback, and / or adjusting the priority or parameters of the corresponding intervention strategies in the strategy library based on the immediate or delayed feedback effects obtained by the executed intervention strategies.
8. An electronic device, characterized in that, It includes a processor and a memory, wherein the memory stores a computer program, and the processor is configured to, when executing the computer program, implement the method as described in any one of claims 1 to 5.
9. A computer-readable storage medium, characterized in that, It stores a computer program thereon, which, when executed by a processor, implements the method as described in any one of claims 1 to 5.
Citation Information
Patent Citations
Remote education data processing system
CN120471275A
English teaching training system and method fusing semantic matching and cognitive evaluation
CN120632398A