A multi-modal virtual reality closed-loop cognitive adaptive psychological assessment method and system based on reinforcement learning
Patent Information
- Application Number
- CN202610765120.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-05-29
- Publication Date
- 2026-09-25
AI Technical Summary
[0012]本发明的目的在于提供一种基于强化学习的多模态虚拟现实闭环认知自适应心理评估方法及系统,以解决现有心理评估系统动态适应能力不足、缺乏长期优化机制以及环境交互固定化的问题
[0061](1)实现实时动态认知感知;
Smart Images

Figure CN122805265A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the fields of artificial intelligence, virtual reality, human-computer interaction, reinforcement learning, and psychological assessment, and particularly to a virtual reality closed-loop cognitive adaptive psychological assessment method and system based on real-time multimodal behavior perception, dynamic cognitive state modeling, and reinforcement learning strategy optimization. Background Technology
[0002] Existing psychological assessment systems mainly rely on static psychological scales, fixed rules, or traditional interview methods to complete psychological state analysis.
[0003] Traditional systems typically only perform data acquisition, status analysis, and result output.
[0004] This type of system has the following problems:
[0005] (1) It lacks real-time dynamic perception capabilities and cannot continuously analyze changes in the user's cognitive state;
[0006] (2) It lacks environmental adaptability and cannot dynamically adjust the interactive environment according to the user's real-time status;
[0007] (3) It lacks a long-term feedback optimization mechanism and cannot form a continuous cognitive adaptation system;
[0008] (4) It is difficult to establish a long-term personalized cognitive behavior model.
[0009] At the same time, existing rule-based virtual reality psychological systems typically rely on fixed thresholds or preset scene logic, making it difficult to handle complex and continuous changes in cognitive states.
[0010] Therefore, there is an urgent need for an intelligent psychological assessment method and system that can perceive the user's cognitive state in real time, dynamically adjust the virtual environment, and form a long-term closed-loop optimization mechanism. Summary of the Invention
[0011] Purpose of the invention
[0012] The purpose of this invention is to provide a multimodal virtual reality closed-loop cognitive adaptive psychological assessment method and system based on reinforcement learning, so as to solve the problems of insufficient dynamic adaptability, lack of long-term optimization mechanism and fixed environmental interaction in existing psychological assessment systems.
[0013] This invention achieves real-time dynamic assessment and continuous cognitive adaptation optimization of users' psychological state through multimodal behavior perception, dynamic cognitive state modeling, reinforcement learning strategy optimization, virtual environment adaptive adjustment, and long-term feedback loop.
[0014] Technical solution
[0015] To achieve the above objectives, the present invention adopts the following technical solution:
[0016] A multimodal virtual reality closed-loop cognitive adaptive psychological assessment method based on reinforcement learning includes the following steps:
[0017] S1: Construct a virtual reality interactive environment and acquire multimodal behavioral data of users in the virtual reality environment;
[0018] S2: Perform real-time preprocessing, denoising, event segmentation, and feature extraction on the multimodal behavioral data;
[0019] S3: Attention features extracted based on the multimodal behavioral data;
[0020] Emotional characteristics, behavioral interaction characteristics,
[0021] Cognitive load characteristics and historical adaptation context characteristics
[0022] Constructing dynamic cognitive state vectors;
[0023] S4: Input the dynamic cognitive state vector into the reinforcement learning policy network to generate an environment-adaptive adjustment policy;
[0024] S5: Dynamically adjust the virtual reality environment parameters according to the environmental adaptive adjustment strategy;
[0025] S6: Continuously acquire user behavior feedback data in the adjusted environment and form a two-way dynamic feedback loop between users and the environment;
[0026] S7: Periodic strategy optimization based on long-term interaction data;
[0027] S8: Output dynamic psychological assessment report.
[0028] Furthermore, the multimodal behavioral data includes at least one of eye movement trajectory, gaze duration, saccade behavior, pupil changes, head movements, behavioral interaction frequency, voice response characteristics, and emotional change information.
[0029] Furthermore, step S2 includes smoothing the eye movement coordinates based on the unscented Kalman filter algorithm and classifying fixation events and saccade events based on the I-DT algorithm.
[0030] Furthermore, the dynamic cognitive state vector includes:
[0031] Attention characteristics;
[0032] Emotional characteristics;
[0033] Behavioral interaction characteristics;
[0034] Cognitive load characteristics;
[0035] Historical adaptation to contextual features.
[0036] Furthermore, the historical adaptation context features include at least one of the following: historical scenario response results, environmental adaptation history, emotional change trends, and intervention effect feedback.
[0037] Furthermore, the reinforcement learning policy network in step S4 is constructed based on the PPO reinforcement learning framework, using dynamic cognitive state vectors as state inputs and user behavior feedback and evaluation targets as reward signals to generate an environment-adaptive adjustment policy.
[0038] Furthermore, the virtual reality environment parameters mentioned in step S5 include:
[0039] Stimulus intensity parameters;
[0040] Interaction difficulty parameters;
[0041] Emotional environment parameters;
[0042] Social stress parameters;
[0043] Intervention strategy parameters.
[0044] Furthermore, the stimulus intensity parameters include at least one of light intensity, sound intensity, environmental complexity, and NPC proximity.
[0045] Furthermore, the periodic policy optimization in step S7 includes:
[0046] Maintain the stability of the main strategy network;
[0047] Continuously accumulate user interaction data;
[0048] Incremental strategy optimization is performed based on long-term behavioral feedback.
[0049] This invention also provides a multimodal virtual reality closed-loop cognitive adaptive psychological assessment system based on reinforcement learning, comprising:
[0050] The data acquisition module is used to acquire multimodal behavioral data of users in the virtual reality environment;
[0051] The data processing module is used to denoise, segment events, and extract features from behavioral data;
[0052] The cognitive state modeling module is used to construct dynamic cognitive state vectors;
[0053] The reinforcement learning decision-making module is used to generate adaptive adjustment strategies for the environment.
[0054] The environment dynamic adjustment module is used to dynamically adjust the parameters of the virtual reality environment;
[0055] The feedback optimization module is used to develop a long-term, continuous optimization mechanism based on user feedback.
[0056] The assessment report generation module is used to output dynamic psychological assessment reports.
[0057] Furthermore, the cognitive state modeling module adopts a temporal modeling structure based on LSTM and Self-Attention.
[0058] Furthermore, the reinforcement learning decision-making module adopts a dynamic policy optimization mechanism based on the reinforcement learning policy gradient.
[0059] Beneficial effects
[0060] Compared with the prior art, the present invention has at least the following beneficial effects:
[0061] (1) Realize real-time dynamic cognitive perception;
[0062] (2) Achieve dynamic coupling between user cognitive state and virtual environment parameters;
[0063] (3) Form a long-term closed-loop cognitive adaptation optimization mechanism;
[0064] (4) Improve the continuity and personalization of psychological assessment;
[0065] (5) Improve environmental interaction adaptability and long-term behavior analysis capabilities. Attached Figure Description
[0066] Figure 1 This is a diagram of the overall system architecture of the present invention;
[0067] Figure 2 This is a flowchart of the closed-loop cognitive adaptation method of the present invention;
[0068] Figure 3 This is a schematic diagram of the closed-loop cognitive adaptation feedback mechanism of the present invention;
[0069] Figure 4 This is a schematic diagram of the dynamic environment adjustment driven by reinforcement learning in this invention. Detailed Implementation
[0070] The present invention will now be described in further detail with reference to the accompanying drawings and specific embodiments.
[0071] The system of this invention runs on virtual reality devices that support eye tracking.
[0072] After a user enters a virtual reality scene, the system collects data in real time on the user's eye movement trajectory, gaze behavior, pupil changes, interaction behavior, head movement, voice response, and emotional changes.
[0073] The system first performs noise reduction and smoothing on the raw eye-tracking data based on the unscented Kalman filter algorithm.
[0074] Subsequently, the system uses the I-DT algorithm to identify gaze events and saccade events.
[0075] Subsequently, the system constructs a dynamic cognitive state vector through a temporal modeling network.
[0076] The cognitive state vector includes:
[0077] Attention stability;
[0078] Degree of emotional fluctuation;
[0079] Cognitive load level;
[0080] Behavioral interaction patterns;
[0081] History adapts to context.
[0082] The reinforcement learning strategy network generates a dynamic adjustment strategy based on the cognitive state vector.
[0083] Environmental adjustments include:
[0084] Adjust the scene complexity;
[0085] Adjust the number of NPCs;
[0086] Adjust the intensity of environmental stimuli;
[0087] Adjust the difficulty of the interaction;
[0088] Adjust the frequency of intervention prompts.
[0089] The system continuously acquires user feedback and forms a long-term closed-loop optimization mechanism.
[0090] During long-term operation, the system continuously accumulates historical user behavior data and periodically retrains and incrementally optimizes the strategies.
[0091] Finally, the system outputs a dynamic psychological assessment report.
[0092] The assessment report includes:
[0093] Overview of psychological state;
[0094] Behavioral pattern analysis;
[0095] Dynamic trend analysis;
[0096] Intervention effectiveness evaluation;
[0097] Personalized suggestions.
[0098] This invention is not limited to the above-described embodiments. Various modifications and substitutions that can be made by those skilled in the art without departing from the core idea of this invention should be included within the protection scope of this invention.
Claims
1. A multimodal virtual reality closed-loop cognitive adaptive psychological assessment method based on reinforcement learning, characterized in that, Includes the following steps: S1: Construct a virtual reality interactive environment and acquire multimodal behavioral data of users in the virtual reality environment; S2: Perform real-time preprocessing, denoising, event segmentation, and feature extraction on the multimodal behavioral data; S3: Construct a dynamic cognitive state vector based on attention features, emotion features, behavioral interaction features, cognitive load features, and historical adaptation context features; S4: Input the dynamic cognitive state vector into the reinforcement learning policy network to generate an environment-adaptive adjustment policy; S5: Dynamically adjust the virtual reality environment parameters according to the environmental adaptive adjustment strategy; S6: Continuously acquire user behavior feedback data in the adjusted environment and form a two-way dynamic feedback loop between the user and the environment; S7: Periodic strategy optimization based on long-term interaction data; S8: Output dynamic psychological assessment report.
2. The method according to claim 1, characterized in that, The multimodal behavioral data includes at least one of the following: eye movement trajectory, fixation duration, saccade behavior, pupil changes, head movements, frequency of behavioral interactions, voice response characteristics, and emotional change information.
3. The method according to claim 1, characterized in that, Step S2 includes smoothing eye movement coordinates based on the unscented Kalman filter algorithm and classifying fixation events and saccade events based on the I-DT algorithm.
4. The method according to claim 1, characterized in that, The dynamic cognitive state vector includes: Attention characteristics; Emotional characteristics; Behavioral interaction characteristics; Cognitive load characteristics; Historical adaptation to contextual features.
5. The method according to claim 4, characterized in that, The historical adaptation context features include at least one of the following: historical scenario response results, environmental adaptation history, emotional change trends, and intervention effect feedback.
6. The method according to claim 1, characterized in that, In step S4, a policy network based on the PPO reinforcement learning framework is used to generate an environment-adaptive adjustment policy according to the dynamic cognitive state vector.
7. The method according to claim 1, characterized in that, The virtual reality environment parameters mentioned in step S5 include: Stimulus intensity parameters; Interaction difficulty parameters; Emotional environment parameters; Social stress parameters; Intervention strategy parameters.
8. The method according to claim 7, characterized in that, The stimulus intensity parameters include at least one of light intensity, sound intensity, environmental complexity, and NPC proximity.
9. The method according to claim 1, characterized in that, The periodic policy optimization in step S7 includes: Maintain the stability of the main strategy network; Continuously accumulate user interaction data; Incremental strategy optimization is performed based on long-term behavioral feedback.
10. A multimodal virtual reality closed-loop cognitive adaptive psychological assessment system based on reinforcement learning, characterized in that, include: The data acquisition module is used to acquire multimodal behavioral data of users in the virtual reality environment; The data processing module is used to denoise, segment events, and extract features from behavioral data; The cognitive state modeling module is used to construct dynamic cognitive state vectors; The reinforcement learning decision-making module is used to generate adaptive adjustment strategies for the environment. The environment dynamic adjustment module is used to dynamically adjust the parameters of the virtual reality environment; The feedback optimization module is used to develop a long-term, continuous optimization mechanism based on user feedback. The assessment report generation module is used to output dynamic psychological assessment reports; The modules are connected through a data communication interface to form a closed-loop cognitive adaptive psychological assessment architecture.
11. The system according to claim 10, characterized in that, The cognitive state modeling module adopts a temporal modeling structure based on LSTM and self-attention.
12. The system according to claim 10, characterized in that, The reinforcement learning decision-making module employs a dynamic policy optimization mechanism based on the reinforcement learning policy gradient.