Secrecy education system based on multi-modal interaction
By collecting user data through multimodal interaction technology and combining deep learning algorithms with virtual simulation to simulate confidentiality scenarios, teaching strategies are dynamically generated. This solves the problems of low user participation and inaccurate feedback in existing confidentiality education systems, and achieves personalized confidentiality education results.
Patent Information
- Application Number
- CN202510913313.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-03
- Publication Date
- 2025-11-11
- Estimated Expiration
- 2045-07-03
AI Technical Summary
Existing confidentiality education systems suffer from insufficient user participation, inaccurate teaching feedback, and low efficiency in knowledge internalization due to one-way information transmission and a single interaction dimension.
The confidentiality education system employs multimodal interaction. It collects users' eye movement, voice commands, operational behaviors, and physiological signal data through multimodal sensors. Combining graph convolutional neural networks and deep reinforcement learning algorithms, it dynamically generates teaching strategies and simulates confidentiality scenarios through virtual simulation, providing real-time feedback to users' operations and forming a closed-loop feedback chain.
It significantly improves user engagement and knowledge internalization efficiency, enables personalized interactive training, enhances users' awareness of confidentiality and risk response capabilities in complex scenarios, and systematically solves the problems of single interaction and one-sided evaluation in traditional systems.
Smart Images

Figure CN120430913B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of data processing technology, and in particular to a confidentiality education system based on multimodal interaction. Background Technology
[0002] The confidentiality education system is an educational tool that uses information technology to enhance users' information security awareness and confidentiality skills. Its core functions are based on data fusion and interactive learning mechanisms. Traditional confidentiality education often adopts a one-way content transmission model, relying on static text, video resources, and standardized tests, making it difficult to dynamically capture users' learning status and environmental interaction characteristics. Modern systems integrate multimodal interaction technologies to simultaneously collect multi-dimensional behavioral data such as user visual attention, voice input, touch operation, and physiological feedback. Combined with machine learning algorithms, it analyzes cognitive characteristics and knowledge gaps, dynamically generating adaptive teaching strategies. The system constructs highly realistic confidentiality scenarios using virtual simulation technology, simulating information leakage risk situations. It guides users to complete decision-making training and behavioral correction through interactive operations. Simultaneously, it constructs a knowledge mastery assessment model based on the fusion of multi-source heterogeneous information, and optimizes the learning path with a real-time feedback mechanism, ultimately improving the efficiency of confidentiality knowledge internalization and the ability to cope with risks in complex scenarios.
[0003] Existing confidentiality education systems mostly employ a one-way information transmission model, primarily relying on learning management systems that integrate multimedia content libraries and online examination functions. They deliver confidentiality knowledge unidirectionally through pre-set text, image, and video resources. These systems typically depend on users selecting courses and completing standardized tests. Their interaction is limited to static content display and fixed-process question-and-answer feedback, lacking dynamic perception and real-time response to user behavior, cognitive state, and environmental information. Due to the singular human-computer interaction dimension, the system struggles to provide adaptive teaching strategies based on individual learning characteristics, resulting in insufficient user participation and limited knowledge internalization efficiency. Furthermore, existing confidentiality education systems rely heavily on single indicators such as exam scores for evaluating learning outcomes, failing to integrate multimodal behavioral data to construct a comprehensive evaluation model, thus limiting the accuracy and personalization of teaching feedback. Summary of the Invention
[0004] To address the shortcomings of existing technologies, this invention provides a confidentiality education system based on multimodal interaction, which solves the problems of insufficient user participation, inaccurate teaching feedback, and low efficiency of knowledge internalization caused by the one-way information transmission and single interaction dimension of existing confidentiality education systems.
[0005] To solve the above-mentioned technical problems, the specific technical solution of the present invention is as follows:
[0006] The confidentiality education system based on multimodal interaction provided by the present invention includes: a data acquisition module, a data processing module, a training module, and an optimization module;
[0007] The data acquisition module is configured to acquire multi-source heterogeneous data through a multi-modal sensor unit. The multi-source heterogeneous data includes the user's eye movement trajectory data, voice command semantic data, operation behavior coordinate data, and physiological signal data. The multi-source heterogeneous data is processed to generate a time-aligned multi-modal data stream through time synchronization processing. A three-dimensional data cube is constructed through a cross-modal feature fusion algorithm and transmitted to the data processing module.
[0008] The data processing module is configured to extract user cognitive features based on the three-dimensional data cube, generate teaching strategy constraints by matching knowledge graph nodes through a multi-objective optimization algorithm, and dynamically generate a teaching instruction set by combining a deep reinforcement learning algorithm and transmit it to the training module.
[0009] The training module is configured to load a confidential scenario model according to the teaching instruction set, adjust the scenario interaction logic based on dynamic disturbance parameters, receive user multimodal operation data and generate compliance path comparison results, trigger real-time feedback and write the interaction data into a distributed database and push it to the optimization module.
[0010] The optimization module is configured to integrate user operation response data, physiological feedback data, and test data to construct a capability evaluation model. It generates a set of policy parameter instructions through a reverse optimization algorithm and updates the reinforcement learning parameters of the data processing module in reverse, forming a closed-loop feedback chain.
[0011] Furthermore, in the confidentiality education system based on multimodal interaction described in this invention, the data acquisition module is further configured to: perform noise reduction processing on the multimodal data stream using an adaptive Kalman filter algorithm to generate filtered multimodal data; input the filtered multimodal data into a tensor decomposition algorithm for spatiotemporal feature fusion to construct a three-dimensional data cube including time dimension, spatial dimension and feature dimension, and transmit the three-dimensional data cube to the data processing module.
[0012] Furthermore, in the confidentiality education system based on multimodal interaction described in this invention, the data processing module is further configured to: extract user attention distribution heatmap, knowledge comprehension probability, and decision tendency vector from the three-dimensional data cube using a graph convolutional neural network, and generate a cognitive feature vector;
[0013] The cognitive feature vectors are matched with knowledge graph nodes using a multi-objective optimization algorithm to generate a Pareto optimal solution set that includes teaching difficulty, cognitive load, and knowledge coverage, thereby constraining the scope of the teaching strategy exploration of the deep reinforcement learning algorithm.
[0014] Furthermore, in the confidentiality education system based on multimodal interaction described in this invention, the training module is further configured to: optimize the dynamic perturbation parameters of the physical engine based on the simulated annealing algorithm, and adjust the file permission change frequency in file processing scenarios and the network attack gradient in network attack and defense scenarios in real time.
[0015] Based on the optimized dynamic disturbance parameters, a preset compliant path is generated through a real-time dynamic path planning algorithm, and the user operation path is compared with the preset compliant path to generate compliant path comparison data.
[0016] Based on the compliant path comparison data, a differentiated feedback animation is triggered, and the user operation path and comparison data are written into a distributed time-series database.
[0017] Furthermore, in the confidentiality education system based on multimodal interaction described in this invention, the optimization module is further configured to: generate multi-dimensional evaluation indicators by fusing response delay data, operation compliance data, physiological stress level data, and test score data in the user operation path using the analytic hierarchy process; input the multi-dimensional evaluation indicators into a Gaussian process regression model constructed by a Bayesian optimization algorithm to quantify the deviation between the user's ability curve and the target ability curve; and generate a policy parameter instruction set by reversely optimizing the association strength of the knowledge graph nodes and the dynamic perturbation parameters using a genetic algorithm, and then transmit the policy parameter instruction set back to the data processing module to update the reward function weights of the preset deep deterministic policy gradient algorithm.
[0018] Furthermore, the confidentiality education system based on multimodal interaction described in this invention also includes a cross-module data collaboration unit, configured to: aggregate the multimodal data streams collected by the data acquisition module through a federated learning framework, and perform distributed collaborative training with the data processing module to generate a global cognitive feature model;
[0019] A differential privacy protection algorithm is used to perform noise injection and desensitization processing on the user interaction data in the multimodal data stream;
[0020] The ant colony optimization algorithm dynamically allocates the computing resources of the microservice instances of the training and optimization modules to adapt to real-time load changes in the system.
[0021] Furthermore, in the confidentiality education system based on multimodal interaction described in this invention, the data processing module is further configured to: input the generated Pareto optimal solution set into a preset depth deterministic strategy gradient algorithm, and dynamically adjust the virtual scene trigger threshold and knowledge point reinforcement sequence based on the real-time user operation data fed back by the training module.
[0022] The adjusted teaching instruction set is transmitted to the training module via a message queue, driving the physics engine to update the file permission change frequency and network attack gradient parameters.
[0023] Furthermore, in the multimodal interaction-based confidentiality education system of the present invention, the training module is further configured to: load virtual simulation models of file processing scenarios and network attack and defense scenarios optimized based on simulated annealing algorithm through a 3D modeling engine;
[0024] The behavior tree model drives the scene NPC decision-making logic, and generates comparison data with the preset compliance path based on user operation data.
[0025] A differentiated feedback animation matching the comparison data is triggered by a real-time rendering engine, and the comparison data is written into a distributed time-series database.
[0026] Furthermore, in the confidentiality education system based on multimodal interaction described in this invention, the optimization module is further configured to: generate multi-dimensional evaluation indicators by fusing response delay data, operation compliance data, physiological stress level data, and test score data in the user operation path generated by the analytic hierarchy process.
[0027] The multi-dimensional evaluation indicators are input into the genetic algorithm to inversely optimize the knowledge graph node association strength and scene perturbation parameters, and generate a set of strategy parameter instructions.
[0028] The policy parameter instruction set is transmitted back to the data processing module through the microservice interface to update the policy parameters of the preset deep deterministic policy gradient algorithm.
[0029] Furthermore, in the confidentiality education system based on multimodal interaction described in this invention, the cross-module data collaboration unit is further configured to: aggregate multimodal data streams from multiple user terminals through a federated learning framework, construct a global cognitive feature model, and synchronize the global cognitive feature model to the data processing module;
[0030] A differential privacy protection algorithm is used to inject noise into the user behavior data in the global cognitive feature model to generate desensitized global model data;
[0031] The microservice instance computing resource allocation strategy of the training module is dynamically adjusted using the ant colony optimization algorithm to respond to changes in the real-time concurrent request volume of the system.
[0032] Beneficial effects of this invention;
[0033] This invention significantly enhances user engagement and knowledge internalization efficiency through multimodal data fusion and dynamic strategy generation mechanisms. The data acquisition module integrates visual tracking, speech recognition, and physiological signal detection units to capture multi-dimensional user interaction behaviors in real time. Adaptive filtering and tensor decomposition are used to construct a spatiotemporally correlated three-dimensional data cube, solving the problem of missing dynamic perception caused by the unidirectional data acquisition in traditional systems. The cognitive feature modeling module extracts user attention distribution and decision-making tendency features based on graph convolutional neural networks. Combined with Pareto optimal solution set-constrained deep reinforcement learning strategy generation, it dynamically adapts to user cognitive load and knowledge coverage requirements, realizing the transformation from fixed teaching processes to personalized interactive training. The virtual simulation module optimizes scene perturbation parameters and compliant path planning through simulated annealing algorithms, combined with a differentiated feedback animation triggering mechanism, enhancing user operational perception and risk response capabilities. The optimization module uses the analytic hierarchy process (AHP) to integrate multi-dimensional evaluation indicators and iteratively updates strategy parameters through a genetic algorithm, forming a closed-loop feedback chain to accurately improve the adaptability of teaching strategies. The cross-module collaborative unit incorporates a federated learning framework and differential privacy protection technology. While aggregating global cognitive feature models, it safeguards user privacy. Combined with ant colony optimization algorithms, it dynamically allocates computing resources to adapt to real-time load fluctuations, balancing system robustness and response efficiency. These technical solutions work synergistically to systematically address the shortcomings of traditional confidentiality education systems, such as simplistic interaction, delayed feedback, and one-sided evaluation, effectively strengthening users' confidentiality awareness and behavioral capabilities in complex scenarios. Attached Figure Description
[0034] To more clearly illustrate the technical solution of the present invention, the drawings used in the embodiments will be briefly introduced below. Obviously, for those skilled in the art, other drawings can be obtained based on the drawings without creative effort.
[0035] Figure 1 This is a system architecture diagram of a confidentiality education system based on multimodal interaction provided in an embodiment of the present invention. Detailed Implementation
[0036] To make the objectives, technical solutions, and advantages of this invention clearer, the technical solutions of this invention will be clearly and completely described below in conjunction with specific embodiments and corresponding drawings. Obviously, the described embodiments are only a part of the embodiments of this invention, and not all of them. Based on the embodiments of this invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this invention. The technical solutions provided by various embodiments of this invention will be described in detail below with reference to the accompanying drawings. To better understand the objectives of this invention, it will be described in further detail below.
[0037] Please see Figure 1The confidentiality education system based on multimodal interaction provided by the present invention includes: a data acquisition module, a data processing module, a training module, and an optimization module;
[0038] The data acquisition module is configured to acquire multi-source heterogeneous data through a multi-modal sensor unit. The multi-source heterogeneous data includes the user's eye movement trajectory data, voice command semantic data, operation behavior coordinate data, and physiological signal data. The multi-source heterogeneous data is processed to generate a time-aligned multi-modal data stream through time synchronization processing. A three-dimensional data cube is constructed through a cross-modal feature fusion algorithm and transmitted to the data processing module.
[0039] The data acquisition module collects user eye-tracking trajectory data, voice command semantic data, operation behavior coordinate data, and physiological signal data through a multimodal sensor unit. The multimodal sensor unit integrates a visual tracking device, a voice acquisition device, a touch interface, and a biosignal sensor to capture the user's attention distribution, semantic input, operation trajectory, and physiological response characteristics in different interaction scenarios. The visual tracking device captures eye movement trajectories based on an infrared light source and an image sensor, generating an eye-tracking data stream containing gaze coordinates, saccade paths, and dwell times. The voice acquisition device collects voice commands through a microphone array, combines this with a speech recognition engine to parse semantic content, and generates structured text data. The touch interface captures the spatial coordinates of user operations through capacitive or optical sensors, recording the time sequence and spatial distribution of touch operations. The biosignal sensor uses non-invasive electrodes to collect electrocardiogram, electrodermal, or electromyographic signals, quantifying the user's physiological stress level in confidential scenarios. Each sensor generates a timestamp based on a unified clock source, achieving initial synchronization of multi-source data through hardware trigger signals or software timestamp alignment mechanisms.
[0040] Multi-source heterogeneous data undergoes time synchronization processing to generate a time-aligned multimodal data stream. The time synchronization process employs network time protocols or hardware clock-based synchronization mechanisms to eliminate time deviations caused by differences in sampling frequencies or transmission delays between sensors. For data segments with significant timestamp discrepancies, interpolation algorithms compensate for missing data, generating continuous time series. The synchronized multimodal data stream is then segmented by time windows, forming parallel data channels containing eye-tracking trajectories, voice commands, operational behaviors, and physiological signals. Data segments within each time window are mapped to a unified time reference through calibration parameters, preserving the temporal correlation between data and providing time-consistent input for subsequent feature fusion.
[0041] A cross-modal feature fusion algorithm performs spatiotemporal correlation modeling on time-aligned multimodal data streams, constructing a 3D data cube. The algorithm first extracts the spatial distribution features of gaze points from eye-tracking data and combines this with the semantic vectors of voice commands to generate an attention-semantic correlation matrix. Operational behavior coordinate data is converted into continuous trajectory paths using a spatial interpolation algorithm and dynamically aligned with the time series of physiological signals. A tensor decomposition algorithm unfolds the multimodal data across time, space, and feature dimensions, extracting cross-modal latent feature vectors through higher-order singular value decomposition. The time dimension characterizes the dynamic characteristics of user behavior evolving with the teaching scenario, the spatial dimension reflects the spatial distribution patterns of data from different sensors, and the feature dimension integrates the coupling relationships between visual, semantic, operational, and physiological signals. The fused 3D data cube is transmitted to the data processing module via a distributed message queue as input data for cognitive feature modeling. The structural design of the 3D data cube supports multidimensional feature extraction by subsequent graph convolutional neural networks, providing a spatiotemporally correlated data foundation for the dynamic generation of teaching strategies.
[0042] The data processing module is configured to extract user cognitive features based on the three-dimensional data cube, generate teaching strategy constraints by matching knowledge graph nodes through a multi-objective optimization algorithm, and dynamically generate a teaching instruction set by combining a deep reinforcement learning algorithm and transmit it to the training module.
[0043] The data processing module extracts user cognitive features based on a three-dimensional data cube. The three-dimensional data cube contains multimodal fusion data across time, space, and feature dimensions. A graph convolutional neural network (Graph Convolutional Neural Network) models the spatiotemporal correlation features. Based on the topological structure of the data cube, the Graph Convolutional Neural Network constructs a spatiotemporal graph model. In the time dimension, it analyzes the shifting patterns of user attention focus; in the spatial dimension, it correlates the physical distribution of visual fixation points, touch operation coordinates, and physiological signal detection areas; and in the feature dimension, it analyzes speech semantic weights, operational compliance markers, and physiological stress levels. The network outputs a heatmap of user attention distribution, quantifying fixation point dwell time and visual focus switching frequency; it generates a knowledge comprehension probability vector reflecting the user's mastery of confidential knowledge points; and it extracts a decision-making tendency vector, representing the user's operational preferences in simulated risk scenarios. These feature vectors collectively constitute a digital representation of the user's cognitive state, providing input data for generating teaching strategies.
[0044] The multi-objective optimization algorithm matches cognitive feature vectors with knowledge graph nodes to generate teaching strategy constraints. Knowledge graph nodes are pre-defined with logical connections to confidential knowledge points, teaching difficulty labels, and cognitive load parameters. A semantic embedding model maps these nodes to a vector space. The algorithm aims to maximize knowledge coverage, minimize cognitive load, and adapt to the teaching difficulty, establishing an objective function under multiple constraints. The algorithm performs similarity matching between cognitive feature vectors and knowledge graph node vectors, generating a Pareto-optimal solution set containing multiple non-dominated solutions. Each solution corresponds to a set of teaching strategy parameters, balancing the dynamic relationship between knowledge point reinforcement intensity and user learning pressure. The Pareto-optimal solution set serves as boundary condition input to the deep reinforcement learning algorithm, limiting the exploration range of the policy network in the action space and preventing the generation of teaching strategies beyond the user's cognitive capabilities.
[0045] A deep reinforcement learning algorithm dynamically generates a teaching instruction set by combining real-time operation data. Based on a deep deterministic policy gradient framework, the algorithm uses the Pareto optimal solution set as the initial policy constraint and real-time user operation data as the environmental state input. Real-time operation data includes path deviation, response latency, and the proportion of compliant operations. The policy network evaluates the suitability of the virtual scene trigger threshold and the knowledge point reinforcement sequence. The policy network adjusts the action value function based on real-time feedback, updates the sensitivity coefficient of the scene trigger threshold and the knowledge point priority weights, and generates a teaching instruction set dynamically matched to the user's cognitive load. The adjusted instruction set is transmitted to the training module via a publish-subscribe message queue, driving the parameter updates of the virtual scene interaction logic. The message queue enables asynchronous communication between modules. The instruction set includes a file permission change frequency adjustment coefficient, a network attack gradient update step size, and a knowledge point reinforcement sequence index, which the physics engine parses and synchronously updates the scene interaction rules.
[0046] The training module is configured to load a confidential scenario model according to the teaching instruction set, adjust the scenario interaction logic based on dynamic disturbance parameters, receive user multimodal operation data and generate compliance path comparison results, trigger real-time feedback and write the interaction data into a distributed database and push it to the optimization module.
[0047] The training module loads confidentiality scenario models based on the teaching instruction set. This instruction set includes file permission change frequency adjustment coefficients, network attack gradient update step sizes, and knowledge point reinforcement sequence indexes. Through the parsing module, it drives the 3D modeling engine to load corresponding virtual simulation models of file processing and network attack / defense scenarios. The file processing scenario constructs a multi-level permission management system, simulating compliant processes for file access, editing, and sharing operations. The network attack / defense scenario pre-configures vulnerability exploitation paths and defense strategy libraries, dynamically generating network attack behavior sequences. During scenario model loading, the physics engine initializes dynamic perturbation rules based on instruction set parameters, establishes permission change trigger conditions and attack gradient calculation logic, forming an interactive environment adapted to the current teaching strategy.
[0048] The dynamic perturbation parameters adjust the scene interaction logic. The simulated annealing algorithm optimizes the physics engine parameters based on the teaching instruction set, iteratively calculating the optimal combination of the file permission change frequency threshold and the network attack gradient step size. A random perturbation factor is introduced during the parameter adjustment process to balance the stability and challenge of the scene interaction. The optimized dynamic perturbation parameters update the permission change triggering conditions in real time, controlling the frequency of permission upgrade requests in the file operation process; simultaneously, the vulnerability exploitation frequency and attack path complexity of the network attack scenario are adjusted to simulate the dynamic threat evolution in a real-world risk environment. After the parameters are updated, the scene interaction rules and user operation feedback mechanism are refreshed synchronously, forming a teaching environment that dynamically adapts to the user's cognitive abilities.
[0049] User multimodal operation data is input into the training module via a touch interface, voice acquisition unit, and eye-tracking device. Touch operation coordinate data is mapped to the virtual scene spatial coordinate system to generate the spatiotemporal trajectory of the user's operation path; voice commands are converted into operation commands by a semantic parsing engine, driving state changes of scene objects; eye-tracking trajectory data extracts visual focus distribution through a gaze point clustering algorithm to identify the user's attention concentration area. After timestamp alignment, the multimodal operation data is input into a behavior tree model to drive the scene's NPC decision-making logic. The behavior tree pre-sets compliance operation judgment nodes and risk response rules, triggering decision branches such as permission overreach interception, attack path blocking, or compliance operation guidance based on the spatiotemporal coordinates and operation type of the user's operation path, generating operation compliance markers and path deviation indicators.
[0050] The compliance path comparison results are generated based on a real-time dynamic path planning algorithm. The algorithm constructs an accessibility graph of authorized operations and a network defense path topology in a virtual scenario based on optimized dynamic perturbation parameters. A fast-expanding random tree algorithm searches for the optimal compliance path under the current parameter constraints, generating a preset path sequence containing key operation nodes and time windows. The user's operation path is compared with the preset path in spatiotemporal trajectory, calculating path deviation distance, operation delay time, and key node matching degree, generating multi-dimensional quantitative compliance data. The comparison results reflect the degree of matching between user operation behavior and teaching objectives, providing a quantitative basis for differentiated feedback.
[0051] The real-time feedback mechanism triggers differentiated animations based on compliance path comparison data. The real-time rendering engine calls upon a pre-built animation resource library, matching feedback types based on path deviation levels. Operations exceeding permission limits trigger dynamic warning borders and high-frequency warning sound effects, while operations matching compliance paths generate progressive guiding light effects and positive incentive prompts. The animation rendering process integrates real-time parameters from the physics engine, adjusting light and shadow intensity and interface element response speed to enhance the immersive experience for users. The triggering logic of feedback animations is embedded in the scene interaction loop, achieving millisecond-level synchronization between user actions and visual feedback, improving the real-time nature and interactivity of the teaching process.
[0052] Interaction data is written to a distributed time-series database and pushed to the optimization module. The spatiotemporal coordinates of user operation paths, compliance markers, and feedback log data are timestamped and then sharded according to scenario type and user ID. The distributed time-series database uses a columnar storage structure, partitioning data by time window, supporting high-concurrency writes and fast time-range queries. After data writing is complete, the message middleware triggers an asynchronous push event, transmitting standardized format interaction data packets to the optimization module's message queue via a publish-subscribe pattern. The data packets contain the original operation trajectory, compliance comparison indicators, and feedback trigger records, providing structured input for subsequent multi-dimensional evaluation and strategy iteration. The data transmission process uses a lightweight serialization protocol to reduce network bandwidth consumption and ensure system real-time requirements.
[0053] The optimization module is configured to integrate user operation response data, physiological feedback data, and test data to construct a capability evaluation model. It generates a set of policy parameter instructions through a reverse optimization algorithm and updates the reinforcement learning parameters of the data processing module in reverse, forming a closed-loop feedback chain.
[0054] The strategy parameter instruction set is transmitted back to the data processing module via a microservice interface. The interface uses a RESTful protocol to encapsulate the instruction data and defines a standardized parameter update format. The data processing module parses the instruction set, updates the reward function weight parameters of the deep deterministic policy gradient algorithm, and strengthens the probability of generating teaching strategies that align with user cognitive characteristics. The updated reward function drives the policy network to re-evaluate the virtual scenario trigger threshold and the priority of the knowledge point reinforcement sequence, forming a dynamic adaptation mechanism for teaching strategies. This closed-loop feedback loop, through continuous parameter iteration and optimization, improves the accuracy of the confidential education system in tracking user capability evolution and the efficiency of strategy adjustment.
[0055] The confidentiality education system based on multimodal interaction provided by this invention achieves a closed-loop teaching process through the collaborative work of multiple modules. The data acquisition module first synchronously collects the user's eye-tracking trajectory data, voice command semantic data, operational behavior coordinate data, and physiological signal data through a multimodal sensor unit. These multi-source heterogeneous data are aligned in time series using timestamp synchronization technology to eliminate temporal discrepancies between sensors, generating a time-aligned multimodal data stream. Subsequently, a cross-modal feature fusion algorithm performs tensor decomposition on the time-aligned data stream, fusing features from different modalities in time, space, and feature dimensions to construct a three-dimensional data cube to represent the user's multi-dimensional interactive behavior characteristics. The three-dimensional data cube is transmitted to the data processing module as the basic input for subsequent cognitive feature modeling.
[0056] Based on the three-dimensional data cube, the data processing module uses a graph convolutional neural network to extract a heatmap of user attention distribution, knowledge comprehension probability, and decision-making tendency vector in the spatiotemporal dimensions. These spatiotemporal feature vectors are matched with knowledge graph nodes using a multi-objective optimization algorithm to generate a Pareto optimal solution set that includes teaching difficulty, cognitive load, and knowledge coverage. This Pareto optimal solution set serves as a constraint input to a preset deep deterministic policy gradient algorithm, dynamically adjusting the virtual scene trigger threshold and knowledge point reinforcement sequence to generate a teaching instruction set. This teaching instruction set is transmitted to the training module via a message queue to drive the updating of subsequent scene interaction logic.
[0057] The training module loads file processing and network attack / defense scenario models from the confidentiality scenario library according to the teaching instruction set. It optimizes the dynamic perturbation parameters of the physics engine using simulated annealing to adjust the frequency of file permission changes and network attack gradients in real time. User operation data triggered by the multimodal interface is driven by a behavior tree model to generate scenario NPC decision logic, generating comparison data between the user's operation path and a preset compliant path. A real-time dynamic path planning algorithm generates compliant paths based on the optimized dynamic perturbation parameters and triggers differentiated feedback animations that match the comparison data through a real-time rendering engine. The user operation path data and comparison data are written to a distributed time-series database and pushed to the optimization module for multi-dimensional evaluation.
[0058] The optimization module, based on the Analytic Hierarchy Process (AHP), integrates response latency data, operational compliance data, physiological stress level data, and test score data from the user's operation path to generate multi-dimensional evaluation indicators. These indicators are input into a Gaussian process regression model constructed using a Bayesian optimization algorithm to quantify the deviation between the user's ability curve and the target ability curve. A genetic algorithm then uses this deviation to inversely optimize the knowledge graph node association strength and scene perturbation parameters, generating a set of policy parameter instructions. This set of policy parameter instructions is transmitted back to the data processing module via a microservice interface to update the reward function weights of the deep deterministic policy gradient algorithm, forming a closed-loop optimization chain for the teaching strategy.
[0059] The cross-module data collaboration unit aggregates multimodal data streams through a federated learning framework, collaboratively training a global cognitive feature model while protecting user privacy. A differential privacy protection algorithm injects noise into the interactive data to desensitize it, ensuring data security. An ant colony optimization algorithm dynamically allocates computing resources for microservice instances in the training and optimization modules, responding to real-time system load changes. The global cognitive feature model is synchronized to the data processing module, enhancing the generalization ability of personalized teaching strategies. Data flow and parameter updates between modules form a self-iterative teaching system, enabling dynamic adaptation between user cognitive states and teaching strategies.
[0060] Specifically, in the confidentiality education system based on multimodal interaction described in this invention, the data acquisition module is further configured to: perform noise reduction processing on the multimodal data stream using an adaptive Kalman filter algorithm to generate filtered multimodal data; input the filtered multimodal data into a tensor decomposition algorithm for spatiotemporal feature fusion to construct a three-dimensional data cube including time dimension, spatial dimension and feature dimension, and transmit the three-dimensional data cube to the data processing module.
[0061] The data acquisition module synchronously collects user eye-tracking trajectory data, voice command semantic data, operational behavior coordinate data, and physiological signal data through a multimodal sensor unit. This multi-source heterogeneous data includes time-series information from different sensors and may contain sensor noise and time deviations. An adaptive Kalman filter algorithm is used to perform noise reduction processing on the multimodal data stream. The filter parameters are dynamically adjusted based on the noise statistical characteristics of each sensor to suppress signal drift and random interference, generating filtered multimodal data. The filtered data preserves the temporal synchronization and signal integrity of user behavior characteristics, providing a preprocessing basis for subsequent feature fusion.
[0062] The filtered multimodal data is input into a tensor decomposition algorithm for cross-modal spatiotemporal feature fusion. The algorithm first slices the time dimension, aligning the time windows of different modalities to eliminate temporal misalignments. In the spatial dimension, a spatial correlation matrix is established based on the physical distribution of sensors or the spatial coordinates of user operations, extracting spatial consistency features of the multimodal data. In the feature dimension, visual attention features of eye-tracking trajectories, semantic features of voice commands, operational pattern features of touch behavior, and stress response features of physiological signals are jointly extracted through higher-order singular value decomposition. The tensor decomposition algorithm maps the multimodal data to a unified three-dimensional data space, constructing a three-dimensional data cube including time, space, and feature dimensions to represent the spatiotemporal correlation of user behavior and the coupling relationship of multimodal features.
[0063] The three-dimensional data cube is transmitted to the data processing module as input for cognitive feature modeling. The temporal dimension of the data cube records dynamic changes in user behavior, the spatial dimension reflects the synergistic effect of multi-sensor data, and the feature dimension integrates cross-modal semantic and physiological information. Through structured data representation, multi-granular and multi-dimensional data support is provided for subsequent spatiotemporal feature extraction by graph convolutional neural networks, forming a complete preprocessing link from raw signal denoising to high-order feature fusion.
[0064] Specifically, in the confidentiality education system based on multimodal interaction described in this invention, the data processing module is further configured to: extract user attention distribution heatmap, knowledge comprehension probability, and decision tendency vector from the three-dimensional data cube using a graph convolutional neural network, and generate a cognitive feature vector;
[0065] The cognitive feature vectors are matched with knowledge graph nodes using a multi-objective optimization algorithm to generate a Pareto optimal solution set that includes teaching difficulty, cognitive load, and knowledge coverage, thereby constraining the scope of the teaching strategy exploration of the deep reinforcement learning algorithm.
[0066] The data processing module receives a three-dimensional data cube from the data acquisition module. This data cube includes multimodal fusion information across time, space, and feature dimensions. A graph convolutional neural network constructs a spatiotemporal graph model based on the topological structure of the data cube. In the time dimension, it captures the dynamic changes in user attention distribution; in the spatial dimension, it correlates the physical locations of different sensors with the user's operational trajectory; and in the feature dimension, it analyzes the visual focus of eye-tracking trajectories, the semantic weights of voice commands, the compliance characteristics of touch operations, and the stress levels of physiological signals. Through multi-layer graph convolution operations, a heatmap of user attention distribution is extracted, quantifying the frequency and duration of visual focus shifts; a knowledge comprehension probability is generated, representing the user's mastery of confidential knowledge points; and a decision-making tendency vector is output, reflecting the user's risk response preferences in simulated scenarios. The heatmap, probability, and vector together constitute a cognitive feature vector, serving as a digital representation of the user's cognitive state.
[0067] The cognitive feature vectors are input into a multi-objective optimization algorithm and dynamically matched with knowledge graph nodes. Knowledge graph nodes are pre-set with confidential knowledge points, teaching difficulty labels, and cognitive load parameters. A semantic embedding model maps the cognitive feature vectors to the knowledge graph vector space. The multi-objective optimization algorithm solves for a Pareto optimal solution set, constrained by maximizing knowledge coverage, minimizing cognitive load, and adapting to the teaching difficulty. This solution set includes multiple non-dominated solutions, each corresponding to a set of teaching strategy parameters used to balance the dynamic relationship between knowledge reinforcement intensity and user learning pressure. The Pareto optimal solution set serves as a boundary condition input to a deep reinforcement learning algorithm, constraining the policy network's exploration range in the action space and preventing excessive deviation from the user's actual cognitive ability. The deep reinforcement learning algorithm adjusts the virtual scene trigger threshold based on real-time interactive feedback, dynamically allocates the priority of knowledge point reinforcement sequences, and generates a teaching instruction set adapted to the user's cognitive characteristics.
[0068] Specifically, in the confidentiality education system based on multimodal interaction described in this invention, the training module is further configured to: optimize the dynamic perturbation parameters of the physical engine based on the simulated annealing algorithm, and adjust the file permission change frequency in file processing scenarios and the network attack gradient in network attack and defense scenarios in real time.
[0069] Based on the optimized dynamic disturbance parameters, a preset compliant path is generated through a real-time dynamic path planning algorithm, and the user operation path is compared with the preset compliant path to generate compliant path comparison data.
[0070] Based on the compliant path comparison data, a differentiated feedback animation is triggered, and the user operation path and comparison data are written into a distributed time-series database.
[0071] The training module receives a set of teaching instructions from the data processing module and optimizes the dynamic perturbation parameters of the physics engine based on the simulated annealing algorithm. The simulated annealing algorithm initializes the initial parameter combinations for file permission change frequency and network attack gradients. Through iterative optimization, random perturbations are introduced to calculate the stability index of the scene interaction logic under different parameter combinations. The algorithm evaluates the adaptability of the parameter combinations based on an energy function, probabilistically accepting deterioration to escape local optima and gradually converge to the globally optimal parameter set. The optimized dynamic perturbation parameters adjust the triggering frequency of permission changes in the file processing scenario in real time and control the gradient change intensity of attack behavior in the network attack and defense scenario, forming scene interaction rules that dynamically adapt to the user's cognitive abilities.
[0072] Based on optimized dynamic perturbation parameters, the real-time dynamic path planning algorithm generates preset compliant paths in a virtual scenario. The algorithm constructs a spatial reachability graph for compliant operations based on scene physical constraints and a confidentiality rule base, combined with permission change frequency and attack gradient parameters. It then searches for the optimal path using a fast-expanding random tree algorithm to generate preset compliant paths that match the current teaching strategy. User operation data triggered through a multimodal interface is captured in real time and compared with the preset paths in both time and space. Path deviation, operation delay time, and compliance indicators for key nodes are calculated to generate multi-dimensional compliant path comparison data. This comparison data reflects the degree of matching between user operation behavior and teaching objectives, providing a quantitative basis for differentiated feedback.
[0073] Compliance path comparison data is input into the real-time rendering engine, triggering differentiated feedback animations adapted to user actions. The feedback system matches pre-set animation templates based on the path deviation level, triggering red warning animations for permission violations and generating green guiding animations for compliant operation paths. During animation rendering, real-time parameters from the physics engine are integrated to dynamically adjust lighting effects and interface element response speed, enhancing user immersion and the immediacy of feedback. The spatiotemporal coordinate data of the user operation path, comparison data, and feedback logs are timestamped and written to a distributed time-series database, stored in data shards according to scene type and time window, providing structured input data for multi-dimensional evaluation of the optimization module.
[0074] Specifically, in the confidentiality education system based on multimodal interaction described in this invention, the optimization module is further configured to: generate multi-dimensional evaluation indicators by fusing response delay data, operation compliance data, physiological stress level data, and test score data from the user operation path using the analytic hierarchy process; input the multi-dimensional evaluation indicators into a Gaussian process regression model constructed by a Bayesian optimization algorithm to quantify the deviation between the user's ability curve and the target ability curve; and generate a strategy parameter instruction set by reversely optimizing the association strength of the knowledge graph nodes and the dynamic perturbation parameters using a genetic algorithm, and then transmit the strategy parameter instruction set back to the data processing module to update the reward function weights of the preset deep deterministic strategy gradient algorithm.
[0075] The optimization module receives user operation path data from the training module, including response latency, operation compliance markers, physiological stress level indicators, and test scores. A multi-dimensional evaluation index system is constructed using the analytic hierarchy process (AHP). The data is standardized and weighted through expert weight allocation and consistency checks. Response latency data is mapped to a user reaction speed index, operation compliance data is converted into a rule compliance score, physiological stress level data is correlated with cognitive stress level, and test scores quantify the degree of knowledge mastery. After normalization, these indicators generate a multi-dimensional evaluation vector representing the difference between the user's overall ability status and the target ability model.
[0076] The multi-dimensional evaluation vector is input into a Bayesian optimization algorithm to construct a Gaussian process regression model. The model takes the user's ability vector as input, maps it to a high-dimensional feature space through a kernel function, and fits the nonlinear relationship between the user's ability curve and a preset target curve. Based on the posterior probability distribution, the ability deviation is calculated to quantify the user's ability gaps in knowledge application, risk decision-making, and stress coping. This deviation serves as the optimization objective function, driving the subsequent inverse parameter optimization process.
[0077] The genetic algorithm initializes the population of knowledge graph node association strength and dynamic perturbation parameters based on the aforementioned ability deviation. The population is iteratively optimized through selection, crossover, and mutation operations. Individual fitness is calculated by incorporating the deviation prediction value from a Gaussian process regression model and the stability constraints of the teaching strategy. During optimization, the Pareto front solution set is retained, and the optimal combination of association strength adjustment coefficients and perturbation parameters is selected to generate a set of strategy parameter instructions. This instruction set is transmitted back to the data processing module via a microservice interface to update the reward function weight parameters of the deep deterministic policy gradient algorithm. This strengthens the probability of generating teaching strategies that align with user ability characteristics, suppresses strategy exploration directions that deviate from the cognitive load threshold, and forms a closed-loop optimization link between teaching strategies and evaluation feedback.
[0078] Specifically, the confidentiality education system based on multimodal interaction described in this invention also includes a cross-module data collaboration unit, which is configured to: aggregate the multimodal data streams collected by the data acquisition module through a federated learning framework, and perform distributed collaborative training with the data processing module to generate a global cognitive feature model;
[0079] A differential privacy protection algorithm is used to perform noise injection and desensitization processing on the user interaction data in the multimodal data stream;
[0080] The ant colony optimization algorithm dynamically allocates the computing resources of the microservice instances of the training and optimization modules to adapt to real-time load changes in the system.
[0081] The cross-module data collaboration unit aggregates multimodal data streams from the data acquisition module using a federated learning framework. These data streams include eye-tracking trajectories, voice commands, operational behaviors, and physiological signal data from different user terminals. The federated learning framework performs preliminary feature extraction and model training on the local device, uploading only model parameters to the central server to avoid raw data transmission. The central server aggregates multi-terminal model parameters and generates a global cognitive feature model through weighted averaging, which is then synchronously updated to the data processing module. This global model integrates multi-user behavioral patterns and cognitive features, enhancing the generalization ability of teaching strategies while protecting user data privacy.
[0082] The differential privacy algorithm desensitizes user interaction data in multimodal data streams. During the local model training phase of federated learning, the algorithm injects Gaussian noise to perturb sensitive information such as eye-track coordinates, voice command text vectors, and operation sequence sequences. The noise injection intensity is dynamically adjusted based on data sensitivity; highly sensitive fields are encrypted using a Laplacian algorithm, while less sensitive fields are processed with random masks. The desensitized data retains the spatial correlation and temporal continuity of multimodal features, meeting the data requirements for cross-module collaborative training and balancing privacy protection with model training accuracy.
[0083] The ant colony optimization algorithm dynamically allocates computing resources for microservice instances in the training and optimization modules. The algorithm initializes virtual resource nodes as ant colony path nodes, using the real-time concurrent request volume as a pheromone concentration indicator. By iteratively simulating the ant foraging path selection process, it optimizes the allocation ratio of CPU, memory, and bandwidth resources for microservice instances. During high-load periods, resources are prioritized for the real-time rendering and physics engines of the training module; during low-load periods, resources are allocated to the genetic algorithm and Bayesian optimization process of the optimization module. This dynamic allocation strategy adapts to real-time system load fluctuations, maintaining a balance between interactive response speed and evaluation computation efficiency in teaching scenarios, supporting the efficient operation of the closed-loop teaching system.
[0084] Specifically, in the confidentiality education system based on multimodal interaction described in this invention, the data processing module is further configured to: input the generated Pareto optimal solution set into a preset depth deterministic strategy gradient algorithm, and dynamically adjust the virtual scene trigger threshold and knowledge point reinforcement sequence based on the real-time user operation data fed back by the training module.
[0085] The adjusted teaching instruction set is transmitted to the training module via a message queue, driving the physics engine to update the file permission change frequency and network attack gradient parameters.
[0086] The data processing module receives a Pareto optimal solution set generated by a multi-objective optimization algorithm. This solution set includes non-dominated solutions for teaching difficulty, cognitive load, and knowledge coverage. A deep deterministic policy gradient algorithm initializes the exploration boundary of the policy network based on this solution set, using real-time user operation data as environmental state input to dynamically evaluate the suitability of the virtual scenario trigger threshold and the knowledge point reinforcement sequence. Real-time operation data includes the user's path deviation, response latency, and compliance operation ratio in the current scenario. The action value function is updated through online policy evaluation, adjusting the sensitivity coefficient of the scenario trigger threshold and the priority weights of the knowledge point reinforcement sequence to generate a teaching strategy dynamically matched to the user's cognitive load.
[0087] The adjusted teaching instruction set is asynchronously transmitted to the training module via a message queue, which employs a publish-subscribe pattern to decouple modules. The instruction set includes adjustment coefficients for file permission change frequency, update steps for network attack gradients, and an index list of knowledge point reinforcement sequences. The physics engine parses the instruction set, dynamically adjusting the trigger frequency of access control rules for file processing scenarios based on the permission change coefficients, and updating the intrusion complexity of network attack and defense scenarios based on the attack gradient step size. After the parameters are updated, the scenario interaction logic is synchronously refreshed, and the compliance verification rules and real-time feedback mechanisms for user operation paths are adapted accordingly, forming a dynamic synergy between teaching strategies and user behavior.
[0088] The parameter update process is embedded in the online training loop of the reinforcement learning policy network. Based on scene state changes and user operation data fed back from the physics engine, the policy network periodically feeds back policy gradients to the data processing module. Gradient data drives policy network weight updates, optimizing the generation efficiency and scene adaptation accuracy of subsequent teaching instruction sets, and constructing a closed-loop iterative chain from policy generation, parameter execution to feedback optimization.
[0089] Specifically, in the confidentiality education system based on multimodal interaction described in this invention, the training module is further configured to: load virtual simulation models of file processing scenarios and network attack and defense scenarios optimized based on simulated annealing algorithm through a 3D modeling engine;
[0090] The behavior tree model drives the scene NPC decision-making logic, and generates comparison data with the preset compliance path based on user operation data.
[0091] A differentiated feedback animation matching the comparison data is triggered by a real-time rendering engine, and the comparison data is written into a distributed time-series database.
[0092] The training module loads virtual simulation models of file processing and network attack / defense scenarios optimized by the simulated annealing algorithm through a 3D modeling engine. The dynamic perturbation parameters generated by the simulated annealing algorithm include file permission change frequency thresholds and network attack gradient adjustment coefficients, which are encoded into the physical engine configuration file of the scenario model. The 3D modeling engine parses the configuration file and constructs a virtual scenario including the permission hierarchy topology, network node vulnerability distribution, and attack path reachability graph. The dynamic perturbation parameters drive real-time changes in the scenario's interaction logic, simulating the risk evolution process in a real confidential environment.
[0093] User operation data is input into the virtual scene via a multimodal interface. The behavior tree model parses the data and drives the scene's NPC decision-making logic. The behavior tree's pre-defined decision nodes include rules for handling permission violations, attack path interception strategies, and compliant operation guidance logic. The NPC triggers the corresponding decision branch based on the user's current operation path's spatiotemporal coordinates and operation type. The decision logic outputs a comparison of the user's operation path and the pre-defined compliant path's spatiotemporal trajectory, with comparison dimensions including path deviation distance, key node operation delay time, and the number of risk triggers, generating multi-dimensional compliance quantification indicators.
[0094] Compliance metrics are input into the real-time rendering engine, triggering differentiated feedback animations that match user actions. The rendering engine calls pre-set animation resource libraries based on path deviation levels; out-of-bounds permissions trigger dynamic warning borders and sound effects; and compliant path matching operations generate progressive guiding light effects. The animation rendering process integrates real-time parameters from the physics engine, adjusting light and shadow intensity and the interaction response speed of interface elements to enhance the realism of user experience. The spatiotemporal coordinate data of user operation paths, compliance metrics, and feedback logs are timestamped and stored in a distributed time-series database, segmented by scene type and user ID, forming structured training data with time-series labels. This provides data support for multi-dimensional evaluation and strategy iteration of the optimization module.
[0095] Specifically, in the confidentiality education system based on multimodal interaction described in this invention, the optimization module is further configured to: generate multi-dimensional evaluation indicators by fusing response delay data, operation compliance data, physiological stress level data and test score data in the user operation path generated by the analytic hierarchy process.
[0096] The multi-dimensional evaluation indicators are input into the genetic algorithm to inversely optimize the knowledge graph node association strength and scene perturbation parameters, and generate a set of strategy parameter instructions.
[0097] The policy parameter instruction set is transmitted back to the data processing module through the microservice interface to update the policy parameters of the preset deep deterministic policy gradient algorithm.
[0098] The optimization module receives user operation path data from the training module, including response latency, operation compliance markers, physiological stress level indicators, and test scores. A multi-dimensional evaluation index system is constructed using the analytic hierarchy process (AHP). The data is standardized and weighted through expert weight allocation and consistency checks. Response latency data is mapped to a user reaction speed index, operation compliance data is converted into a rule compliance score, physiological stress level data is correlated with cognitive stress level, and test scores quantify the degree of knowledge mastery. After normalization, these indicators generate a multi-dimensional evaluation vector representing the difference between the user's overall ability status and the target ability model.
[0099] The multi-dimensional evaluation vector is input into the genetic algorithm to initialize the population of knowledge graph node association strength and dynamic perturbation parameters. The algorithm iteratively optimizes the population through roulette wheel selection, single-point crossover, and Gaussian mutation operations. When calculating individual fitness, it incorporates the bias prediction value of the Gaussian process regression model and the stability constraints of the teaching strategy. During optimization, the Pareto front solution set is preserved, and the optimal combination of association strength adjustment coefficients and perturbation parameters is selected to generate a policy parameter instruction set. This instruction set includes knowledge graph node weight update coefficients, file permission change frequency adjustment step sizes, and network attack gradient correction values, adapting to the user's real-time capability status.
[0100] The policy parameter instruction set is transmitted back to the data processing module via a microservice interface. The interface uses a RESTful protocol to encapsulate the instruction data, enabling asynchronous communication across modules. The data processing module parses the instruction set, updates the reward function weight parameters of the deep deterministic policy gradient algorithm, strengthens the probability of generating teaching strategies that align with user ability characteristics, and suppresses policy exploration directions that deviate from the cognitive load threshold. The updated reward function drives the policy network to re-evaluate the priority of the virtual scenario trigger threshold and the knowledge point reinforcement sequence, forming a closed-loop optimization link between teaching strategies and evaluation feedback, thus improving the dynamic adaptability of the confidential education system.
[0101] Specifically, in the confidentiality education system based on multimodal interaction described in this invention, the cross-module data collaboration unit is further configured to: aggregate multimodal data streams from multiple user terminals through a federated learning framework, construct a global cognitive feature model, and synchronize the global cognitive feature model to the data processing module;
[0102] A differential privacy protection algorithm is used to inject noise into the user behavior data in the global cognitive feature model to generate desensitized global model data;
[0103] The microservice instance computing resource allocation strategy of the training module is dynamically adjusted using the ant colony optimization algorithm to respond to changes in the real-time concurrent request volume of the system.
[0104] The cross-module data collaboration unit aggregates multimodal data streams from multiple user terminals using a federated learning framework. These data streams include heterogeneous data such as eye-tracking trajectories, voice commands, touch operations, and physiological signals. The federated learning framework performs preliminary feature extraction and model training on the local terminal, uploading only model parameters to the central server, thus avoiding the transmission of raw user data across devices. The central server performs weighted averaging and gradient aggregation on the model parameters uploaded from multiple terminals to generate a global cognitive feature model. This global model integrates behavioral patterns and cognitive feature distributions across user groups and is synchronized to the knowledge graph nodes of the data processing module, enhancing the generalization ability and scenario adaptability of personalized teaching strategies.
[0105] The differential privacy-preserving algorithm performs noise injection desensitization on user behavior data in the global cognitive feature model. During the parameter aggregation stage of federated learning, the algorithm introduces Laplace noise to perturb the feature vectors in the model parameters that are strongly correlated with user identity. High-sensitivity fields employ an adaptive noise injection strategy, dynamically adjusting the noise intensity based on the importance of the feature dimension, while low-sensitivity fields are obfuscated using random masks. The desensitized global model data retains the statistical distribution characteristics of multimodal features, eliminating individual user identity information, thus satisfying privacy requirements while maintaining the effectiveness of model training.
[0106] The ant colony optimization algorithm dynamically adjusts the resource allocation strategy for microservice instances in the training module. The algorithm abstracts the CPU, memory, and bandwidth resources of microservice instances as virtual path nodes, using the real-time concurrent request volume as an indicator of path pheromone concentration. By simulating the path-finding mechanism of an ant colony, it iteratively optimizes the resource allocation ratio. During high-concurrency periods, resources are prioritized for the real-time rendering engine and physics engine to ensure smooth interaction in the virtual scene; during low-load periods, resources are allocated to compliant paths for comparison computation and feedback animation generation, improving data processing throughput efficiency. This dynamic allocation strategy adapts to system load fluctuations, maintaining a resource balance between teaching, training, and evaluation optimization.
[0107] The key technical features involved in the technical solution of this invention are explained as follows:
[0108] A multimodal sensor unit refers to a composite sensor system that integrates visual tracking, voice recognition, touch interaction, and physiological signal detection. It is used to simultaneously collect multi-source heterogeneous data such as user eye movement trajectory, voice commands, operation behavior coordinates, and heart rate variability.
[0109] The adaptive Kalman filter algorithm is a dynamic noise cancellation technique that suppresses sensor signal drift and random interference by adjusting the filter parameters in real time, while preserving the timing synchronization of data.
[0110] Tensor decomposition algorithms are used for cross-modal feature fusion. By jointly modeling time, space and feature dimensions, a three-dimensional data cube representing the spatiotemporal correlation of user behavior is constructed.
[0111] Graph convolutional neural networks are spatiotemporal feature extraction models that extract user attention distribution heatmaps, knowledge comprehension probability, and decision-making tendency vectors based on the topological structure of a three-dimensional data cube, generating cognitive feature vectors.
[0112] The Pareto optimal solution set refers to the non-dominated solution set output by a multi-objective optimization algorithm, which balances teaching difficulty, cognitive load, and knowledge coverage, and constrains the policy generation boundary.
[0113] The deep deterministic policy gradient algorithm is a reinforcement learning policy generation method that dynamically adjusts the virtual scene trigger threshold and knowledge point reinforcement sequence by combining real-time operation data to generate a teaching instruction set.
[0114] Simulated annealing algorithm is used to optimize the dynamic perturbation parameters of the physics engine. It breaks out of local optima by probabilistically accepting deterioration and generates file permission change frequency and network attack gradient that are adapted to user capabilities.
[0115] The federated learning framework enables cross-terminal collaborative training of data. Local devices extract features and upload model parameters, while the central server aggregates and generates a global cognitive feature model, protecting user privacy.
[0116] Differential privacy protection algorithms desensitize sensitive user behavior data by injecting Laplace noise into the model parameters.
[0117] Ant colony optimization algorithm dynamically allocates microservice resources, using concurrent request volume as a pheromone indicator to optimize the allocation ratio of CPU, memory and bandwidth to adapt to the real-time load of the system.
[0118] These features work together to achieve multimodal data fusion, dynamic strategy generation, closed-loop feedback optimization, and privacy protection, systematically solving the technical defects of traditional systems such as one-way interaction, delayed feedback, and one-sided evaluation, and improving the efficiency of internalizing confidential knowledge and the ability to cope with complex scenarios.
[0119] Adaptive Kalman filter algorithm: used for real-time noise cancellation of multimodal data streams. It dynamically adjusts the filtering parameters according to the statistical characteristics of sensor noise, suppresses signal drift and random interference, and improves the data signal-to-noise ratio.
[0120] Tensor decomposition algorithm: performs cross-dimensional fusion of time-aligned multimodal data, and constructs a three-dimensional data cube including time, space and feature dimensions through high-order singular value decomposition to represent the spatiotemporal correlation of user behavior.
[0121] Graph Convolutional Neural Networks: Based on the topological structure of a three-dimensional data cube, it extracts spatiotemporal features and outputs attention distribution heatmaps, knowledge comprehension probability, and decision tendency vectors to quantify the user's cognitive state.
[0122] Multi-objective optimization algorithm (NSGA-II): With the constraints of maximizing knowledge coverage and minimizing cognitive load, it generates a Pareto optimal solution set, defines the feasible domain of teaching strategies, and balances user learning pressure and knowledge reinforcement intensity.
[0123] Deep Deterministic Policy Gradient Algorithm (DDPG): Combining real-time interactive data with Pareto optimal solution sets, it dynamically adjusts the virtual scene trigger threshold and knowledge point reinforcement sequence to generate a teaching instruction set.
[0124] Simulated annealing algorithm: Optimizes dynamic perturbation parameters of virtual scenarios (such as file permission change frequency and network attack gradient), avoids local optima by probabilistically accepting degraded solutions, and generates highly realistic risk scenarios.
[0125] Real-time Dynamic Path Planning (RRT): Generates a preset compliant path based on physics engine parameters, compares it with the user-operated path, outputs path deviation and compliance indicators, and supports a feedback mechanism.
[0126] Bayesian optimization algorithm: Construct a Gaussian process regression model to quantify the deviation between the user's capability curve and the target curve, and identify the weak links in knowledge application and risk decision-making.
[0127] Genetic Algorithm: It reverse-optimizes the node association strength and scene perturbation parameters of the knowledge graph, and generates a set of strategy parameter instructions through selection, crossover and mutation operations to drive the iterative update of teaching strategies.
[0128] Federated learning framework: Enables collaborative training of data from multiple terminals. Local devices extract features and only upload model parameters, while the central server aggregates and generates a global cognitive feature model, avoiding leakage of raw data.
[0129] Differential privacy protection algorithm: Injecting Laplace noise during the parameter aggregation stage of federated learning obfuscates sensitive features of user behavior data and meets privacy protection requirements.
[0130] Ant colony optimization algorithm: Dynamically allocates computing resources for microservice instances, using concurrent request volume as a pheromone indicator to optimize the allocation ratio of CPU, memory and bandwidth, and adapt to real-time load fluctuations of the system.
[0131] The confidentiality education system based on multimodal interaction provided by this invention achieves a closed-loop teaching process through the collaborative work of multiple modules. The data acquisition module synchronously collects user eye-tracking trajectory data, voice command semantic data, operational behavior coordinate data, and physiological signal data through multimodal sensor units. These multi-source heterogeneous data are aligned with time series using timestamp synchronization technology to eliminate temporal discrepancies between sensors and generate a time-aligned multimodal data stream. An adaptive Kalman filter algorithm performs noise cancellation processing on the multimodal data stream, dynamically adjusting filter parameters to suppress signal drift and random interference. The filtered data is input into a tensor decomposition algorithm, which performs cross-modal feature fusion in time, space, and feature dimensions to construct a three-dimensional data cube, representing the spatiotemporal correlation and multimodal feature coupling relationship of user behavior. The three-dimensional data cube is transmitted to the data processing module as the basic input for cognitive feature modeling.
[0132] The data processing module, based on a 3D data cube, utilizes a graph convolutional neural network to extract a heatmap of user attention distribution, knowledge comprehension probability, and decision-making tendency vectors across the spatiotemporal dimensions. A multi-objective optimization algorithm matches the cognitive feature vectors with knowledge graph nodes, generating a Pareto-optimal solution set that includes teaching difficulty, cognitive load, and knowledge coverage. This solution set is input into a pre-defined deep deterministic policy gradient algorithm, which dynamically adjusts the virtual scene trigger threshold and knowledge point reinforcement sequence based on real-time user operation data, generating a teaching instruction set. This instruction set is transmitted to the training module via a message queue, driving updates to the scene interaction logic.
[0133] The training module loads a confidential scenario model based on the teaching instruction set, optimizes the dynamic perturbation parameters of the physics engine using simulated annealing, and adjusts the frequency of file permission changes in file processing scenarios and the network attack gradient in network attack and defense scenarios in real time. User operation data triggered by the multimodal interface is driven by a behavior tree model to generate scenario NPC decision logic, producing comparison data between the user's operation path and the preset compliant path. A real-time dynamic path planning algorithm generates compliant paths based on the optimized parameters, and a real-time rendering engine triggers differentiated feedback animations to enhance the immersive experience and immediacy of user feedback. User operation path data and comparison data are written to a distributed time-series database and pushed to the optimization module for multi-dimensional evaluation.
[0134] The optimization module employs the Analytic Hierarchy Process (AHP) to integrate response latency, compliance, physiological stress, and test data from the user's operation path, generating multi-dimensional evaluation metrics. A Bayesian optimization algorithm constructs a Gaussian process regression model to quantify the deviation between the user's capability curve and the target curve. A genetic algorithm, based on this deviation, inversely optimizes the knowledge graph node association strength and scenario perturbation parameters, generating a policy parameter instruction set. This instruction set is transmitted back to the data processing module via a microservice interface, updating the reward function weights of the deep deterministic policy gradient algorithm, forming a closed-loop iterative chain from policy generation and execution to feedback optimization.
[0135] The cross-module data collaboration unit aggregates data streams from multiple terminals using a federated learning framework. Feature extraction and model training are performed on local devices, while only model parameters are uploaded to the central server to generate a globally recognized feature model, which is then synchronized to the data processing module to enhance policy generalization capabilities. A differential privacy-preserving algorithm injects Laplacian noise during the parameter aggregation stage to desensitize sensitive feature vectors, balancing privacy protection and model accuracy. An ant colony optimization algorithm dynamically allocates microservice resources for the training module, prioritizing real-time rendering and the physics engine during high-concurrency periods, and allocating resources to the evaluation computation process during low-load periods, adapting to real-time system load fluctuations.
[0136] Each module overcomes the limitations of traditional systems' one-way transmission and single interaction through multimodal data fusion, dynamic strategy generation, and closed-loop evaluation and optimization. Users enhance their ability to apply confidentiality knowledge through highly realistic scenario training in multimodal interaction. The system dynamically adjusts teaching strategies based on real-time feedback, enabling rapid internalization of risk response capabilities in complex scenarios. The proposed solution, through data collaboration and privacy protection mechanisms, ensures data security while improving user engagement and feedback accuracy, meeting the intelligent and personalized needs of modern confidentiality education systems.
[0137] This invention addresses the issue of insufficient user engagement in existing systems through multimodal data fusion and dynamic strategy generation mechanisms. The data acquisition module integrates visual tracking, speech recognition, touch interaction, and physiological signal detection units to capture multi-dimensional data in real time, including user eye movement trajectories, operational behaviors, and physiological feedback. A spatiotemporally correlated three-dimensional data cube is constructed using a tensor decomposition algorithm. The cognitive feature modeling module extracts user attention distribution, knowledge comprehension, and decision-making tendency features based on graph convolutional neural networks. Combined with a multi-objective optimization algorithm, it generates a Pareto optimal solution set adapted to the user's cognitive state, driving the dynamic adjustment of virtual scene trigger thresholds and knowledge point reinforcement sequences. This transforms the teaching model from one-way instruction to multimodal interaction, enhancing the depth of user engagement.
[0138] To address the issue of inaccurate teaching feedback, the system improves feedback accuracy through a closed-loop evaluation and reverse optimization mechanism. The virtual simulation module optimizes dynamic perturbation parameters based on simulated annealing, generates compliant path comparison data matching real-time teaching strategies, and triggers differentiated feedback animations by driving NPC decision-making logic through a behavior tree model. The optimization module uses the analytic hierarchy process (AHP) to integrate user operation responses, physiological stress, and test data to construct multi-dimensional evaluation indicators. Through Bayesian optimization and a genetic algorithm, it reverse-adjusts the knowledge graph node association strength and scene perturbation parameters, generating a policy parameter instruction set to update the reinforcement learning model, thus forming a precise feedback chain from user behavior collection and multi-dimensional evaluation to policy iteration.
[0139] To improve knowledge internalization efficiency, the system constructs a cross-module collaboration and dynamic resource allocation system. The federated learning framework aggregates multi-terminal data streams to generate a global cognitive feature model, and combines this with a differential privacy protection algorithm to achieve data anonymization, enhancing the generalization ability of teaching strategies. The ant colony optimization algorithm dynamically allocates microservice resources, adapting to real-time load changes and ensuring a balance between virtual scene interaction response speed and evaluation computation efficiency. The deep deterministic policy gradient algorithm continuously optimizes the reward function weights based on closed-loop feedback data, dynamically balancing knowledge coverage and cognitive load constraints. Through a teaching instruction set, it drives users to complete knowledge application training in highly realistic scenarios, achieving rapid internalization of risk response capabilities in complex scenarios.
Claims
1. A confidentiality education system based on multimodal interaction, characterized in that: include: Data acquisition module, data processing module, training module, and optimization module; The data acquisition module is configured to acquire multi-source heterogeneous data through a multi-modal sensor unit. The multi-source heterogeneous data includes the user's eye movement trajectory data, voice command semantic data, operation behavior coordinate data, and physiological signal data. The multi-source heterogeneous data is processed to generate a time-aligned multi-modal data stream through time synchronization processing. A three-dimensional data cube is constructed through a cross-modal feature fusion algorithm and transmitted to the data processing module. The three-dimensional data cube includes a time dimension, a spatial dimension, and a feature dimension. The data processing module is configured to extract user cognitive features based on the three-dimensional data cube, generate teaching strategy constraints by matching knowledge graph nodes through a multi-objective optimization algorithm, and dynamically generate a teaching instruction set by combining a deep reinforcement learning algorithm and transmit it to the training module. The training module is configured to load a confidential scenario model according to the teaching instruction set, adjust the scenario interaction logic based on dynamic disturbance parameters, receive user multimodal operation data and generate compliant path comparison results. The compliant path comparison results include a spatiotemporal trajectory comparison between the user operation path and a preset path, calculation of path deviation distance, operation delay time and key node matching degree, triggering real-time feedback and writing the interaction data into a distributed database and pushing it to the optimization module. The optimization module is configured to integrate user operation response data, physiological feedback data, and test data to construct a capability evaluation model, and to generate a policy parameter instruction set through a reverse optimization algorithm and update the reinforcement learning parameters of the data processing module in reverse.
2. The confidentiality education system based on multimodal interaction according to claim 1, characterized in that, The data acquisition module is further configured to: perform noise cancellation processing on the multimodal data stream using an adaptive Kalman filter algorithm to generate filtered multimodal data; input the filtered multimodal data into a tensor decomposition algorithm for spatiotemporal feature fusion to construct a three-dimensional data cube including time dimension, spatial dimension and feature dimension, and transmit the three-dimensional data cube to the data processing module.
3. The confidentiality education system based on multimodal interaction according to claim 2, characterized in that; The data processing module is further configured to: extract user attention distribution heatmap, knowledge comprehension probability, and decision tendency vector from the three-dimensional data cube using a graph convolutional neural network, and generate cognitive feature vector; The cognitive feature vectors are matched with knowledge graph nodes using a multi-objective optimization algorithm to generate a Pareto optimal solution set that includes teaching difficulty, cognitive load, and knowledge coverage, thereby constraining the scope of the teaching strategy exploration of the deep reinforcement learning algorithm.
4. The confidentiality education system based on multimodal interaction according to claim 3, characterized in that... ; The training module is also configured to: optimize the dynamic perturbation parameters of the physics engine based on the simulated annealing algorithm, and adjust the file permission change frequency in the file processing scenario and the network attack gradient in the network attack and defense scenario in real time. Based on the optimized dynamic disturbance parameters, a preset compliant path is generated through a real-time dynamic path planning algorithm, and the user's operation path is compared with the preset compliant path to generate compliant path comparison data. Based on the compliance path comparison data, a differentiated feedback animation is triggered, and the user operation path and comparison data are written into a distributed database.
5. The confidentiality education system based on multimodal interaction according to claim 4, characterized in that... ; The optimization module is further configured to: integrate response delay data, operational compliance data, physiological stress level data, and test score data from the user operation path based on the analytic hierarchy process (AHP) to generate multi-dimensional evaluation indicators; input the multi-dimensional evaluation indicators into a Gaussian process regression model constructed by a Bayesian optimization algorithm to quantify the deviation between the user's capability curve and the target capability curve; and reverse-optimize the association strength and dynamic perturbation parameters of the knowledge graph nodes through a genetic algorithm to generate a policy parameter instruction set, and transmit the policy parameter instruction set back to the data processing module to update the reward function weights of the deep reinforcement learning algorithm, wherein the deep reinforcement learning algorithm is a deep deterministic policy gradient algorithm.
6. The confidentiality education system based on multimodal interaction according to claim 5, characterized in that, It also includes a cross-module data collaboration unit, which is configured to aggregate multimodal data streams collected by the data acquisition module through a federated learning framework, and perform distributed collaborative training with the data processing module to generate a global cognitive feature model; A differential privacy protection algorithm is used to perform noise injection and desensitization processing on user interaction data in multimodal data streams; The ant colony optimization algorithm dynamically allocates computing resources for microservice instances in the training and optimization modules to adapt to real-time load changes in the system.
7. The confidentiality education system based on multimodal interaction according to claim 6, characterized in that; The data processing module is also configured to: input the generated Pareto optimal solution set into the deep deterministic policy gradient algorithm, and dynamically adjust the virtual scene trigger threshold and knowledge point reinforcement sequence based on the real-time user operation data fed back by the training module. The adjusted teaching instruction set is transmitted to the training module via a message queue, driving the physics engine to update the file permission change frequency and network attack gradient parameters.
8. The confidentiality education system based on multimodal interaction according to claim 7, characterized in that... ; The training module is also configured to load virtual simulation models of file processing and network attack and defense scenarios optimized based on the simulated annealing algorithm through a 3D modeling engine. The behavior tree model drives the scene NPC decision-making logic, and generates comparison data with the preset compliance path based on user operation data. The real-time rendering engine triggers differential feedback animations that match the comparison data, and the comparison data is written to a distributed database.
9. The confidentiality education system based on multimodal interaction according to claim 8, characterized in that... ; The optimization module is also configured to: generate multi-dimensional evaluation indicators by fusing response delay data, operation compliance data, physiological stress level data and test score data in the generated user operation path using the analytic hierarchy process (AHP). By inputting multi-dimensional evaluation metrics into the genetic algorithm, the association strength of knowledge graph nodes and scene perturbation parameters are optimized in reverse to generate a set of policy parameter instructions. The policy parameter instruction set is transmitted back to the data processing module through the microservice interface to update the policy parameters of the deep deterministic policy gradient algorithm.
Citation Information
Patent Citations
Digital management system based on VR technology
CN119762306A
Classroom learning dynamic evaluation and management system
CN119991377A