AR (Augmented Reality) system cognitive security interaction method based on multi-modal big language model reasoning and bidirectional individuation

By combining a personalized attention perception model and a multimodal large language model, the problem of accurately identifying user cognitive distraction and understanding risk preferences in dynamic scenes in AR systems is solved, realizing efficient and personalized safe interaction of AR systems and improving user experience and security.

CN121979387APending Publication Date: 2026-05-05UNIV OF ELECTRONICS SCI & TECH OF CHINA
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
UNIV OF ELECTRONICS SCI & TECH OF CHINA
Filing Date
2026-01-20
Publication Date
2026-05-05

AI Technical Summary

Technical Problem

Existing AR systems cannot accurately capture moments of cognitive distraction in dynamic scenes, nor can they understand users' subjective risk preferences, resulting in high false alarm rates, alarm fatigue, and damage to human-machine trust.

Method used

By constructing a personalized attention perception model, using IMU signals to eliminate brainwave motion artifacts, and combining it with a multimodal large language model to understand user risk preferences, personalized scene cognitive reasoning and adaptive alarms with human-machine alignment are achieved.

Benefits of technology

Accurately capture users' moments of distraction, reduce false alarm rates, avoid overreaction, enhance user trust, and achieve safe interaction that adapts to physiological characteristics and psychological preferences.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121979387A_ABST
    Figure CN121979387A_ABST
Patent Text Reader

Abstract

The invention discloses an AR (Augmented Reality) system cognitive security interaction method based on multi-modal large language model reasoning and bidirectional individuation, which utilizes a cross attention mechanism guided by an inertial measurement unit signal to explicitly model and eliminate electroencephalogram signal motion artifacts, and compared with the existing method of directly using electroencephalogram signals or simple filtering, the method has the advantages that the method is simple and convenient to implement, and the efficiency is high. The anti-interference capability of personalized attention perception is high, attention recognition is more accurate, the distraction moment of the user can be accurately captured, and the false alarm rate is reduced. Besides, personalized preference configuration information is generated through man-machine alignment and is embedded into the reasoning process of the multi-modal large language model, so that the subjective preference of the user to the risk is read, the problem of alarm fatigue is solved, and user trust is established. According to the method, a real agent closed loop is realized, and the method is a complete agent architecture with perception (physiological denoising)-cognition (LLM reasoning combined with preference)-action (self-adaptive interaction), and can adapt to physiological features and psychological preferences of different users.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of mobile augmented reality cognitive interaction technology, and more specifically, it relates to a cognitive safety interaction method for AR (Augmented Reality) systems based on multimodal large language model reasoning and bidirectional personalization. Background Technology

[0002] As mobile augmented reality (AR) devices (such as HoloLens and Apple Vision Pro) gradually integrate into daily life, it has become commonplace for users to use AR devices in dynamic scenarios such as walking and climbing. However, the virtual information presented by AR devices often preempts users' visual attention, leading to "inattentional blindness." This means that although users' eyes are on the real environment, they fail to process the risk information in the environment at the cognitive level, which can easily lead to safety accidents.

[0003] To address this issue, existing AR system security interactions have evolved from "passive defense" to "general intelligent assistance," with key technical solutions including:

[0004] (1) Passive defense based on explicit behavior triggers:

[0005] Early AR systems primarily relied on explicit user behavior patterns, such as gaze duration or head posture changes. For example, an AR system would only initiate an environmental scan when it detected that a user was gazing at virtual content for more than a certain threshold, such as one second. While these approaches were logically simple, they often exhibited lag and failed to predict the user's cognitive state.

[0006] (2) Assistive systems based on general physiological computing and computer vision:

[0007] More advanced solutions, such as AttentionAR, are beginning to incorporate multimodal perception technology. Its typical workflow is as follows:

[0008] Attention state monitoring: Physiological signals are collected through wearable devices, such as EEG headbands and IMU sensors, and general machine learning models, such as SVM or general BiLSTM, are used to determine whether the user is in a state of "focused on the outside world" or "cognitively distracted".

[0009] Environmental risk identification: Once "cognitive distraction" is detected, the AR system calls upon a common object detection model, such as YOLO or a visual language model (VLM / MLLM), to analyze the environmental image.

[0010] Alarm feedback: Based on preset objective rules, such as "an alarm will be triggered when an obstacle is detected", a prompt will be sent to the user.

[0011] This type of cognitive safety interaction scheme attempts to build a safety net by combining physiological perception (user awareness) and scene awareness, and it is currently the mainstream direction of technology.

[0012] Despite the introduction of multimodal perception in existing technologies, serious problems remain when facing complex mobile scenarios and highly personalized user needs, including a mismatch between model generality and individual specificity, and a misalignment between objective algorithms and subjective perception. Specifically:

[0013] 1. Physiological computation level: Unable to cope with the dual challenges of "individual heterogeneity" and "movement artifacts".

[0014] Individual differences leading to generalization failure: Human electroencephalogram (EEG) signals exhibit extremely high inter-subject variability. Different users may display drastically different neural activation patterns when performing the same cognitive task. Current technologies typically employ a "one-size-fits-all" generic model, meaning the model is trained on group data to serve all users. Experiments have shown that this generic model's accuracy drops significantly when faced with users whose physiological characteristics deviate from the "average person," resulting in an inability to accurately capture moments of user distraction.

[0015] Signal contamination from motion noise: In dynamic scenarios such as walking and climbing stairs, body movements generate strong electromyography (EMG) and motion artifacts. This noise often overlaps with EEG frequency bands (such as Theta and Alpha waves) that reflect attention. While existing technologies use simple filtering (such as bandpass filtering), they lack a **deep denoising mechanism** that integrates IMU motion data, making it difficult to separate "cognitive signals" from "motion noise," resulting in a high false alarm rate.

[0016] 2. Cognitive Interaction Level: The "Semantic Gap" Between Objective Risk Assessment and Subjective Psychological Thresholds

[0017] Misalignment in risk perception: Risk is a subjectively constructed concept, rather than a purely physical fact. Existing technologies, including general multimodal models, typically define risk levels (Low / Moderate / High) based on objective criteria such as object category and distance. However, users' risk preferences vary greatly: for the same scenario, such as someone 3 meters ahead, a conservative user might want to immediately call the police, while an aggressive user might consider it an irrelevant disturbance.

[0018] Cognitive biases in a state of distraction are overlooked: Current technologies fail to quantify the amplifying effect of cognitive distraction on risk perception. Research shows that when users are cognitively distracted, their sensitivity to risk undergoes a non-linear change (distraction amplification). General models cannot perceive this subtle psychological shift.

[0019] Alarm fatigue and trust crisis: Due to a lack of understanding of user intent and preferences (i.e., a lack of human-agent alignment), existing AR systems often exhibit either "overreaction" or "slow response." Frequent invalid alarms can lead to "alarm fatigue," causing users to ignore or even disable accessibility features; while missed alarms at critical moments can completely undermine human-machine trust, ultimately causing the secure interaction system to fail. Summary of the Invention

[0020] The purpose of this invention is to overcome the shortcomings of the prior art and provide a cognitive safety interaction method for AR systems based on multimodal large language model reasoning and bidirectional personalization. This method can not only understand whether the user is "cognitively distracted" through physiological signals and accurately capture the moment of the user's distraction, but also remove "motion noise" and reduce the false alarm rate. Furthermore, it can understand the user's "subjective preference" for risk through multimodal large model reasoning, thereby avoiding the breakdown of human-computer trust caused by "overreaction" or "slow reaction".

[0021] To achieve the above-mentioned objectives, this invention provides a cognitive safety interaction method for AR systems based on multimodal large-model reasoning and bidirectional personalization, characterized by the following steps:

[0022] (1) Construct and train a personalized attention perception model that is resistant to motion interference.

[0023] Using inertial measurement unit signals as a reference noise source, motion artifacts are precisely removed from EEG signals through a cross-attention mechanism, thereby training a personalized attention perception model that is resistant to motion interference specific to the user and judging in real time whether the user is in a state of focused environment or cognitive distraction.

[0024] (2) Constructing a human-computer aligned personalized scene cognition reasoning layer based on a multimodal large language model

[0025] The user's subjective risk preferences are captured in advance during the calibration phase to generate personalized preference configuration information. This information is then embedded into the prompts of the multimodal big language model as a core knowledge base. During cognitive reasoning based on egocentric perspective images captured by the AR system's real-time camera, the embedded personalized preference configuration information is forcibly invoked for comparison. This allows the multimodal big language model to understand the user's unique view of danger, guiding it to reason and generate the necessity of alarming, i.e., whether to alarm. This results in a human-machine aligned personalized scene cognitive reasoning layer.

[0026] (3) Implement adaptive personalized interventions

[0027] When the personalized attention perception model detects that the user is in a state of cognitive distraction during the use of the AR system, it triggers an environmental risk assessment: the human-machine aligned personalized scene cognitive reasoning layer performs cognitive reasoning based on the egocentric viewpoint image captured by the AR system's real-time camera and generates an alarm necessity, i.e., whether to alarm.

[0028] If an alarm is required, a personalized alarm with low interference and high perception will be presented in the AR system interface based on the user's preset interaction habits.

[0029] The objective of this invention is achieved as follows.

[0030] This invention, based on a multimodal large language model (MLLM) reasoning and a bidirectional personalized AR system cognitive safety interaction method, utilizes a cross-attention mechanism guided by inertial measurement unit (IMU) signals to explicitly model and eliminate motion artifacts in EEG signals. Compared to existing methods that directly use EEG signals or simple filtering, this method exhibits stronger anti-interference capabilities and more accurate attention recognition, precisely capturing users' distraction moments and reducing false alarm rates. Furthermore, this invention generates personalized preference configuration information through human-AI alignment and embeds it into the reasoning process of the MLLM to understand the user's "subjective preference" for risk, solving the problem of "alarm fatigue" and establishing user trust. This invention achieves a true agent closed loop, a complete agent architecture with "perception (physiological denoising) - cognition (LLM reasoning combined with preferences) - action (adaptive interaction)," capable of adapting to the physiological characteristics and psychological preferences of different users. Attached Figure Description

[0031] Figure 1 This is a flowchart of a specific implementation of the cognitive safety interaction method for AR systems based on multimodal large language model reasoning and bidirectional personalization of the present invention;

[0032] Figure 2This is a schematic diagram illustrating the principle of a specific implementation of the AR system cognitive safety interaction method based on multimodal large language model reasoning and bidirectional personalization of the present invention;

[0033] Figure 3 This is a schematic diagram of the motion-resistant personalized attention perception model structure constructed in the AR system cognitive safety interaction method based on multimodal large model reasoning and bidirectional personalization of the present invention. Detailed Implementation

[0034] The specific embodiments of the present invention will now be described with reference to the accompanying drawings to enable those skilled in the art to better understand the invention. It should be particularly noted that in the following description, detailed descriptions of known functions and designs that might obscure the main content of the invention will be omitted here.

[0035] I. Theoretical Basis: Theoretical basis for AR attention state definition and multimodal detection

[0036] To clarify the specific targets and feasibility of "distraction" monitoring in this invention, the internal / external attention and its physiological detection basis are explained as follows:

[0037] 1. Definitions and differences between internal and external attention

[0038] This invention, based on the classification system of cognitive psychology, mainly divides users' attention states into two categories:

[0039] • External Attention: This refers to a user focusing their cognitive resources on processing external sensory input. In the mobile AR scenario of this invention, typical external attention behaviors include "visual search" (such as finding road signs or observing pedestrians) or "reading AR information" (such as scanning virtual text). In this case, the user maintains a high level of awareness of environmental risks.

[0040] • Internal Attention: This refers to the user's shift from processing external sensory information to focusing on internal mental activities. In this invention, this is defined as a major source of "distraction." For example, when a user is performing complex mental arithmetic tasks or deep thinking while walking, even if their gaze is directed forward, their cognitive focus has already shifted, making them highly susceptible to "inattentive blindness."

[0041] 2. Theoretical basis for distinguishing between EEG and IMU

[0042] This invention utilizes electroencephalography (EEG) and inertial measurement unit (IMU) signals to distinguish between the two states mentioned above, and its scientific basis is as follows:

[0043] • Neurophysiological basis of EEG: Cognitive neuroscience research shows that the directionality of attention (internal vs. external) is closely related to the activity of the prefrontal cortex of the brain and significantly modulates the power of specific EEG frequency bands.

[0044] References support this study: Chun et al. proposed a taxonomy of internal and external attention in the *Annual Review of Psychology*; further research by Magosso et al. confirmed that when attention shifts from external visual stimuli to internal cognitive tasks, the power of alpha waves (8-13 Hz) and theta waves (4-8 Hz) in EEG signals undergoes significant modulation changes. This provides direct physiological evidence for this invention to identify cognitive distraction states using power changes at different EEG frequencies.

[0045] • Behavioral evidence from IMU: In addition to brain electrical activity, head movement patterns under different attention states also showed significant statistical differences.

[0046] Behavioral characteristics: When a user is in a state of "external attention" (such as searching the environment), it is usually accompanied by more active head rotation (such as a large change in the angular velocity of the yaw axis); while in a state of "internal attention" (such as contemplation), head movements tend to be static or exhibit unconscious micro-movements. This invention captures these subtle kinematic features through an IMU, not only as an independent classification criterion, but also as a key "reference signal" for removing motion artifacts in EEG (see step S1.3).

[0047] II. Technical Solution: A Cognitive and Safe Interaction Method for AR Systems Based on Multimodal Large Language Model Reasoning and Two-Way Personalization

[0048] This invention presents a closed-loop AR system safety interaction method based on multimodal large language model reasoning and bidirectional personalized approach, encompassing "physiological perception - cognitive alignment - adaptive action." This method aims to make the AR system act like an "intelligent co-pilot" that understands the user, not only interpreting physiological signals to determine if the user is cognitively distracted, but also using a multimodal large language model to understand the user's "subjective preference" for risk. The overall process comprises three core levels.

[0049] In this embodiment, as Figure 1 , 2 As shown, the present invention provides a cognitive safety interaction method for AR systems based on multimodal large language model reasoning and bidirectional personalization, comprising the following steps:

[0050] Step S1: Construct and train a motion-resistant personalized attention perception model

[0051] To address the pain point of high physiological signal noise in mobile scenarios, this paper utilizes inertial sensor (IMU) signals as a reference noise source and uses a cross-attention mechanism to accurately remove motion artifacts from electroencephalogram (EEG) signals. This allows for the training of a personalized attention perception model that is resistant to motion interference, enabling real-time judgment of whether the user is in a focused environment or a cognitively distracted state.

[0052] The core technical problem addressed in this step is how to extract pure cognitive features from physiological signals heavily contaminated by body movement during dynamic processes such as walking and climbing stairs. To this end, this embodiment designs a "dual-stream signal denoising and fusion network," which is a motion-resistant, personalized attention perception model specific to the user, guided by inertial measurement unit (IMU) signals and designed to facilitate cross-attention. Figure 3 As shown.

[0053] Step S1.1: Synchronous acquisition of multimodal heterogeneous signals

[0054] The user wears an AR display device and a physiological signal acquisition device. The AR system simultaneously acquires two signals with millisecond-level timestamps:

[0055] Target Signal: The user's electroencephalogram (EEG) signal, which can be a single-channel or multi-channel EEG signal. In this embodiment, the focus is on the Delta to Gamma frequency band, which reflects the brain's cognitive load and attentional state.

[0056] Reference Signal: This contains inertial measurement unit (IMU) data with nine characteristic components: triaxial acceleration, triaxial angular velocity, and triaxial Euler angles. It reflects the user's head and body motion and is the primary source of noise.

[0057] Step S1.2: Data preprocessing and time-frequency feature construction

[0058] The two acquired signals are preprocessed and feature-engineered to construct a feature sequence for model input.

[0059] Signal decomposition and filtering: For user EEG signals, bandpass filtering is performed to retain the effective frequency band of 0.5-45Hz, and the signals are decomposed into power characteristics of six frequency bands, namely the energy values ​​of Delta, Theta, Alpha, Low Beta, High Beta, and Gamma. For inertial measurement unit data containing nine characteristic components, a moving average filter is applied for smoothing to reduce random noise.

[0060] Sliding window feature extraction and normalization: The sliding window technique is used to segment the power characteristics of the six frequency bands and the smoothed inertial measurement unit data containing nine feature components. The window length is set to 10 seconds and the movement step is 1 second. Normalization is performed on the signal in each window to eliminate the influence of different signal dimensions. In this way, a set of EEG feature sequences of six frequency bands and an inertial measurement unit data feature sequence containing nine feature components are obtained every second.

[0061] Downsampling and Class Balance: To reduce data dimensionality and improve computational efficiency, the extracted EEG feature sequences from the six frequency bands and the inertial measurement unit (IMU) data feature sequences containing nine feature components were downsampled to 1 Hz. Furthermore, during the training phase of the personalized attention perception model, to address the potential issue of inconsistent sample numbers for different attention states (i.e., class imbalance), a class weighting strategy was introduced into the loss function to ensure the model's sensitivity to minority class samples.

[0062] The downsampled six-band EEG feature sequences constitute the EEG signal features, or EEG features. The downsampled inertial measurement unit (IMU) data feature sequence containing nine feature components constitutes the IMU signal feature, or IMU feature. .

[0063] Step S1.3: Motion artifact stripping based on cross-attention

[0064] To "clean" the EEG signal, we built a dynamic filter using the attention mechanism in deep learning. The core logic is to find components in the EEG signal that are highly similar to the IMU motion pattern and treat them as noise to be subtracted—that is, trajectory removal based on IMU features.

[0065] Encoding: A bidirectional long short-term memory (BiLSTM) network is used to encode the EEG features and IMU features respectively to capture the dynamic changes over time.

[0066] Role Assignment:

[0067] Query item (Query, Encoded EEG features , representing "the goal we want to purify".

[0068] Key item (Key, ) and value item (Value, ): Encoded IMU features , representing "known motion interference patterns".

[0069] Attention weight calculation:

[0070] Calculate the query items and key items dot product (Dot Product) analyzes the similarity between the EEG signal at each moment and the IMU motion signal. A higher similarity indicates that the change in the EEG signal at that moment is more likely caused by motion than by brain activity. The dot product is then used to analyze the similarity. As attention weights.

[0071] Motion artifact synthesis:

[0072] Based on the calculated similarity, i.e., the attention weight, from the value item... That is, a weighted synthesis of IMU features yields an estimated motion artifact. :

[0073]

[0074] Differential denoising:

[0075] Subtract this estimated motion artifact component from the EEG features. This yields pure EEG features that retain cognitive information but remove motion interference. :

[0076]

[0077] Step S1.4: Personalized Attention State Classification

[0078] The purified pure EEG characteristics Input a fully connected neural network classifier and output the probability of the user's current cognitive distraction state. When the probability is greater than a set threshold, the user is determined to be in a cognitive distraction state (i.e., an internal attention state, that is, immersed in thinking or virtual content and ignoring the external environment). If the probability is not greater than the set threshold, the user is in a focused environment state (i.e., an external attention state).

[0079] Step S2: Construct a human-computer aligned personalized scene cognition and reasoning layer based on a multimodal large language model

[0080] The user's subjective risk preferences are captured in advance during the calibration phase to generate personalized preference configuration information. This information is then embedded as a core knowledge base into the prompts of a multimodal large language model. During cognitive reasoning based on egocentric viewpoint images captured in real-time by an AR system's camera, the embedded personalized preference configuration information is forcibly invoked for comparison. This allows the multimodal large language model to understand the user's unique perspective on danger, guiding it to reason and generate the necessity of an alarm (i.e., whether to issue an alarm). This results in a human-machine aligned personalized scene cognitive reasoning layer. Specifically, this includes the following steps:

[0081] When the perception model sends a trigger signal indicating that the user is in a state of "cognitive distraction," the AR system activates the visual analysis module. The core of this step is to enable the general multimodal large language model (MLLM) to "understand" the specific user's unique perception of danger.

[0082] Step S2.1: Construct personalized preference configuration information for achieving human-machine cognitive alignment.

[0083] When the AR system is used for the first time or during periodic calibration, an interactive process is used to capture the user's subjective preferences and generate structured data.

[0084] Scene demonstration: The AR system shows users a series of typical video clips containing different types of risks (such as steps, pedestrians, and vehicles) and levels of risk.

[0085] General AI assessment: The AR system first displays a general large model's objective assessment of the scene (e.g., an obstacle is detected 3 meters ahead, risk level: medium).

[0086] User subjective correction: Users modify the assessment of general AI based on their own feelings. For example, a user may think, "An obstacle 3 meters away is not a risk to me, and I don't need to call the police," or "Although the risk is low, I am very afraid of dogs and must call the police."

[0087] Differentiated modeling: The AR system records the difference between general AI assessments and user subjective corrections, generating a personalized preference profile. This profile includes the user's list of sensitive risk sources (what must be reported), subjective risk thresholds (what level constitutes danger), and willingness to trigger an alarm.

[0088] Step S2.2: Construct a chain-of-thought prompting that includes preference constraints.

[0089] The egocentric viewpoint image captured by the camera is input into the multimodal large language model, and a prompt word containing instructions is constructed to guide the multimodal large language model to reason step by step:

[0090] Role setting: Command model: "You are not a general observer, you are a personalized security assistant specifically designed to serve this user";

[0091] Context injection: The personalized preference configuration information generated in step S2.1 is embedded into the prompt words as the core knowledge base;

[0092] Visual perception: The multimodal large language model first identifies objects, distances, and dynamic relationships in the image;

[0093] Preference matching and reasoning (key step): The multimodal large language model forces the call to the embedded preference configuration information for comparison; logical example: "Although the risk of this step is objectively low, according to the configuration information, the user is extremely sensitive to steps when reading, so I must increase the risk weight."

[0094] Decision output: Generate final conclusion: the necessity of alarm, i.e. whether to issue an alarm.

[0095] Step S3: Implement adaptive personalized intervention

[0096] When the personalized attention perception model detects that the user is in a state of cognitive distraction during the use of the AR system, it triggers an environmental risk assessment: the human-machine aligned personalized scene cognitive reasoning layer performs cognitive reasoning based on the egocentric viewpoint image captured by the AR system's real-time camera and generates an alarm necessity, i.e., whether to alarm.

[0097] If an alarm is required, a personalized alarm with low interference and high perceptibility will be presented in the AR system interface based on the user's preset interaction habits. The personalized alarm includes:

[0098] Step S3.1: Intent-driven decision execution:

[0099] The alarm decision-making of the AR system is no longer based on rigid objective rules (such as "alarm if risk level > 2"), but directly executes the alarm necessity conclusions inferred by a multimodal large language model. This mechanism effectively avoids invalid alarms like "crying wolf" and reduces disturbance to users.

[0100] Step S3.2: Adaptive Multimodal Presentation

[0101] Once an alarm is detected, the AR system will dynamically select the alarm method based on the user's interaction habits recorded in the personalized preference configuration information:

[0102] For visually sensitive users, a red flashing border or directional arrow is projected onto the AR glasses;

[0103] For users with hearing sensitivity, play specific prompts or voice commands;

[0104] For minimalist users, only a tiny icon is displayed.

[0105] This ensures that the alarm is perceptible to the user without causing additional fright or cognitive burden.

[0106] III. Technical Effects of the Invention (Advantages Compared with Existing Technologies)

[0107]

[0108] Table 1

[0109] 1. Strong anti-interference ability and more accurate attention recognition

[0110] Compared to existing methods that directly use raw EEG or simple filtering, this invention utilizes an IMU-guided cross-attention mechanism to explicitly model and eliminate motion artifacts. Table 1 compares the performance with existing methods. As can be seen from Table 1, in dynamic scenarios such as user walking, this method can significantly improve the accuracy of attention state recognition (significantly improving compared to general models and traditional methods such as SVM), effectively solving the problem of physiological signal noise in mobile AR.

[0111] 2. Address alarm fatigue and build user trust.

[0112] Unlike existing technologies that use fixed risk thresholds, this invention generates personalized preference records through human-AI alignment and incorporates them into the reasoning process of a large model-based learning (MLLM). This means that the agent can "think like a user." Experimental results show that this method significantly improves the accuracy of predicting "alarm necessity" in relation to the user's true intentions, and user subjective ratings show that their trust in the system and perceived safety are both superior to those of general models.

[0113] 3. Achieved a true closed-loop intelligent agent system.

[0114] The AR system in this invention is not a simple detector, but a complete intelligent agent architecture with "perception (physiological denoising) - cognition (LLM reasoning combined with preferences) - action (adaptive interaction)", which can adapt to the physiological characteristics and psychological preferences of different users.

[0115] Although the illustrative specific embodiments of the present invention have been described above to enable those skilled in the art to understand the invention, it should be understood that the invention is not limited to the scope of the specific embodiments. For those skilled in the art, various changes are obvious as long as they are within the spirit and scope of the invention as defined and determined by the appended claims, and all inventions utilizing the concept of the present invention are protected.

Claims

1. A cognitive safety interaction method for AR systems based on multimodal large-model reasoning and bidirectional personalization, characterized in that, Includes the following steps: (1) Construct and train a personalized attention perception model that is resistant to motion interference; Using inertial measurement unit signals as a reference noise source, motion artifacts are precisely removed from EEG signals through a cross-attention mechanism, thereby training a personalized attention perception model that is resistant to motion interference specific to the user and judging in real time whether the user is in a state of focused environment or cognitive distraction. (2) Construct a human-computer aligned personalized scene cognition reasoning layer based on a multimodal large language model; The user's subjective risk preferences are captured in advance during the calibration phase to generate personalized preference configuration information. This information is then embedded into the prompts of the multimodal big language model as a core knowledge base. During cognitive reasoning based on egocentric perspective images captured by the AR system's real-time camera, the embedded personalized preference configuration information is forcibly invoked for comparison. This allows the multimodal big language model to understand the user's unique view of danger, guiding it to reason and generate the necessity of alarming, i.e., whether to alarm. This results in a human-machine aligned personalized scene cognitive reasoning layer. (3) Implement adaptive, personalized interventions; When the personalized attention perception model detects that the user is in a state of cognitive distraction during the use of the AR system, it triggers an environmental risk assessment: the human-machine aligned personalized scene cognitive reasoning layer performs cognitive reasoning based on the egocentric viewpoint image captured by the AR system's real-time camera and generates an alarm necessity, i.e., whether to alarm. If an alarm is required, a personalized alarm with low interference and high perception will be presented in the AR system interface based on the user's preset interaction habits.

2. The AR system cognitive safety interaction method based on multimodal large model reasoning and bidirectional personalization according to claim 1, characterized in that, Step (1) utilizes the inertial measurement unit signal as a reference noise source and precisely removes motion artifacts from the EEG signal through a cross-attention mechanism, thereby training a personalized attention perception model specific to the user that resists motion interference. This model determines in real time whether the user is in a focused state or a cognitively distracted state. 1.1) Synchronous acquisition of multimodal heterogeneous signals; The AR system simultaneously acquires two signals at millisecond-level timestamps: the main signal is the user's EEG signal, and the reference signal is the inertial measurement unit data containing nine feature components, namely, three-axis acceleration, three-axis angular velocity, and three-axis Euler angles. 1.2) Data preprocessing and time-frequency feature construction; The two acquired signals are preprocessed and feature-engineered to construct a feature sequence for model input: Signal decomposition and filtering: For user EEG signals, bandpass filtering is performed to retain the effective frequency band of 0.5-45Hz, and it is decomposed into the power characteristics of six frequency bands, namely the energy values ​​of Delta, Theta, Alpha, Low Beta, High Beta, and Gamma. For inertial measurement unit data containing nine characteristic components, a moving average filter is applied for smoothing to reduce random noise. Sliding window feature extraction and normalization: The sliding window technique is used to segment the power characteristics of the six frequency bands and the smoothed inertial measurement unit data containing nine feature components. The window length is set to 10 seconds and the movement step is 1 second. Normalization processing is performed on the signal in each window. In this way, a set of EEG feature sequences of six frequency bands and an inertial measurement unit data feature sequence containing nine feature components are obtained every second. Downsampling and category balancing: The extracted EEG feature sequences of the six frequency bands and the inertial measurement unit data feature sequences containing nine feature components are all downsampled to 1 Hz. In addition, during the training phase of the personalized attention perception model, a category weighting strategy is introduced into the loss function. The downsampled six-band EEG feature sequences constitute the EEG signal features, or EEG features. The downsampled inertial measurement unit (IMU) data, containing nine characteristic components, constitutes the IMU signal characteristics. ; 1.3) Motion artifact stripping based on cross-attention; Encoding: A bidirectional long short-term memory network is used to encode EEG features and IMU features respectively to capture dynamic changes over time; Role Assignment: Query Item Encoded EEG features Key items Sum of values All are encoded IMU features ; Attention weight calculation: Calculate query terms and key items dot product ; Motion artifact synthesis: Based on the calculated similarity, i.e., the attention weight, from the value terms... That is, a weighted synthesis of IMU features yields an estimated motion artifact. : ; Differential denoising: Subtracting motion artifact components from EEG features This yields pure EEG features that retain cognitive information but remove motion interference. : ; 1.4) Personalized attention state classification; The purified pure EEG characteristics Input a fully connected neural network classifier and output the probability of the user's current cognitive distraction state. If the probability is greater than a set threshold, the user is determined to be in a cognitive distraction state; if it is not greater than the set threshold, the user is in a focused state.

3. The AR system cognitive safety interaction method based on multimodal large model reasoning and bidirectional personalization according to claim 1, characterized in that, Step (2) involves pre-capturing the user's subjective risk preferences during the calibration phase, generating personalized preference configuration information, and embedding this information as a core knowledge base into the prompts of the multimodal large language model. During cognitive reasoning based on egocentric perspective images captured by the AR system's real-time camera, the embedded personalized preference configuration information is forcibly invoked for comparison, allowing the multimodal large language model to understand the user's unique view of danger. This guides the multimodal large language model to reason and generate the necessity of an alarm, i.e., whether to issue an alarm. Thus, the human-machine aligned personalized scene cognitive reasoning layer is obtained as follows: 2.1) Construct personalized preference configuration information to achieve human-machine cognitive alignment; When the AR system is used for the first time or during periodic calibration, an interactive process is used to capture the user's subjective preferences and generate structured data. Scene demonstration: The AR system displays a series of scenarios to users, including different types of risks; General AI Assessment: The AR system first demonstrates an objective assessment of the scene using a general large model; User subjective correction: Users correct their assessments of general AI based on their own feelings; Differential modeling: The AR system records the difference between general AI evaluation and user subjective corrections to generate personalized preference configuration information; 2.2) Construct a thought chain that includes preference constraints; The egocentric viewpoint image captured by the camera is input into the multimodal large language model, and a prompt word containing instructions is constructed to guide the multimodal large language model to reason step by step: Role setting: Command model: "You are not a general observer, you are a personalized security assistant specifically designed to serve this user"; Context injection: The personalized preference configuration information generated in step 2.1) is embedded into the prompt words as the core knowledge base; Visual perception: The multimodal large language model first identifies objects, distances, and dynamic relationships in the image; Preference matching and inference: The multimodal large language model forces the invocation of embedded preference configuration information for comparison; Decision output: Generate alarm necessity, i.e., whether to issue an alarm.