Multi-mode man-machine interaction method and system for augmented reality environment of AI glasses
By conducting reliability assessments and dynamic weight adjustments on the interactive input information of AI glasses, and identifying and distinguishing compensatory fine-tuning behaviors, the problem of decreased accuracy and reliability of AI glasses in complex environments is solved. This improves the accuracy of user intent recognition and system robustness, making it suitable for scenarios such as industrial manufacturing and medical surgery.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- SHENZHEN DINGHAODA TECHNOLOGY CO LTD
- Filing Date
- 2026-01-29
- Publication Date
- 2026-05-12
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
In long-term, high-intensity use, AI glasses suffer from a decline in the accuracy and reliability of interaction due to environmental factors and user compensation behaviors, as well as the difficulty in accurately distinguishing the user's true intentions.
By acquiring operator input information, environmental parameters, and motion sensing component status, reliability assessment is performed, and the weight of the input information is dynamically adjusted to identify and distinguish compensatory fine-tuning behaviors. Combined with laser pulse time stamps, time correlation analysis is conducted to accurately identify user intent.
It significantly improves the accuracy and robustness of AI glasses in complex environments, reduces misoperation, and enhances user experience, making it particularly suitable for precision operation scenarios such as industrial manufacturing and medical surgery.
Smart Images

Figure CN122018688A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of human-computer interaction technology, and in particular to a multimodal human-computer interaction method and system for AI glasses augmented reality environments. Background Technology
[0002] AI glasses, as a system enabling multimodal human-computer interaction in augmented reality environments, are fundamentally based on the integration of eye tracking, gesture recognition, and speech semantic understanding to establish a three-dimensional spatial interaction correspondence, thereby achieving natural operation of the augmented reality interface. This technology has significant application value in scenarios requiring highly precise operations, such as industrial manufacturing and medical surgery, effectively addressing the shortcomings of traditional single-point interaction devices in terms of information transmission efficiency and operational complexity. However, in actual long-term, high-intensity applications, the performance of AI glasses may be subtly affected by environmental factors, leading to a decrease in their interaction accuracy and reliability.
[0003] For example, in the cleanrooms of semiconductor manufacturing, technicians wear AI glasses to perform optical alignment calibration on high-precision lithography machines. Ideally, the AI glasses can precisely overlay virtual alignment crosshairs and rulers, and achieve efficient and accurate adjustments through voice commands and gestures. However, during long-term operation, subtle temperature fluctuations in the workshop's air conditioning system can cause minute changes in the local ambient temperature around the lithography machine, leading to extremely slight thermal expansion and contraction of the physical components. This continuous and subtle change in ambient temperature, along with the natural fine-tuning of the technician's head posture, can cause the motion sensing components inside the AI glasses to accumulate minute deviations, resulting in a spatial misalignment between the virtual alignment crosshairs in the augmented reality interface and the actual alignment marks on the physical optical components that is imperceptible to the naked eye.
[0004] This virtual-physical spatial misalignment causes the eye-tracking and gesture recognition systems of AI glasses to deviate in their judgment of user intentions. Even if the technician's eyes are accurately focused on the physical calibration point, the system may still misinterpret their intention due to a misunderstanding of the virtual target's location. Similarly, feedback from gesture operations may become sluggish and inaccurate. Faced with this subtle visual deviation and uncertainty in interaction, technicians instinctively try to "correct" the virtual overlay by slightly adjusting their head posture or gaze angle to realign it with the physical target. However, this natural head movement made by the user to compensate for system deviations further interferes with the AI glasses' head posture tracking system, making it difficult for the system to accurately distinguish the user's true intentions. Sometimes, these subtle head adjustments are misinterpreted as commands to pan or zoom the interface, causing the virtual interface to inadvertently jump or flicker, further distracting the technician and exacerbating the sense of confusion during operation.
[0005] In complex scenarios where subtle spatial discrepancies exist and multiple interactive inputs become ambiguous, the integration and judgment mechanism of various interaction methods in AI glasses encounters difficulties. The system cannot effectively unify the judgment of these uncertain input information, nor can it determine which input source is more reliable, let alone accurately understand the true intention of the technician. As a result, the system may fail to execute instructions or erroneously execute operations inconsistent with the user's intentions, seriously threatening the success of optical alignment tasks in semiconductor manufacturing.
[0006] To address the aforementioned issues, existing technologies urgently need improvement. Summary of the Invention
[0007] This invention provides a multimodal human-computer interaction method and system for AI glasses in augmented reality environments, aiming to solve the problems of decreased interaction accuracy and reliability caused by environmental factors and user compensation behaviors in long-term, high-intensity applications of AI glasses, as well as the difficulty of the system in accurately distinguishing the user's true intentions.
[0008] The technical solution of this application is as follows: In a first aspect, this application discloses a multimodal human-computer interaction method for AI glasses in augmented reality environments, including: Acquire the operator's interactive input information, including eye position information, hand gesture information, and voice command information; Acquire parameters related to the AI glasses' operating environment and their own state, including environmental parameters and the state information of motion sensing components; Obtain information on the operator's work stage; Based on interactive input information, environmental parameters, state information of motion sensing components, and working stage information, the reliability of interactive input information is evaluated, and the reliability evaluation results of each type of interactive input information are obtained. Based on the reliability assessment results, adjust the weight of each type of interactive input information when judging the operator's intention.
[0009] This technical solution effectively addresses the problem of inaccurate user intent judgment caused by the inconsistent reliability of various interactive input information in complex environments. By dynamically adjusting the weights of different modal interactive information, the accuracy and robustness of human-computer interaction are significantly improved, thereby overcoming the limitations of existing interactive systems that are susceptible to interference from the environment and user behavior.
[0010] Furthermore, this application also proposes a multimodal human-computer interaction method for AI glasses augmented reality environments, wherein recognizing the fine-tuning behavior of the operator to compensate for system deviations includes: Obtain the time stamp of the laser pulse occurrence; Acquire the operator's head posture fine-tuning movement characteristics and gaze movement characteristics; Determine the first temporal correlation between the head posture fine-tuning action and the laser pulse, and identify whether the head posture fine-tuning action is a compensation behavior based on the first temporal correlation, and identify whether the head posture fine-tuning action is a micro-operation command. Determine the second temporal correlation between the gaze movement and the laser pulse, and identify whether the gaze movement is a compensation behavior and whether the gaze movement is a micro-operation command based on the second temporal correlation.
[0011] Through this technical solution, this application can accurately identify the compensatory fine-tuning behavior of the operator caused by the transient micro-displacement induced by the laser pulse, distinguish it from the real operation command, avoid the system's misjudgment of the user's intention, and thus improve the accuracy of the interaction.
[0012] Building upon this, this application further proposes a multimodal human-computer interaction method for AI glasses augmented reality environments, wherein recognizing the fine-tuning behavior of operators to compensate for system biases includes: The head posture fine-tuning motion features are obtained, including the amplitude, speed, and duration of the head posture fine-tuning motion. Acquire gaze movement features, which include the amplitude, speed, and duration of gaze movements; Obtain the time stamp of the laser pulse occurrence; Determine the first time interval between the occurrence time of the head posture fine-tuning action and the occurrence time of the laser pulse, as well as the duration of the head posture fine-tuning action; Determine the second time interval between the occurrence of the gaze movement and the occurrence of the laser pulse, as well as the duration of the gaze movement; Based on the first time interval, the second time interval, and the duration threshold, it is determined whether the head posture fine-tuning action is a compensatory behavior, and the weight of the interactive input information related to the head posture fine-tuning action is adjusted according to the recognition result. The system determines whether the gaze movement falls within a preset physiological response window based on its duration, and whether the duration of the gaze movement is shorter than a duration threshold. It also identifies whether the gaze movement is a compensatory behavior and adjusts the weight of interactive input information related to the gaze movement based on the identification results.
[0013] Through this technical solution, this application can more accurately identify the operator's compensation behavior by performing detailed analysis of the characteristics (amplitude, speed, duration) of head posture fine-tuning movements and gaze movements, as well as their temporal correlation with laser pulses. Based on this, the interaction weights can be dynamically adjusted, thereby effectively improving the system's recognition accuracy of micro-operation commands and reducing misoperations.
[0014] In some preferred embodiments, this application also discloses a multimodal human-computer interaction method for AI glasses augmented reality environments, wherein recognizing the fine-tuning behavior of the operator to compensate for system deviations includes: Obtain information on the relative positional changes between the virtual overlay and the physical world reference point; Identify transient visual drift events and extract their features; Acquire the operator's head posture fine-tuning movement characteristics and gaze movement characteristics; Determine the third temporal correlation between head posture fine-tuning movements and transient visual drift events, and determine the fourth temporal correlation between gaze movements and transient visual drift events; Based on the judgment results of the third time correlation, identify whether the head posture fine-tuning action is a compensatory behavior and whether it is a micro-operation command; Based on the judgment results of the fourth time association, identify whether the gaze movement is a compensatory behavior and whether it is a micro-operation command; The weights of relevant interactive input information are adjusted based on the recognition results of head posture fine-tuning movements and whether gaze movements are compensatory behaviors.
[0015] Through this technical solution, this application can effectively identify the fine-tuning behavior of the user to compensate for visual drift by analyzing the time correlation between the operator's head posture fine-tuning and gaze movement and the event, thereby avoiding misjudging these compensation behaviors as operation commands and improving the system's interaction accuracy in complex visual environments.
[0016] Furthermore, this application also proposes a multimodal human-computer interaction method for AI glasses augmented reality environments, wherein recognizing the fine-tuning behavior of the operator to compensate for system deviations includes: Acquire the operator's head posture fine-tuning movement characteristics and gaze movement characteristics; The system acquires information on various existing system deviations, including persistent spatial misalignment between the virtual overlay and the physical world caused by ambient temperature fluctuations, transient micro-displacement of physical targets caused by laser pulses, and instantaneous visual drift of the virtual overlay caused by local electromagnetic interference or airflow. Obtain the type, occurrence time, intensity, and duration characteristics of each type of systematic deviation information; The head posture fine-tuning action features are analyzed for temporal causal correlation with each systematic deviation event to obtain the fifth temporal correlation between head posture fine-tuning action and systematic deviation. Based on the fifth temporal correlation, the first matching score between each head posture fine-tuning action and each systematic deviation is calculated. The causal relationship between gaze action features and each type of systematic deviation event is analyzed over time to obtain the sixth temporal relationship between gaze action and systematic deviation. Based on the sixth temporal relationship, the second matching score between each gaze action and each systematic deviation is calculated. Based on the first matching score, the head posture fine-tuning action is associated with each system deviation to identify whether the head posture fine-tuning action is a compensatory behavior or a micro-operation command, and the weight of the interactive input information related to the head posture fine-tuning action is adjusted. Based on the second matching score, gaze actions are associated with each system bias for identification, identifying whether gaze actions are compensatory behaviors and micro-operation commands, and adjusting the weights of interactive input information related to gaze actions.
[0017] By comprehensively considering various types and characteristics of system biases, and conducting refined temporal causal correlation analysis and matching score calculation between the operator's head posture fine-tuning and gaze movements and these biases, this application can more accurately identify the root cause of user compensation behavior, thereby achieving more precise interaction weight adjustment and significantly improving the system's robustness and user intent recognition ability under multi-source complex interference.
[0018] As a technological improvement, this application also proposes a multimodal human-computer interaction method for AI glasses augmented reality environments, wherein adjusting the weights of interactive input information related to head posture fine-tuning movements includes: Obtain the task priority of the precision physical operation currently being performed by the operator; Obtain the cognitive load status of the operator currently performing multitasking; Determine the importance of head posture fine-tuning movements in the current task based on task priority; Based on cognitive load status, assess the operator's control precision over fine-tuning of head posture movements; Adjust the weight of head posture fine-tuning movements according to their importance and control precision.
[0019] Through this technical solution, when adjusting interaction weights, this application not only considers the identification of compensation behavior, but also incorporates the assessment of task priority and operator cognitive load status, making weight adjustment more intelligent and contextualized, and able to more accurately reflect the operator's true intentions. Especially in precision operation and multi-task switching scenarios, it significantly improves the adaptability and accuracy of interaction.
[0020] To enhance functionality, this application also proposes a multimodal human-computer interaction method for AI glasses in augmented reality environments, wherein obtaining the cognitive load state of the operator's current multitasking switching includes: Acquire the operator's physiological signals, including heart rate variability data, eye movement data, and electroencephalogram (EEG) data; Obtain the task identifiers of multiple tasks currently being performed by the operator; Based on the task identifier, the cognitive load characteristic parameters of each task are extracted from the preset task characteristic database. The cognitive load characteristic parameters include the complexity of the task, the urgency of the task, and the interaction mode of the task. Based on physiological signals and cognitive load characteristic parameters, a preliminary assessment of the operator's cognitive load status is conducted, and preliminary assessment results are obtained. Determine if the operator is switching tasks; When it is determined that the operator is switching tasks, the weight of the preliminary assessment results is adjusted according to the time when the task switch occurs.
[0021] This technical solution combines physiological signals and task characteristic parameters to comprehensively assess cognitive load, and takes into account the impact of task switching on cognitive load. It can more accurately and in real time reflect the cognitive state of operators, providing a more reliable basis for subsequent adjustment of interaction weights, thereby optimizing the adaptability and efficiency of human-computer interaction.
[0022] Building upon the above, this application further proposes acquiring the operator's physiological signals, including heart rate variability data, eye movement data, and electroencephalogram (EEG) data, including: Continuously monitor the signal quality parameters of each physiological signal acquisition channel, including the signal-to-noise ratio, baseline drift, and signal integrity. When the signal quality parameter of any physiological signal acquisition channel is lower than the preset quality threshold, the abnormal channel is identified. When an abnormal signal is detected, real-time compensation processing of the physiological signals of the abnormal channel is triggered. Adjust the weight of physiological signals from abnormal channels in cognitive load assessment based on the type and degree of signal abnormality in abnormal channels; When multiple physiological signal acquisition channels have signal anomalies, data from other physiological signal channels with better signal quality should be used first for cognitive load assessment. When all physiological signal acquisition channels show abnormal signals that cannot be effectively compensated, the operator is prompted to replace or check the sensors, and the cognitive load assessment results are temporarily frozen until the signal quality returns to normal.
[0023] This technical solution effectively solves the signal anomaly problem that may occur during the physiological signal acquisition process by introducing physiological signal quality monitoring and anomaly compensation mechanism. It ensures the accuracy and reliability of cognitive load assessment and can maintain stable system operation even when the sensor performance is poor or interfered with, and provide timely feedback to users.
[0024] More specifically, in some implementation schemes, this application also proposes a multimodal human-computer interaction method for AI glasses augmented reality environments, wherein when an abnormal signal is detected, real-time compensation processing of the physiological signals of the abnormal channel is triggered, including: Obtain physiological signals from abnormal channels; Spectral analysis of physiological signals from abnormal channels is performed to obtain the spectral characteristics of the physiological signals; Based on the spectral characteristics of physiological signals, high-frequency interference components in physiological signals are identified. Analyze the frequency characteristics of high-frequency interference components; Based on the frequency characteristics of high-frequency interference components, match the type and operating mode of the interference source; Adjust the center frequency and bandwidth parameters of the adaptive filter according to the type and operating mode of the interference source; Based on the adjusted adaptive filter, dynamic interference noise in the physiological signal is removed; The waveform characteristics of physiological signals after removing dynamically changing interference noise are compared to obtain the comparison results of waveform characteristics; Based on the comparison results of waveform characteristics, the amplitude or phase of the physiological signal is adjusted.
[0025] Through this technical solution, this application can effectively remove dynamically changing interference noise by performing refined spectrum analysis and adaptive filtering compensation on abnormal physiological signals, and further adjust the amplitude or phase of the signal. This allows for the maximum restoration of the authenticity of physiological signals even when signals are abnormal, ensuring the accuracy of cognitive load assessment and significantly improving the robustness and data processing capabilities of the system.
[0026] Secondly, this application also discloses a multimodal human-computer interaction system for AI glasses augmented reality environments, including: The input end is used to acquire interactive input information from the operator, including gaze position information, hand movement information, and voice command information; acquire parameters related to the AI glasses' operating environment and its own state, including environmental parameters and the state information of motion sensing components; and acquire information about the operator's work stage. The evaluation end is used to evaluate the reliability of interactive input information based on interactive input information, environmental parameters, status information of motion sensing components, and working stage information, and to obtain the reliability evaluation result of each type of interactive input information. The adjustment end is used to adjust the weight of each type of interactive input information in determining the operator's intention based on the reliability assessment results.
[0027] This application provides a structured system that can efficiently acquire multimodal interaction information and environmental status, conduct reliability assessments, and dynamically adjust interaction weights. This system addresses the problem of insufficient interaction accuracy of AI glasses in complex environments at the system level, providing an integrated hardware and software solution for achieving smarter and more reliable human-computer interaction. Beneficial effects
[0028] This application discloses a multimodal human-computer interaction method for AI glasses in augmented reality environments. It acquires interactive input information such as the operator's gaze position, hand gestures, and voice commands, and combines this information with parameters related to the AI glasses' operating environment and its own state (such as environmental parameters and the state information of motion sensing components), as well as the operator's work stage information, to conduct a reliability assessment of these interactive inputs. Based on the assessment results, the system can dynamically adjust the weight of each interactive input when determining the operator's intention.
[0029] This method effectively solves the problems in existing technologies where AI glasses, under long-term high-intensity use, suffer from decreased interaction accuracy and reliability due to environmental factors (such as spatial misalignment caused by temperature fluctuations, transient micro-displacement caused by laser pulses, visual drift caused by electromagnetic interference or airflow) and fine-tuning behavior by operators to compensate for system deviations, as well as the difficulty in accurately distinguishing the user's true intentions.
[0030] Specifically, this application, by conducting a reliability assessment of multimodal interaction information, can identify which interaction modalities are interfered with and which are relatively reliable in complex environments. For example, when the system detects a slight spatial misalignment between the virtual overlay and the physical world, it may reduce the weight of gaze position information and hand gesture information while increasing the weight of voice command information to avoid misoperation due to inaccurate visual or gesture input. Simultaneously, by recognizing and distinguishing the fine-tuning behaviors (such as head posture adjustments and gaze movements) produced by the operator to compensate for system deviations from the actual operation commands, the system avoids misinterpreting these compensatory behaviors as user intentions, thereby reducing interference such as virtual interface jitter or flickering and improving the smoothness and accuracy of operation.
[0031] In summary, this application significantly improves the accuracy, robustness, and user experience of AI glasses in augmented reality environments by dynamically and intelligently adjusting the weights of multimodal interaction information. It is particularly suitable for scenarios such as industrial manufacturing and medical surgery that require highly precise operations, effectively overcoming the limitations of existing interactive systems that are susceptible to interference from the environment and user behavior. Attached Figure Description
[0032] Figure 1 This is a flowchart of a method for multimodal human-computer interaction in an augmented reality environment using AI glasses, provided by an embodiment of the present invention. Figure 2 This is a schematic diagram of the structure of a multimodal human-computer interaction system for AI glasses augmented reality environment provided in an embodiment of this application. Detailed Implementation
[0033] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0034] Reference Figure 1 , Figure 1 This is a flowchart of a method for multimodal human-computer interaction in an augmented reality environment using AI glasses, provided by an embodiment of the present invention, including: S11, Obtain the operator's interactive input information, which includes eye position information, hand movement information, and voice command information; S12, acquire parameters related to the operating environment and state of the AI glasses, including environmental parameters and state information of motion sensing components; S13, Obtain the work stage information of the operator; S14, based on the interactive input information, the environmental parameters, the state information of the motion sensing component, and the working stage information, perform a reliability assessment on the interactive input information to obtain a reliability assessment result for each type of interactive input information; S15, Based on the reliability assessment results, adjust the weight of each interactive input information in determining the operator's intention.
[0035] This application effectively solves the problem of decreased interaction accuracy caused by environmental influences and system deviations in the prior art by conducting reliability assessments of multimodal interactive input information and dynamically adjusting its weights based on the assessment results. This significantly improves the interaction reliability and user intent recognition accuracy of AI glasses in complex and precise operation scenarios.
[0036] To better understand the technical solution proposed in this application, the following will provide a detailed explanation of some key terms and implementation environments involved.
[0037] The "AI glasses" referred to in this application are an augmented reality device that integrates multiple sensors and computing units, which can overlay virtual information in the operator's field of vision and support multimodal interaction methods.
[0038] "Multimodal human-computer interaction" refers to the system's ability to simultaneously receive and process input information from different modalities, such as gaze, gestures, and voice, to comprehensively determine the operator's intentions. "Interactive input information" refers to the data generated when the operator interacts with the system through AI glasses, including "gaze position information," "hand movement information," and "voice command information." "Gaze position information" can be obtained through eye-tracking technology, reflecting the operator's gaze point and gaze trajectory; "hand movement information" can be obtained through gesture recognition technology, reflecting the operator's hand posture and movements; and "voice command information" can be obtained through the microphone array built into the AI glasses, and after noise reduction processing, it is converted into text by the speech recognition module, and then semantically understood by the natural language processing module.
[0039] "Parameters related to the operating environment and the AI glasses' own state" refer to external environmental factors and internal hardware states that affect the performance and interaction accuracy of AI glasses, including "environmental parameters" and "state information of motion sensing components." "Environmental parameters" can include ambient temperature, humidity, light intensity, etc.; "state information of motion sensing components" can include the drift of the inertial measurement unit (IMU), camera calibration parameters, etc. "Work phase information" refers to the specific stage of the task the operator is currently in, such as the start, execution, or end of the task, or the switching of a specific subtask.
[0040] The core of the multimodal human-computer interaction method for AI glasses augmented reality environments proposed in this application lies in the reliability assessment of interactive input information and the dynamic adjustment of weights to improve the system's accuracy in recognizing the operator's intentions.
[0041] Firstly, regarding "acquiring interactive input information from operators," various methods can be employed. For instance, gaze position information can be acquired through an eye-tracking sensor integrated into the AI glasses. This sensor can monitor the operator's eye movements in real time and calculate the coordinates of the gaze point within the augmented reality interface. Hand movement information can be acquired through a depth camera or infrared sensor mounted on the AI glasses. These sensors can capture the three-dimensional posture and movement trajectory of the operator's hands and recognize predefined gestures through image processing algorithms. Voice commands can be acquired through the microphone array built into the AI glasses. After noise reduction, the speech recognition module converts the commands into text, which is then semantically understood by the natural language processing module.
[0042] Secondly, regarding "acquiring parameters related to the operating environment and the AI glasses' own state," environmental parameters can be obtained through environmental sensors integrated into the AI glasses, such as temperature sensors, humidity sensors, and light sensors. These sensors can monitor the environmental conditions around the AI glasses in real time. The state information of motion sensing components can be obtained by analyzing data from the inertial measurement unit (IMU) inside the AI glasses. For example, algorithms such as Kalman filtering can be used to fuse IMU data to estimate its drift and calibration error.
[0043] Secondly, regarding "obtaining information about the operator's work stage," this can be achieved through various means. For example, the system can automatically identify the current work stage based on a preset task flow, using the operator's voice commands or gesture confirmations. Additionally, it can obtain the progress and stage information of the current task by interacting with an external task management system.
[0044] Next, regarding the aspect of "evaluating the reliability of interactive input information based on interactive input information, environmental parameters, state information of motion sensing components, and operational stage information, and obtaining the reliability evaluation result for each type of interactive input information," a multi-factor fusion evaluation model can be constructed. For example, when the ambient temperature fluctuates significantly, it may affect the stability of optical components, leading to a decrease in the accuracy of gaze tracking, and the reliability evaluation result of gaze position information will decrease accordingly. When the motion sensing components experience significant drift, the accuracy of hand gesture recognition will be affected, and the reliability evaluation result of hand gesture information will also decrease. Furthermore, during critical operational stages where operators perform precise operations, the system may assign higher reliability expectations to certain interactive input information (such as fine-tuning gestures). Specifically, machine learning models, such as support vector machines (SVM) or neural networks, can be used to train the model using the aforementioned multiple parameters as input features, outputting a reliability score for each type of interactive input information.
[0045] Finally, regarding "adjusting the weight of each type of interactive input information in determining the operator's intention based on the reliability assessment results," a dynamic weight adjustment strategy can be adopted. For example, when the reliability assessment result of gaze position information is low, the system will reduce its weight in determining the operator's intention, relying more on hand gesture information or voice command information. Conversely, when the reliability of a certain type of interactive input information is high, its weight will be increased accordingly. This adjustment can be linear or non-linear, depending on the mapping relationship between the assessment results and the weights. For example, a threshold can be set; when the reliability is below this threshold, the weight drops sharply; when it is above this threshold, the weight increases slowly.
[0046] The multimodal human-computer interaction method for AI glasses in augmented reality environments proposed in this application effectively solves the problem of decreased accuracy and reliability of AI glasses interaction in complex environments through the synergistic effect of the aforementioned technical features. In the cleanrooms of semiconductor manufacturing, when environmental temperature fluctuations cause slight misalignment between the virtual alignment crosshairs and physical markers, traditional systems may fail to accurately identify the true intentions of technicians. However, the method of this application first acquires environmental parameters (such as temperature fluctuations) and the state information of motion sensing components (such as IMU drift), and combines this with the technician's gaze position information, hand gesture information, and voice command information, as well as the current alignment stage of the lithography machine. Based on this multi-dimensional data, the system can perform reliability assessments on each type of interactive input information. For example, when the environmental temperature fluctuates significantly, the system determines that the reliability of gaze tracking and gesture recognition may decrease; while when the technician issues a clear voice command, the reliability of the voice command may be higher. Based on these reliability assessment results, the system dynamically adjusts the weight of different interactive input information in judging the operator's intentions. For example, when the reliability of eye tracking and gesture recognition is low, the system will reduce their weight and rely more on voice commands to understand the technician's intentions, or adjust the weight of head posture adjustments when judging fine-tuning operations. Therefore, even in the presence of system bias and environmental interference, the method of this application can more accurately identify the operator's true intentions, avoid misoperations, and thus significantly improve the interaction efficiency and safety of AI glasses in precision operation scenarios.
[0047] The core innovation of this application lies in introducing a reliability assessment mechanism for multimodal interactive input information and dynamically adjusting its weights based on the assessment results. Compared to the fixed weights or simple fusion strategies commonly used in existing technologies, the method of this application can more intelligently and adaptively handle the uncertainty of interaction in complex environments. For example, in traditional systems, when environmental temperature fluctuations cause virtual-physical space misalignment, the system may be unable to distinguish whether the technician intends to perform interface translation or compensate for system deviations, leading to misoperation. This application, however, acquires environmental parameters, motion sensing component status information, and operational stage information to conduct reliability assessments of interactive inputs such as gaze, gestures, and voice, and dynamically adjusts their weights based on the assessment results. For example, when gaze tracking reliability decreases due to environmental factors, the system reduces the weight of gaze information and relies more on gestures or voice commands to determine the user's intent. This dynamic adjustment mechanism allows the system to better adapt to environmental changes and system deviations, thereby significantly improving the accuracy and reliability of human-computer interaction in augmented reality environments, especially in industrial and medical scenarios requiring highly precise operations, where its advantages are even more pronounced.
[0048] In some embodiments described above, this application proposes a method for reliability assessment and weighting of operator input information. However, in augmented reality environments, operators performing precision tasks may unconsciously make minor head posture adjustments or eye movements to compensate for slight system deviations. If these compensatory adjustments are incorrectly identified as operator intentions, the reliability assessment of the input information will be biased, affecting the accuracy of judging the operator's true intentions and reducing the efficiency of human-computer interaction and user experience. Therefore, this application further proposes a method for identifying operator adjustments made to compensate for system deviations, thereby improving the accuracy of reliability assessment of input information by accurately distinguishing between compensatory adjustments and operational instructions.
[0049] The above adjustments to the weight of each type of interactive input in determining the operator's intent include: Obtain the time stamp of the laser pulse occurrence; Acquire the operator's head posture fine-tuning action characteristics and gaze action characteristics; Determine the first temporal correlation between the head posture fine-tuning action and the laser pulse, and identify whether the head posture fine-tuning action is a compensation behavior based on the first temporal correlation, and identify whether the head posture fine-tuning action is a micro-operation command. Determine the second temporal correlation between the gaze action and the laser pulse, and identify whether the gaze action is a compensation behavior based on the second temporal correlation, and identify whether the gaze action is a micro-operation command.
[0050] Specifically, the timestamp of the laser pulse refers to the timestamp recorded at a specific moment when the laser pulse emitted by the system is generated in the augmented reality environment of the AI glasses. This laser pulse can be used for target localization, distance measurement, or system calibration, and its occurrence time serves as an important reference point for determining whether the operator's subsequent fine-tuning actions are compensatory. The operator's head posture fine-tuning characteristics can be understood as changes in the operator's head posture that occur with extremely small amplitude and short duration, such as those monitored and collected in real time by the inertial measurement unit integrated within the AI glasses or by external posture sensors. The gaze movement characteristics refer to minute movements or changes in gaze point that occur in the operator's eyes within an extremely short time, such as data on gaze direction and gaze point position obtained through eye-tracking sensors. These fine-tuning actions are typically small in amplitude and short in duration, and may not possess explicit semantic instructions.
[0051] In practical applications, determining the first temporal correlation between the head posture fine-tuning action and the laser pulse refers to analyzing the time interval between the occurrence of the head posture fine-tuning action and the occurrence of the laser pulse. For example, if the head posture fine-tuning action occurs within a preset very short time window (e.g., tens to hundreds of milliseconds) after the laser pulse occurs, a first temporal correlation is considered to exist. Based on this first temporal correlation, it can be identified whether the head posture fine-tuning action is a compensatory behavior, that is, whether it is an unconscious adjustment made in response to system feedback or target changes caused by the laser pulse. Simultaneously, it can also be identified whether the head posture fine-tuning action is a micro-operation command, that is, whether it is a conscious command issued by the operator for precise control.
[0052] Similarly, determining the second temporal correlation between the gaze action and the laser pulse refers to analyzing the time interval between the occurrence of the gaze action and the occurrence of the laser pulse. If the gaze action occurs within a preset, extremely short time window after the laser pulse occurs, a second temporal correlation is considered to exist. Based on this second temporal correlation, it can be identified whether the gaze action is a compensatory behavior, such as a rapid gaze adjustment to address laser point drift; and whether the gaze action is a micro-operation command, such as confirmation or selection through rapid scanning.
[0053] This application's solution, by introducing a time stamp of laser pulse occurrence as a reference, enables precise temporal causal correlation analysis of operator head posture adjustments and gaze movements with specific system events (such as laser pulses). When operators perform precision operations in an augmented reality environment, laser pulses emitted by the system may cause instantaneous changes in visual feedback or slight drifts in target position. The operator's perceptual system quickly captures these changes and instinctively compensates for these deviations with minute head or gaze adjustments to maintain precise alignment or focus on the target. By judging the temporal correlation between these adjustments and the laser pulse—for example, if the adjustment occurs immediately after the laser pulse—it is highly suspected of being compensatory behavior. This mechanism allows the system to distinguish between unconscious, reactive adjustments caused by external stimuli or system biases and proactive, intentional micro-operation commands issued by the operator. This avoids misinterpreting compensatory behavior as operational commands, thereby improving the accuracy of identifying the operator's true intentions.
[0054] Through the above technical solution, this application effectively solves the problem of distinguishing between operator compensatory fine-tuning behaviors and actual operation instructions in traditional methods. By introducing laser pulse time stamps as a reference and establishing a temporal correlation between head posture fine-tuning actions and gaze movements and laser pulses, the system can more accurately identify operator fine-tuning behaviors caused by system deviations. This precise recognition capability allows for lower weighting of these compensatory behaviors in subsequent reliability assessments of interactive input information, or even excluding them from valid instructions, thereby significantly improving the accuracy of judging operator intentions. Ultimately, this contributes to building a more intelligent, robust, and user-friendly AI glasses augmented reality multimodal human-computer interaction system, with its advantages being particularly evident in scenarios requiring high-precision operation.
[0055] In some preferred embodiments, suppose an operator wearing AI glasses is performing a precision industrial assembly task that requires aligning a laser pointer with a tiny component. When the AI glasses system emits a laser pulse to mark a target location, the laser point may momentarily shift due to environmental factors or minor system jitter. The operator's visual system immediately detects this shift and may unconsciously make a tiny head posture adjustment or a rapid gaze scan within tens of milliseconds after the laser pulse to precisely realign the laser point. The AI glasses system records the time of the laser pulse. Simultaneously, through its built-in inertial sensors and eye-tracking module, the system captures the characteristics of the operator's head posture adjustment and gaze movements. The system then determines the time interval between these adjustments and the laser pulse. If the head posture adjustment or gaze movement occurs within a preset physiological response window (e.g., 50 to 200 milliseconds) after the laser pulse, and its amplitude, speed, and other characteristics conform to a compensatory behavior pattern, the system identifies it as compensatory behavior rather than a new operational command from the operator. For example, if the operator's gaze quickly shifts slightly towards the laser point after the laser pulse, and the duration is extremely short, it is judged as a compensatory gaze adjustment. In this way, the system can avoid misinterpreting these unintentional compensatory fine-tunings as operational commands, thereby ensuring that these fine-tuning behaviors will not interfere with the judgment of the operator's true intentions in subsequent reliability assessments of interactive input information, thus guaranteeing the accuracy and smoothness of precision operations.
[0056] In some of the embodiments described above in this application, although the reliability of the operator's interactive input information has been assessed and weights adjusted accordingly, in augmented reality environments, operators may unconsciously produce some fine-tuning behaviors to compensate for subtle deviations in the AI glasses system or the environment, such as visual transient changes caused by laser pulses. These fine-tuning behaviors, such as minor adjustments to head posture or rapid eye movements, if incorrectly identified as intentional commands from the operator, may lead to misjudgment of the operator's intent by the system, thereby affecting the accuracy and fluency of the interaction.
[0057] In response, this application further proposes a method for adjusting the weight of each type of interactive input information when determining the operator's intent, the steps of which include: The head posture fine-tuning action features are obtained, including the amplitude, speed and duration of the head posture fine-tuning action. The gaze movement features are obtained, including the amplitude, speed, and duration of the gaze movement. Obtain the time stamp of the laser pulse occurrence; Determine the first time interval between the occurrence time of the head posture fine-tuning action and the occurrence time of the laser pulse, and the duration of the head posture fine-tuning action; Determine the second time interval between the occurrence time of the gaze movement and the occurrence time of the laser pulse, and the duration of the gaze movement; Based on the first time interval, the second time interval, and the duration threshold, it is determined whether the head posture fine-tuning action is a compensatory behavior, and the weight of the interactive input information related to the head posture fine-tuning action is adjusted according to the recognition result. The system determines whether the gaze action falls within a preset physiological response window based on its duration, and whether the duration of the gaze action is shorter than a duration threshold. It then identifies whether the gaze action is a compensatory behavior and adjusts the weight of interactive input information related to the gaze action based on the identification results.
[0058] Specifically, the head posture fine-tuning action characteristics refer to subtle changes in head posture captured by the inertial measurement unit (IMU) built into the AI glasses or an external tracking system. Their amplitude, speed, and duration are all within a specific threshold range, typically much smaller than the action characteristics of intentional commands. The gaze action characteristics refer to rapid, small-range movements of the gaze direction or fixation point captured by eye-tracking sensors; their amplitude, speed, and duration also exhibit typical characteristics of fine-tuning behavior. The laser pulse occurrence time stamp refers to the precise timestamp of the laser pulse emitted by the AI glasses or external device. This pulse may be used for target indication, distance measurement, or depth perception, and may visually elicit an instantaneous reaction from the operator.
[0059] The first time interval refers to the time difference between the start time of the head posture fine-tuning movement and the laser pulse occurrence time, and the second time interval refers to the time difference between the start time of the gaze movement and the laser pulse occurrence time. These time intervals are used to assess whether there is a causal relationship between the fine-tuning movement and the laser pulse event. The duration threshold is a preset time length used to distinguish between brief compensatory fine-tuning and intentional operations with a longer duration. The preset physiological response window refers to the time range within which humans produce an unconscious physiological response to a specific stimulus (such as a laser pulse). For example, the rapid reflexive movement of the gaze in response to a momentary strong light stimulus typically occurs within tens to hundreds of milliseconds.
[0060] This application's solution, by meticulously acquiring head posture fine-tuning and gaze movement characteristics and combining them with the time stamp of laser pulse occurrence, enables in-depth analysis of the operator's fine-tuning behavior. Through this technical solution, the application can accurately identify the operator's fine-tuning behavior in augmented reality environments to compensate for system biases, such as unconscious reactions to instantaneous visual changes caused by laser pulses. This significantly improves the accuracy of AI glasses in recognizing the operator's true intentions, avoiding the misinterpretation of compensatory fine-tuning movements as intentional commands, thereby effectively reducing the occurrence of misoperations. Consequently, the reliability and user experience of multimodal human-computer interaction are significantly improved, enabling operators to obtain a smoother and more natural interactive experience when performing precise tasks, further enhancing the practicality and robustness of AI glasses in complex augmented reality environments.
[0061] In some preferred embodiments, suppose an operator is using AI glasses to perform an industrial assembly task requiring precise aiming, where the AI glasses periodically emit laser pulses to assist in positioning. After a laser pulse occurs, the system detects a minor head movement within the next 100 milliseconds, with an amplitude less than 0.5 degrees and a duration of 50 milliseconds. Simultaneously, the eye-tracking system detects a rapid, less than 1-degree jerk in the operator's gaze within 30 milliseconds after the laser pulse, lasting 20 milliseconds. At this point, the system determines that the time interval between the head posture adjustment and the laser pulse is 0 to 100 milliseconds, and the duration is 50 milliseconds, both less than a preset duration threshold (e.g., 150 milliseconds). For the gaze movement, the time interval between the laser pulse and the gaze movement is 30 milliseconds, falling within a preset physiological response window (e.g., 20-80 milliseconds), and the duration of 20 milliseconds is also shorter than the duration threshold. Based on these determinations, the system identifies both the head posture adjustment and the gaze movement as compensatory behaviors of the operator in response to the laser pulse stimulus, rather than intentionally issued commands. Therefore, in the subsequent intent determination, the weight of head posture information and gaze position information related to these fine-tuning actions will be significantly reduced to ensure that the system does not misinterpret these unconscious compensatory behaviors as operation commands, thereby ensuring the accuracy of operation and the stability of the system.
[0062] In some embodiments described above, the reliability of the operator's interactive input information is assessed, and weights are adjusted based on the assessment results to determine the operator's intent. However, in augmented reality environments, there may be system deviations such as instantaneous visual drift between the virtual overlay of AI glasses and the physical world. Operators may unconsciously make minor head posture adjustments or eye movements to compensate for these deviations. If these compensatory minor adjustments are incorrectly identified as operation commands, the system may misjudge the operator's true intent, thereby affecting the accuracy and fluency of the interaction. To address this, this application further proposes a method for identifying minor adjustments made by operators to compensate for system deviations. By introducing the detection and analysis of instantaneous visual drift events and temporally correlating them with the operator's minor adjustments, the method can more accurately distinguish between compensatory behaviors and micro-operation commands, thereby optimizing the weight adjustment of interactive input information.
[0063] Specifically, the aforementioned adjustments to the weight of each type of interactive input in determining the operator's intent include: Obtain information on the relative positional changes between the virtual overlay and the physical world reference point; Identify transient visual drift events and extract features from these events; Acquire the operator's head posture fine-tuning movement characteristics and gaze movement characteristics; Determine the third temporal correlation between the head posture fine-tuning action and the instantaneous visual drift event, and determine the fourth temporal correlation between the gaze action and the instantaneous visual drift event; Based on the judgment result of the third time correlation, identify whether the head posture fine-tuning action is a compensation behavior and whether it is a micro-operation command; Based on the judgment result of the fourth time association, identify whether the gaze action is a compensation behavior and whether it is a micro-operation command; Based on the recognition results of the head posture fine-tuning movements and whether the gaze movements are compensation behaviors, the weights of the relevant interactive input information are adjusted.
[0064] Obtaining information on the relative positional changes between the virtual overlay and a physical world reference point involves continuously monitoring the alignment between the projected position of the virtual content in the physical world and the actual physical reference point using sensors in the AI glasses (such as cameras and depth sensors). Specifically, this can be achieved using Simultaneous Localization and Mapping (SLAM) technology or feature-based tracking algorithms to calculate the spatial deviation between the virtual overlay and the physical world reference point in real time. The purpose of this is to provide basic data for subsequent identification of system deviations.
[0065] Furthermore, identifying and extracting the features of instantaneous visual drift events means that when the aforementioned relative position change information exhibits a rapid, unexpected displacement exceeding a preset threshold within a short period of time, it is identified as an instantaneous visual drift event. The features of the instantaneous visual drift event may include the amplitude, duration, direction, and frequency of the drift, with the aim of quantifying the characteristics of the visual drift to facilitate subsequent correlation analysis with the operator's fine-tuning actions.
[0066] Acquiring the operator's head posture fine-tuning and gaze movement characteristics refers to using the inertial measurement unit (IMU) built into the AI glasses or external tracking devices to capture subtle changes in the operator's head posture, and using eye-tracking sensors to capture subtle movements of the operator's gaze. The head posture fine-tuning characteristics can include the amplitude, speed, and duration of the head posture fine-tuning movements; the gaze movement characteristics can include the amplitude, speed, and duration of the gaze movement. The purpose is to capture unconscious or semi-conscious compensatory movements that the operator may produce when perceiving system deviations.
[0067] Determining the third temporal correlation between the head posture fine-tuning action and the instantaneous visual drift event, and determining the fourth temporal correlation between the gaze movement and the instantaneous visual drift event, refers to analyzing whether the occurrence times of the head posture fine-tuning action and the gaze movement are closely related to the occurrence time of the instantaneous visual drift event within a specific time window. For example, a time threshold can be set to determine whether the fine-tuning action occurs within a specific time window after the visual drift event, with the aim of establishing a causal relationship between system deviation and operator response.
[0068] Based on the judgment result of the third temporal correlation, it is determined whether the head posture fine-tuning action is a compensatory behavior and whether it is a micro-operation command; based on the judgment result of the fourth temporal correlation, it is determined whether the gaze movement is a compensatory behavior and whether it is a micro-operation command. Specifically, if the head posture fine-tuning action or gaze movement has a strong third or fourth temporal correlation with the instantaneous visual drift event, the action is identified as a compensatory behavior; conversely, if there is no significant temporal correlation, the action is more likely to be identified as a micro-operation command issued by the operator. The purpose is to distinguish the operator's intention and avoid misinterpreting compensatory actions as commands.
[0069] Therefore, based on the recognition results of whether the head posture fine-tuning action and the gaze action are compensating behaviors, the weights of the relevant interactive input information are adjusted. If it is identified as a compensating behavior, the weight of this action in judging the operator's intention will be reduced or even ignored; if it is identified as a micro-operation command, its weight may be maintained or increased to ensure that the system accurately responds to the operator's true intention.
[0070] This application's solution, by real-time monitoring of the relative positional changes between the virtual overlay and the physical world reference point, can promptly identify instantaneous visual drift events and extract their features. Simultaneously, the system acquires the operator's head posture fine-tuning movements and gaze movements. Crucially, by determining the temporal correlation between these fine-tuning movements and instantaneous visual drift events, this application can effectively distinguish between unconscious or semi-conscious fine-tuning behaviors caused by the operator compensating for system biases and actively issued micro-operation commands. It is precisely this precise distinguishing ability that allows the system to more accurately reflect the operator's true intentions when subsequently adjusting the weights of interactive input information, avoiding misjudgments caused by system biases.
[0071] Through the above technical solution, this application can significantly improve the accuracy and robustness of multimodal human-computer interaction in augmented reality environments using AI glasses. Especially in the presence of system biases such as momentary visual drift, the system no longer misinterprets the operator's compensatory fine-tuning movements as operational commands, thus avoiding unnecessary system responses and a decline in user experience. This allows operators to interact with the augmented reality environment more smoothly and naturally, improving interaction efficiency and satisfaction, with its advantages being particularly pronounced in application scenarios requiring high-precision operations.
[0072] In some preferred embodiments, it is assumed that an operator is using AI glasses to perform a precise virtual assembly task. During the task, due to slight electromagnetic interference in the environment, a momentary visual drift event occurs on the physical workbench, causing the overlay of virtual assembly parts to shift instantaneously. For example, the virtual part suddenly shifts 2 mm to the left, lasting approximately 150 milliseconds. Unconsciously, the operator makes a 0.3-degree adjustment to the right to realign the virtual part, while simultaneously moving their gaze quickly to the right to track it. The solution of this application first obtains information on the relative positional change between the virtual overlay and a physical world reference point, and identifies the momentary visual drift event and its characteristics (e.g., a 2 mm shift and a 150-millisecond duration).
[0073] Simultaneously, the system acquires the operator's head posture fine-tuning motion characteristics (e.g., a 0.3-degree rightward head turn lasting 200 milliseconds) and gaze motion characteristics (e.g., a quick rightward glance lasting 100 milliseconds). The system then determines whether there is a close third and fourth temporal correlation between the timing of these head posture fine-tuning and gaze motions and the timing of the instantaneous visual drift event. If these actions all occur within 250 milliseconds after the visual drift event, the system, based on these temporal correlation determinations, identifies the operator's head posture fine-tuning and gaze motions as compensatory behaviors rather than proactive micro-operation commands. Ultimately, based on this identification result, the system reduces or ignores the weight of these head posture and gaze inputs when judging the operator's intentions, thereby preventing the system from incorrectly interpreting these compensatory actions as commands for moving virtual components or selecting an option, ensuring the accuracy of the interaction.
[0074] Traditional multimodal human-computer interaction methods in AI glasses augmented reality environments often struggle to accurately distinguish between the operator's voluntary intentions and unconscious compensatory fine-tuning behaviors caused by system biases when assessing the reliability of interactive input information. For example, in augmented reality environments, when a brief or persistent misalignment occurs between the virtual overlay and the physical world, the operator may instinctively make minor adjustments through head posture or gaze to correct the perceptual deviation, rather than issuing explicit commands. If this problem is not addressed, these compensatory fine-tuning behaviors may be misinterpreted as the operator's intentions, leading to inaccurate system responses and reduced interaction efficiency and user experience. To address this, this application proposes a more refined method for identifying operator fine-tuning behaviors that compensate for system biases. By deeply analyzing the causal relationship between system biases and operator fine-tuning actions, the accuracy of interaction intention judgment is improved.
[0075] In some embodiments described above in this application, adjusting the weight of each type of interactive input information in determining the operator's intent specifically includes: Acquire the operator's head posture fine-tuning movement characteristics and gaze movement characteristics; The system acquires information on various existing system deviations, including persistent spatial misalignment between the virtual overlay and the physical world caused by ambient temperature fluctuations, transient micro-displacement of physical targets caused by laser pulses, and instantaneous visual drift of the virtual overlay caused by local electromagnetic interference or airflow. Obtain the type, occurrence time, intensity, and duration characteristics of each type of system deviation information; The head posture fine-tuning action features are analyzed for temporal causal correlation with each of the system deviation events to obtain the fifth temporal correlation between the head posture fine-tuning action and the system deviation. Based on the fifth temporal correlation, the first matching score between each head posture fine-tuning action and each system deviation is calculated. The gaze movement features are analyzed for temporal causal correlation with each of the system deviation events to obtain a sixth temporal correlation between the gaze movement and the system deviation. A second matching score between each gaze movement and each system deviation is calculated based on the sixth temporal correlation. Based on the first matching score, the head posture fine-tuning action is associated with each system deviation to identify whether the head posture fine-tuning action is a compensation behavior and whether it is a micro-operation command, and the weight of the interactive input information related to the head posture fine-tuning action is adjusted. Based on the second matching score, the gaze action is associated with each system deviation to identify whether the gaze action is a compensation behavior and whether it is a micro-operation command, and the weight of the interactive input information related to the gaze action is adjusted.
[0076] Specifically, the head posture fine-tuning motion characteristics can be understood as the operator's subtle head movements or rotations within a very small range, the amplitude, speed, and duration of which are usually much smaller than actively issued head commands. The gaze motion characteristics refer to the operator's subtle scanning, gazing, or following movements of the eyes within a short period of time, characterized by their rapid, small-range nature and often accompanied by brief focusing on a specific visual target. These characteristics can be precisely acquired and analyzed using the inertial measurement unit and eye-tracking sensors integrated within the AI glasses.
[0077] The various system deviation information aims to comprehensively cover all factors in the augmented reality environment that may affect the alignment between the virtual overlay and the physical world. For example, persistent spatial misalignment between the virtual overlay and the physical world caused by ambient temperature fluctuations refers to the long-term, slow drift of the virtual image relative to the real scene caused by minute deformations of the AI glasses' optical components or display screen due to temperature changes. Transient micro-displacement of physical targets induced by laser pulses refers to the momentary mechanical vibration or thermal effect that laser equipment may produce on the surrounding environment or target when emitting pulses in certain industrial or scientific applications, resulting in a brief, imperceptible displacement of the physical target. Instantaneous visual drift of the virtual overlay caused by local electromagnetic interference or airflow refers to the momentary impact of external electromagnetic fields or airflow disturbances on the internal sensors or display system of the AI glasses, causing a brief, rapid jitter or shift in the virtual image. Obtaining the type, occurrence time, intensity, and duration characteristics of these system deviation information helps to build a complete profile of system deviation events, providing accurate temporal and contextual basis for subsequent causal correlation analysis.
[0078] In practical applications, performing temporal causal correlation analysis between the head posture fine-tuning action features and each of the aforementioned systematic deviation events refers to determining whether a reasonable causal chain exists between the occurrence time of the fine-tuning action and the occurrence time of the systematic deviation event by analyzing the temporal relationship between the two. For example, if a head fine-tuning action occurs immediately after a systematic deviation event, and its duration matches the perceived correction time of the deviation, a causal relationship may exist. The fifth and sixth temporal correlations are quantitative descriptions of this temporal and causal relationship. Based on this, the first and second matching scores are calculated to quantify the strength and probability of the correlation between head posture fine-tuning actions or gaze movements and specific systematic deviation events. The matching score can comprehensively consider multiple dimensions such as time interval, the proportional relationship between the action amplitude and the deviation intensity, and the consistency between the action direction and the deviation direction.
[0079] Furthermore, based on the first and second matching scores, the head posture fine-tuning movements and gaze movements are associated with each system deviation to determine whether these fine-tuning movements are indeed intended to compensate for a specific system deviation. If the matching score reaches a preset threshold, the fine-tuning movement is identified as a compensatory behavior. Simultaneously, it is also necessary to identify whether these fine-tuning movements are micro-operation commands; that is, although they are minor movements, they are not compensatory behaviors but rather precise control commands issued by the operator. Through this precise identification, it is possible to avoid misjudging compensatory behaviors as operational intentions, thereby adjusting the weight of interactive input information related to the head posture fine-tuning movements and gaze movements, reducing the influence of compensatory behaviors in intention judgment, and increasing the weight of genuine operational commands.
[0080] This application's solution acquires and analyzes in detail the operator's head posture fine-tuning action features and gaze movement features, and combines this with comprehensive perception and feature extraction of multiple system deviation information to establish a precise temporal and causal relationship between the operator's fine-tuning behavior and system deviation events. Through this technical solution, this application can significantly improve the accuracy of multimodal human-computer interaction in AI glasses augmented reality environments. Specifically, by introducing refined perception and analysis of multiple system deviation information, and combining temporal causal relationship analysis and matching score calculation between the operator's fine-tuning actions and system deviation events, the system can more accurately identify the operator's compensatory behavior and distinguish it from genuine micro-operation instructions. This allows for more precise adjustment of the weights of different interactive input information during reliability assessment, effectively reducing the interference of compensatory fine-tuning caused by system deviations on the operator's intention judgment, thereby avoiding misoperation and improving the efficiency and user experience of human-computer interaction. Furthermore, this solution enhances the system's adaptability and stability in complex and ever-changing augmented reality environments.
[0081] In some embodiments described above, this application proposes identifying fine-tuning behaviors by operators to compensate for system deviations and adjusting the weights of interactive input information related to these head posture fine-tuning actions based on the identification results. However, in practical applications, simply adjusting the weights based on the identification results of the fine-tuning behaviors may not adequately consider the complexity of the operator's working environment and personal state. For example, when performing tasks of different precision or under conditions of high cognitive load, the same head posture fine-tuning action may represent different operational intentions or have different reliability. Failure to further consider these contextual factors may lead to misjudgments of the operator's intentions, thereby affecting the accuracy and efficiency of human-computer interaction.
[0082] In response, this application further proposes a method for adjusting the weights of interactive input information related to the head pose fine-tuning action, which includes: Obtain the task priority of the precision physical operation currently being performed by the operator; Obtain the cognitive load status of the operator currently performing multitasking; Based on the task priority, determine the importance of the head posture fine-tuning action in the current task; Based on the cognitive load state, assess the operator's control accuracy of the head posture fine-tuning movements; The weight of the head posture fine-tuning action is adjusted according to the importance and control precision.
[0083] Specifically, acquiring the task priority of the precision physical operation currently being performed by the operator refers to the system acquiring the urgency, importance, and precision requirements of the task being performed by the operator. For example, task priorities can be divided into three levels: high, medium, and low. High-priority tasks may involve safety-critical operations or high-precision assembly, while low-priority tasks may only be routine checks or data entry. The purpose is to provide a basis for subsequently assessing the importance of fine-tuning head posture movements.
[0084] The acquisition of the cognitive load status of the operator's current multitasking switching can be understood as assessing the psychological stress and attention allocation experienced during multitasking by monitoring the operator's physiological or behavioral indicators. For example, cognitive load status can be assessed based on physiological signals such as heart rate variability data, eye movement data, or electroencephalogram (EEG) data, or indirectly by analyzing the frequency and complexity of task switching. The aim is to quantify the operator's psychological resource consumption at a specific moment, thereby indirectly reflecting the reliability of their operation.
[0085] In practical applications, determining the importance of the head pose fine-tuning action in the current task based on task priority refers to matching the acquired task priority with preset rules or models to determine the criticality of the fine-tuning action for completing the current task. For example, in high-precision calibration tasks, minor head pose adjustments may be considered highly important instructions; while in non-critical observation tasks, their importance is relatively low. The purpose is to distinguish the semantic value of fine-tuning actions in different task contexts.
[0086] Furthermore, assessing the operator's control accuracy of the head posture fine-tuning movements based on cognitive load refers to predicting the accuracy and stability of the operator's execution of these movements by considering their current cognitive load level. For example, when the operator's cognitive load is high, their control over fine motor skills may decrease, leading to reduced accuracy in the fine-tuning movements; conversely, when the cognitive load is low, the control accuracy may be higher. The purpose is to evaluate the reliability of the fine-tuning movements as effective instructions.
[0087] Therefore, adjusting the weight of the head posture fine-tuning action based on the importance and control precision means dynamically adjusting the influence of the head posture fine-tuning action in judging the operator's intention by comprehensively considering the two evaluation results. For example, when the importance and control precision are high, the weight of the fine-tuning action will be significantly increased; when the importance and control precision are low, its weight may be reduced, or even considered noise. The purpose is to enable the system to understand the operator's true intention more intelligently and accurately.
[0088] This application's solution addresses the limitations of basic solutions in adjusting the weights of interactive input information by introducing two key contextual pieces of information: task priority and cognitive load state. It is precisely because operators exhibit varying degrees of emphasis and precision in their fine-tuning actions across different task scenarios that it is difficult to accurately determine their intent based solely on the fine-tuning action itself.
[0089] This solution first obtains the priority of the operator's current task, thereby objectively determining the semantic importance of head posture fine-tuning actions within that task. For example, when performing a task requiring extremely high precision, even a tiny head posture adjustment may represent a precise instruction from the operator, thus its importance should be assigned a high value. Secondly, by acquiring the operator's cognitive load state, this solution can assess the operator's control precision of head posture fine-tuning actions in the current state. When the operator is under high cognitive load, their ability to control fine movements may decrease; at this time, fine-tuning actions may be more of a physiological reaction than a precise instruction, thus their control precision assessment value will be lower. Finally, the system comprehensively considers these two assessment results—importance and control precision—to dynamically adjust the weight of head posture fine-tuning actions. This context-based weight adjustment mechanism enables the system to more intelligently distinguish between the operator's compensatory behavior and micro-operation instructions, thereby avoiding misjudgments caused by a single recognition result and significantly improving the accuracy and robustness of human-computer interaction.
[0090] Through the above technical solution, this application enables more refined and intelligent adjustment of the weights of operators' head posture fine-tuning movements. Compared to adjusting weights solely based on the recognition results of fine-tuning behaviors, this solution introduces task priority and cognitive load status, allowing the system to more accurately understand the operator's true intentions in specific situations. This effectively avoids interaction misjudgments caused by changes in task importance or operator cognitive state in complex and ever-changing work environments, thereby significantly improving the accuracy, reliability, and user experience of multimodal human-computer interaction in AI glasses augmented reality environments. Especially when high-precision operations are required or operators are in high-pressure, multi-tasking environments, this solution ensures that the system responds correctly to the operator's subtle but crucial interaction commands, thereby improving the overall system performance and safety.
[0091] In some embodiments of this application, in order to adjust the weights of interactive input information related to head posture fine-tuning movements, it is necessary to obtain the cognitive load state of the operator's current multitasking switching. However, simply obtaining this cognitive load state may not fully reflect the operator's true cognitive load level in a complex multitasking environment, especially in high cognitive load scenarios such as task switching. This could lead to inaccurate assessment of the operator's control precision, thereby affecting the reasonable adjustment of the interactive input information weights. Therefore, this application further proposes a method for obtaining the cognitive load state of the operator's current multitasking switching to improve the accuracy and real-time performance of cognitive load assessment.
[0092] Specifically, obtaining the cognitive load status of the operator's current multitasking switching includes: Acquire the operator's physiological signals, including heart rate variability data, eye movement data, and electroencephalogram (EEG) data; Obtain the task identifiers of the multiple tasks currently being performed by the operator; Based on the task identifier, the cognitive load characteristic parameters of each task are extracted from the preset task characteristic database. The cognitive load characteristic parameters include the complexity of the task, the urgency of the task, and the interaction mode of the task. Based on the physiological signals and the cognitive load characteristic parameters, a preliminary assessment of the operator's cognitive load status is conducted to obtain preliminary assessment results; Determine whether the operator is switching tasks; When it is determined that the operator is switching tasks, the weight of the preliminary assessment results is adjusted according to the time of the task switch.
[0093] The physiological signals referred to here are bioelectrical signals or physiological parameters that can reflect changes in the operator's physiological state, such as heart rate variability data, eye movement data, and electroencephalogram (EEG) data. Heart rate variability data can reflect the activity of the autonomic nervous system and is related to stress and cognitive load levels; eye movement data, such as pupil size, blink frequency, and fixation point changes, can reveal visual attention and cognitive effort; EEG data, especially EEG activity in different frequency bands, can directly reflect the brain's cognitive state and activity intensity. These physiological signals are collected in real time by sensors integrated into the AI glasses, with the aim of providing objective indicators of the operator's internal cognitive state.
[0094] The task identifiers are used to uniquely identify the tasks currently being performed by the operator. In practical applications, these task identifiers can be automatically identified and assigned by the AI glasses system based on the operator's interaction behavior, the environment, or a preset workflow. The task characteristic database is a pre-built knowledge base that stores cognitive load characteristic parameters for various tasks, such as task complexity, task urgency, and task interaction mode. Task complexity can be quantified as the number of steps, decision points, or information processing volume required to complete the task; task urgency can be defined based on the task's time limit or its impact on system stability; and task interaction mode describes the main interaction methods relied upon by the task, such as voice, gestures, or eye contact. By querying this database, a cognitive load benchmark based on the task itself can be provided for the current task, aiming to provide task-level contextual information for cognitive load assessment.
[0095] The preliminary assessment results are obtained through comprehensive analysis of physiological signals and task characteristic parameters. For example, a machine learning model can be used, taking physiological signal features (such as heart rate variability, eye movement parameters, and EEG power spectral density) and task characteristic parameters as input, and outputting a quantified cognitive load score. This preliminary assessment aims to provide a comprehensive, multi-dimensional, integrated cognitive load baseline. Furthermore, determining whether an operator is switching tasks can be done by monitoring changes in the operator's interaction patterns (e.g., switching from one application interface to another, or changing from executing one sequence of operations to executing another), the content of voice commands, or significant shifts in gaze focus. When a task switch is detected, the weight of the preliminary assessment results is adjusted based on the timing of the task switch. For example, within a preset time window after a task switch, the weight of the cognitive load assessment results can be appropriately increased to reflect the instantaneous increase in cognitive load during the task switch, aiming to capture the impact of this high cognitive load event on the operator's state.
[0096] This application's solution constructs a more refined and real-time cognitive load assessment mechanism by integrating the operator's physiological signals, the inherent cognitive load characteristics of the current task, and the dynamic perception of task switching events. Physiological signals directly reflect the operator's internal state, while task characteristic parameters provide the cognitive demand context of the task itself. More importantly, by identifying task switching and dynamically adjusting the assessment weights, this application can capture the instantaneous fluctuations in the operator's cognitive load when switching between different tasks, which is crucial for understanding the operator's true cognitive state in complex environments. This multi-dimensional and dynamic assessment method makes the assessment of the operator's control precision more accurate, thus providing a more reliable basis for adjusting the weights of subsequent interactive input information.
[0097] The aforementioned technical solution enables more accurate acquisition of the operator's cognitive load status in multi-task switching scenarios. This precise cognitive load assessment allows for a more comprehensive consideration of the operator's actual cognitive burden when adjusting the weights of interactive input information related to head posture fine-tuning movements. This prevents misoperations or compensatory behaviors caused by excessive cognitive load from being misidentified as intentional commands. Consequently, the robustness and intelligence level of the multimodal human-computer interaction system in AI glasses augmented reality environments are improved, enabling the system to more intelligently understand the operator's true intentions. This is particularly evident in high-precision, high-load work scenarios, significantly enhancing the efficiency and safety of human-computer interaction.
[0098] In some embodiments described above, this application proposes acquiring physiological signals of operators to assess their cognitive load status. These physiological signals include heart rate variability data, eye movement data, and electroencephalogram (EEG) data. However, in practical applications, the acquisition of physiological signals is susceptible to interference from various factors, such as environmental noise, poor sensor contact, or operator physiological activities, leading to decreased signal quality and consequently affecting the accuracy of cognitive load assessment. If these problems are not addressed, cognitive load assessments based on low-quality physiological signals may not accurately reflect the operator's actual state, resulting in biased judgments of operator intentions by the multimodal human-computer interaction system and reduced reliability and efficiency of the interaction. Therefore, this application further proposes a method to optimize the physiological signal acquisition process to ensure the reliability of the acquired physiological signals, thereby improving the accuracy of cognitive load assessment.
[0099] The acquisition of the operator's physiological signals includes heart rate variability data, eye movement data, and electroencephalogram (EEG) data, including: The signal quality parameters of each physiological signal acquisition channel are continuously monitored, including the signal-to-noise ratio, baseline drift, and signal integrity. When the signal quality parameter of any physiological signal acquisition channel is lower than the preset quality threshold, the abnormal channel is identified. When an abnormality in the signal is detected, real-time compensation processing of the physiological signal in the abnormal channel is triggered. The weight of the physiological signals of the abnormal channels in the cognitive load assessment is adjusted according to the type and degree of signal abnormality of the abnormal channels. When multiple physiological signal acquisition channels have signal anomalies, data from other physiological signal channels with better signal quality should be used first for cognitive load assessment. When all physiological signal acquisition channels show abnormal signals that cannot be effectively compensated, a prompt is issued to the operator to replace or check the sensors, and the cognitive load assessment results are temporarily frozen until the signal quality returns to normal.
[0100] Specifically, continuous monitoring of signal quality parameters for each physiological signal acquisition channel refers to the system's real-time analysis of the raw data streams from physiological signal acquisition channels such as heart rate variability, eye movement, and electroencephalogram (EEG). Among these, the signal-to-noise ratio (SNR) quantifies the relative intensity of the effective signal to background noise; baseline drift assesses slow, unexpected deviations from the signal baseline; and signal integrity refers to the presence of missing, interrupted, or severe artifacts in the signal data. Preset quality thresholds are the minimum acceptable standards set for each signal quality parameter. When the signal quality parameter of any acquisition channel falls below its corresponding preset quality threshold, that channel is identified as an abnormal channel.
[0101] When an anomaly is detected, the system triggers real-time compensation processing for the physiological signals of the abnormal channel. This compensation processing can be adaptively adjusted according to the type and severity of the signal anomaly. For example, low-pass filtering can be used for high-frequency noise interference; high-pass filtering or baseline correction algorithms can be used for baseline drift; and wavelet denoising or independent component analysis (ICA) can be used to remove transient artifacts. After compensation processing, the weight of the physiological signal of the abnormal channel in the cognitive load assessment is adjusted according to the type and severity of the signal anomaly. For example, the weight of a channel with a slight anomaly that is successfully compensated may be slightly reduced; while the weight of a channel with a severe anomaly or poor compensation effect will be significantly reduced, or its data may even be temporarily excluded.
[0102] Furthermore, when multiple physiological signal acquisition channels simultaneously exhibit signal anomalies, the system prioritizes using data from other physiological signal channels with better signal quality for cognitive load assessment. This means the system possesses a certain degree of redundancy and robustness; even if some sensor data quality is poor, assessment can still be performed using other reliable data sources. As a final safeguard, when all physiological signal acquisition channels exhibit signal anomalies that cannot be effectively recovered through real-time compensation, the system will prompt the operator to replace or inspect the sensors and temporarily freeze the cognitive load assessment results until signal quality returns to normal, thus avoiding erroneous judgments based on unreliable data.
[0103] This application's solution continuously monitors the signal quality parameters of the physiological signal acquisition channels and identifies abnormal channels based on preset quality thresholds, thereby enabling timely detection and handling of various problems that occur during the physiological signal acquisition process. When a signal abnormality is detected, the system can trigger real-time compensation processing to effectively remove or reduce interference and artifacts in the signal, ensuring high reliability of the physiological signals input to the cognitive load assessment module. Furthermore, by adjusting signal weights according to the type and severity of the abnormality, and prioritizing the use of high-quality data when multiple channels are abnormal, this application's solution maximizes the utilization of effective information and avoids the negative impact of low-quality data on the cognitive load assessment results. When all channels cannot be effectively compensated, the system will promptly alert the operator and pause the assessment, fundamentally eliminating the risk of making decisions based on erroneous data.
[0104] Through the above technical solution, this application significantly improves the reliability and accuracy of physiological signal acquisition in multimodal human-computer interaction methods in AI glasses augmented reality environments. This solution effectively addresses common problems such as noise, drift, and artifacts during physiological signal acquisition, ensuring the quality of input data for cognitive load assessment. Therefore, cognitive load assessment based on high-quality physiological signals will more accurately reflect the operator's true cognitive state, enabling the AI glasses system to more precisely adjust the weight of interactive input information. This achieves more intelligent and adaptable multimodal human-computer interaction, greatly improving system robustness and user experience.
[0105] Specifically, when an abnormal signal is detected, real-time compensation processing of the physiological signals in the abnormal channel can include the following steps: Obtain the physiological signals of the abnormal channel; Spectral analysis is performed on the physiological signals of the abnormal channels to obtain the spectral characteristics of the physiological signals; Based on the spectral characteristics of the physiological signal, high-frequency interference components in the physiological signal are identified. Analyze the frequency characteristics of the high-frequency interference components; Based on the frequency characteristics of the high-frequency interference components, match the type and operating mode of the interference source; Adjust the center frequency and bandwidth parameters of the adaptive filter according to the type and operating mode of the interference source; The dynamically changing interference noise in the physiological signal is removed using the adjusted adaptive filter. The waveform features of the physiological signal after removing the dynamically changing interference noise are compared to obtain the comparison results of the waveform features; Based on the comparison results of the waveform characteristics, the amplitude or phase of the physiological signal is adjusted.
[0106] Acquiring the physiological signals from the abnormal channels refers to reading the raw physiological data stream from physiological signal acquisition channels identified as having signal quality below a preset quality threshold. The physiological signals may include heart rate variability data, eye movement data, and electroencephalogram (EEG) data, etc.
[0107] Furthermore, spectral analysis is performed on the physiological signals from the abnormal channels to obtain the spectral characteristics of the physiological signals. Specifically, signal processing techniques such as Fast Fourier Transform (FFT) can be used to convert the time-domain physiological signals into a frequency-domain representation. The spectral characteristics may include the signal's power spectral density, dominant frequency, and harmonic components.
[0108] Identifying high-frequency interference components in the physiological signal based on its spectral characteristics and analyzing their frequency characteristics involves analyzing the spectrum to identify high-frequency components with specific frequency ranges or peaks that do not conform to the normal physiological signal pattern. These frequency characteristics may include parameters such as the center frequency, bandwidth, and amplitude of these interference components.
[0109] In practical applications, matching the type and operating mode of the interference source based on the frequency characteristics of the high-frequency interference components can be done based on a pre-established interference source database or expert rules. For example, narrowband interference at a specific frequency might be matched to power line noise, while broadband random interference might be matched to electromagnetic interference or motion artifacts. The operating mode can indicate whether the interference is continuous, intermittent, or pulsed.
[0110] Based on this, adjusting the center frequency and bandwidth parameters of the adaptive filter according to the type and operating mode of the interference source means dynamically configuring an adaptive filter based on the identified interference characteristics. For example, if a 50Hz power line interference is identified, the center frequency of the filter is adjusted to 50Hz, and an appropriate bandwidth is set to effectively suppress this frequency component. The adaptive filter can employ techniques such as the Least Mean Square (LMS) algorithm, Kalman filtering, or Wiener filtering.
[0111] Subsequently, the dynamically changing interference noise in the physiological signal is removed according to the adjusted adaptive filter. This is done by applying the configured adaptive filter to the physiological signal data stream to filter out the interference noise that changes over time in real time, thereby obtaining a purer physiological signal.
[0112] The waveform characteristics of the physiological signal after removing the dynamic interference noise are compared to obtain the waveform characteristic comparison result. This means comparing the waveform characteristics (e.g., peak amplitude, waveform duration, morphological characteristics, etc.) of the filtered physiological signal with the expected normal physiological waveform characteristics or the original signal to evaluate the effect of noise removal and whether there is any new distortion.
[0113] Finally, based on the comparison results of the waveform characteristics, the amplitude or phase of the physiological signal is adjusted. This means that if the comparison results show that there is still a deviation or distortion in the amplitude or phase, further fine-tuning is performed to ensure that the compensated physiological signal can accurately reflect the operator's real physiological activity state.
[0114] The proposed solution employs refined spectral analysis of the physiological signals from abnormal channels to accurately identify and analyze the frequency characteristics of high-frequency interference components, thereby matching the specific interference source type and operating mode. This allows for targeted adjustment of the adaptive filter parameters, effectively removing dynamically changing interference noise. Furthermore, by comparing the waveform characteristics of the compensated physiological signals and adjusting their amplitude or phase, the accuracy of the compensation results and the authenticity of the physiological signals are ensured.
[0115] The above technical solutions can significantly improve the compensation accuracy and real-time performance of abnormal channel physiological signals, effectively restoring the quality of interfered physiological signals. This provides a more reliable data foundation for subsequent cognitive load assessment, thereby enhancing the overall robustness and accuracy of multimodal human-computer interaction methods in AI glasses augmented reality environments. This enables the system to accurately determine the operator's intentions and make reasonable adjustments to interaction weights even in complex and changing environments.
[0116] refer to Figure 2 , Figure 2 This is a schematic diagram of the structure of a multimodal human-computer interaction system for AI glasses augmented reality environments provided in an embodiment of this application, which includes: The input terminal is used to acquire interactive input information from the operator, including gaze position information, hand movement information, and voice command information; acquire parameters related to the operating environment and the user's own state of the AI glasses, including environmental parameters and state information of motion sensing components; and acquire information about the operator's work stage. The evaluation end is used to perform a reliability evaluation on the interactive input information based on the interactive input information, the environmental parameters, the state information of the motion sensing component, and the working stage information, and to obtain a reliability evaluation result for each type of interactive input information. The adjustment end is used to adjust the weight of each type of interactive input information in judging the operator's intention based on the reliability assessment results.
[0117] This system effectively addresses the issue of decreased accuracy and reliability in multimodal human-computer interaction in augmented reality environments caused by environmental factors and system biases, through the collaborative work of its input, evaluation, and adjustment ends. Specifically, the input end comprehensively collects information on the operator's interaction intent, as well as environmental and system state data affecting system performance. The evaluation end then quantitatively assesses the reliability of different interaction modalities based on this multidimensional information. Finally, the adjustment end dynamically optimizes the contribution weight of each interaction modality in judging the operator's intent based on the evaluation results. As a result, this system can more intelligently and adaptively understand the operator's true intent, significantly improving the accuracy and efficiency of human-computer interaction in complex and precise operational scenarios.
[0118] The specific methods for acquiring operator interaction input information, acquiring parameters related to the AI glasses' operating environment and its own state, and acquiring operator work stage information have been described in the above embodiments, and will not be repeated here. It should be emphasized that the system proposed in this application implements these functions through its structured components.
[0119] Specifically, the input terminal can be configured to integrate multiple sensors and interface modules. For example, gaze position information can be acquired through an eye-tracking sensor, hand movement information can be obtained through a depth camera or inertial measurement unit (IMU), and voice command information can be received through a microphone array. Furthermore, environmental parameters can be acquired through independent temperature, humidity, and light sensors, and the status information of motion sensing components can be provided through an internal diagnostic module or calibration unit. Operator work stage information can be obtained through a data interface with the task management system, or input through manual confirmation by the operator on a specific interface. Alternatively, the input terminal can be designed to allow the operator to manually input or confirm some information through a simplified user interface, such as manually selecting the current work stage before the task begins, or manually calibrating certain parameters when prompted by the system.
[0120] Furthermore, the evaluation end can be implemented as a data processing module that receives multi-source data from the input end. This evaluation end can use a preset rule set or simple threshold judgment logic to evaluate the reliability of the interactive input information. For example, a fixed temperature threshold can be set; when the ambient temperature exceeds this threshold, the reliability of the gaze position information is considered reduced; or, when the drift of the motion sensing component exceeds a certain fixed upper limit, the reliability of the hand movement information is reduced. In some embodiments, the evaluation end can also directly obtain the corresponding interactive input information reliability evaluation result by consulting a pre-stored reliability lookup table, based on the current environmental parameters and the state information of the motion sensing component.
[0121] Therefore, the adjustment terminal can be configured as a weight management module, which adjusts the weights of each interactive input information based on the reliability assessment results provided by the evaluation terminal. For example, the adjustment terminal can use a simple linear mapping relationship to directly convert the reliability assessment results into weight values, or when the reliability assessment results are below a certain fixed threshold, the weight of the corresponding interactive input information can be directly set to a lower fixed value, and when it is above the threshold, it can be set to a higher fixed value. As a specific implementation, the adjustment terminal can preset a set of weight configuration schemes and select the corresponding weight scheme to apply based on the reliability level (e.g., "high", "medium", "low") output by the evaluation terminal.
[0122] The core innovation of the AI glasses augmented reality multimodal human-computer interaction system proposed in this application lies in the introduction of structured components such as input, evaluation, and adjustment terminals, which enables reliable assessment and dynamic weight adjustment of multimodal interactive input information. In traditional systems, the fusion of multimodal interactive information often adopts fixed weights or simple rules based on experience, lacking the ability to adapt to environmental changes and the system's own state. As a result, in complex and precise operation scenarios, such as optical alignment tasks in semiconductor manufacturing, the system has difficulty accurately distinguishing the operator's true intentions, which can easily lead to misoperations.
[0123] In contrast, this system comprehensively perceives environmental, system status, and operator interaction information through the input terminal, and quantitatively analyzes the reliability of this information through the evaluation terminal. Finally, the adjustment terminal intelligently adjusts the weights of each interaction modality based on the analysis results. This modular design allows the system to adapt more flexibly and accurately to dynamically changing work environments and operator states, effectively avoiding interaction uncertainties caused by virtual-physical space misalignment or system deviations. Therefore, this system significantly improves the accuracy and reliability of human-computer interaction in augmented reality environments, providing more stable and efficient interaction support for high-precision tasks.
[0124] The above description is merely an embodiment of this application and is not intended to limit the scope of protection of this application. Various modifications and variations can be made to this application by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the scope of protection of this application.
Claims
1. A multimodal human-computer interaction method for AI glasses in augmented reality environments, characterized in that, include: The system acquires interactive input information from the operator, including eye position information, hand gesture information, and voice command information. Acquire parameters related to the operating environment and the state of the AI glasses, including environmental parameters and state information of motion sensing components; Obtain the work stage information of the operator; Based on the interactive input information, the environmental parameters, the state information of the motion sensing component, and the working stage information, the reliability of the interactive input information is evaluated to obtain the reliability evaluation result of each type of interactive input information. Based on the reliability assessment results, adjust the weight of each type of interactive input information when determining the operator's intention.
2. The multimodal human-computer interaction method for AI glasses augmented reality environments according to claim 1, characterized in that, The adjustment of the weight of each type of interactive input information in determining the operator's intent includes: Obtain the time stamp of the laser pulse occurrence; Acquire the operator's head posture fine-tuning action characteristics and gaze action characteristics; Determine the first temporal correlation between the head posture fine-tuning action and the laser pulse, and identify whether the head posture fine-tuning action is a compensation behavior based on the first temporal correlation, and identify whether the head posture fine-tuning action is a micro-operation command. Determine the second temporal correlation between the gaze action and the laser pulse, and identify whether the gaze action is a compensation behavior based on the second temporal correlation, and identify whether the gaze action is a micro-operation command.
3. The multimodal human-computer interaction method for AI glasses augmented reality environments according to claim 1, characterized in that, The adjustment of the weight of each type of interactive input information in determining the operator's intent includes: The head posture fine-tuning action features are obtained, including the amplitude, speed and duration of the head posture fine-tuning action. The gaze movement features are obtained, including the amplitude, speed, and duration of the gaze movement. Obtain the time stamp of the laser pulse occurrence; Determine the first time interval between the occurrence time of the head posture fine-tuning action and the occurrence time of the laser pulse, and the duration of the head posture fine-tuning action; Determine the second time interval between the occurrence time of the gaze movement and the occurrence time of the laser pulse, and the duration of the gaze movement; Based on the first time interval, the second time interval, and the duration threshold, it is determined whether the head posture fine-tuning action is a compensatory behavior, and the weight of the interactive input information related to the head posture fine-tuning action is adjusted according to the recognition result. The system determines whether the gaze action falls within a preset physiological response window based on its duration, and whether the duration of the gaze action is shorter than a duration threshold. It then identifies whether the gaze action is a compensatory behavior and adjusts the weight of interactive input information related to the gaze action based on the identification results.
4. The multimodal human-computer interaction method for AI glasses augmented reality environments according to claim 1, characterized in that, The adjustment of the weight of each type of interactive input information in determining the operator's intent includes: Obtain information on the relative positional changes between the virtual overlay and the physical world reference point; Identify transient visual drift events and extract features from these events; Acquire the operator's head posture fine-tuning movement characteristics and gaze movement characteristics; Determine the third temporal correlation between the head posture fine-tuning action and the instantaneous visual drift event, and determine the fourth temporal correlation between the gaze action and the instantaneous visual drift event; Based on the judgment result of the third time correlation, identify whether the head posture fine-tuning action is a compensation behavior and whether it is a micro-operation command; Based on the judgment result of the fourth time association, identify whether the gaze action is a compensation behavior and whether it is a micro-operation command; Based on the recognition results of the head posture fine-tuning movements and whether the gaze movements are compensation behaviors, the weights of the relevant interactive input information are adjusted.
5. A multimodal human-computer interaction method for AI glasses augmented reality environments according to claim 1, characterized in that, The adjustment of the weight of each type of interactive input information in determining the operator's intent includes: Acquire the operator's head posture fine-tuning movement characteristics and gaze movement characteristics; The system acquires information on various existing system deviations, including persistent spatial misalignment between the virtual overlay and the physical world caused by ambient temperature fluctuations, transient micro-displacement of physical targets caused by laser pulses, and instantaneous visual drift of the virtual overlay caused by local electromagnetic interference or airflow. Obtain the type, occurrence time, intensity, and duration characteristics of each type of system deviation information; The head posture fine-tuning action features are analyzed for temporal causal correlation with each of the system deviation events to obtain the fifth temporal correlation between the head posture fine-tuning action and the system deviation. Based on the fifth temporal correlation, the first matching score between each head posture fine-tuning action and each system deviation is calculated. The gaze movement features are analyzed for temporal causal correlation with each of the system deviation events to obtain a sixth temporal correlation between the gaze movement and the system deviation. A second matching score between each gaze movement and each system deviation is calculated based on the sixth temporal correlation. Based on the first matching score, the head posture fine-tuning action is associated with each system deviation to identify whether the head posture fine-tuning action is a compensation behavior and whether it is a micro-operation command, and the weight of the interactive input information related to the head posture fine-tuning action is adjusted. Based on the second matching score, the gaze action is associated with each system deviation to identify whether the gaze action is a compensation behavior and whether it is a micro-operation command, and the weight of the interactive input information related to the gaze action is adjusted.
6. A multimodal human-computer interaction method for AI glasses augmented reality environments according to claim 5, characterized in that, The adjustment of the weights of the interactive input information related to the head posture fine-tuning action includes: Obtain the task priority of the precision physical operation currently being performed by the operator; Obtain the cognitive load status of the operator currently performing multitasking; Based on the task priority, determine the importance of the head posture fine-tuning action in the current task; Based on the cognitive load state, assess the operator's control accuracy of the head posture fine-tuning movements; The weight of the head posture fine-tuning action is adjusted according to the importance and control precision.
7. A multimodal human-computer interaction method for AI glasses augmented reality environments according to claim 6, characterized in that, The acquisition of the cognitive load status of the operator's current multitasking switching includes: Acquire the operator's physiological signals, including heart rate variability data, eye movement data, and electroencephalogram (EEG) data; Obtain the task identifiers of the multiple tasks currently being performed by the operator; Based on the task identifier, the cognitive load characteristic parameters of each task are extracted from the preset task characteristic database. The cognitive load characteristic parameters include the complexity of the task, the urgency of the task, and the interaction mode of the task. Based on the physiological signals and the cognitive load characteristic parameters, a preliminary assessment of the operator's cognitive load status is conducted to obtain preliminary assessment results; Determine whether the operator is switching tasks; When it is determined that the operator is switching tasks, the weight of the preliminary assessment results is adjusted according to the time of the task switch.
8. A multimodal human-computer interaction method for AI glasses augmented reality environments according to claim 7, characterized in that, The acquisition of the operator's physiological signals includes heart rate variability data, eye movement data, and electroencephalogram (EEG) data, including: The signal quality parameters of each physiological signal acquisition channel are continuously monitored, including the signal-to-noise ratio, baseline drift, and signal integrity. When the signal quality parameter of any physiological signal acquisition channel is lower than the preset quality threshold, the abnormal channel is identified. When an abnormality in the signal is detected, real-time compensation processing of the physiological signal in the abnormal channel is triggered. The weight of the physiological signals of the abnormal channels in the cognitive load assessment is adjusted according to the type and degree of signal abnormality of the abnormal channels. When multiple physiological signal acquisition channels have signal anomalies, data from other physiological signal channels with better signal quality should be used first for cognitive load assessment. When all physiological signal acquisition channels show abnormal signals that cannot be effectively compensated, a prompt is issued to the operator to replace or check the sensors, and the cognitive load assessment results are temporarily frozen until the signal quality returns to normal.
9. A multimodal human-computer interaction method for AI glasses in augmented reality environments according to claim 8, characterized in that, When an abnormal signal is detected, real-time compensation processing of the physiological signal in the abnormal channel is triggered, including: Obtain the physiological signals of the abnormal channel; Spectral analysis is performed on the physiological signals of the abnormal channels to obtain the spectral characteristics of the physiological signals; Based on the spectral characteristics of the physiological signal, high-frequency interference components in the physiological signal are identified. Analyze the frequency characteristics of the high-frequency interference components; Based on the frequency characteristics of the high-frequency interference components, match the type and operating mode of the interference source; Adjust the center frequency and bandwidth parameters of the adaptive filter according to the type and operating mode of the interference source; The dynamically changing interference noise in the physiological signal is removed using the adjusted adaptive filter. The waveform features of the physiological signal after removing the dynamically changing interference noise are compared to obtain the comparison results of the waveform features; Based on the comparison results of the waveform characteristics, the amplitude or phase of the physiological signal is adjusted.
10. A multimodal human-computer interaction system for AI glasses augmented reality environments, characterized in that: include: The input terminal is used to acquire interactive input information from the operator, including gaze position information, hand movement information, and voice command information; acquire parameters related to the operating environment and the user's own state of the AI glasses, including environmental parameters and state information of motion sensing components; and acquire information about the operator's work stage. The evaluation end is used to perform a reliability evaluation on the interactive input information based on the interactive input information, the environmental parameters, the state information of the motion sensing component, and the working stage information, and to obtain a reliability evaluation result for each type of interactive input information. The adjustment end is used to adjust the weight of each type of interactive input information in judging the operator's intention based on the reliability assessment results.