Multi-modal intelligent glasses interaction control method based on fusion of electroencephalogram and eye movement
Patent Information
- Application Number
- CN202610886400.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-18
- Publication Date
- 2026-09-29
AI Technical Summary
在需要保持安静或隐蔽的场合(例如,会议室)下,用户无法使用语音或手势实现自然、流畅的操作,导致用户体验感较差
[0011]本公开的上述各个实施例具有如下有益效果:通过本公开的一些实施例的一种基于脑电和眼动融合的多模态智能眼镜交互控制方法,可以降低计算资源的消耗且缩短控制指令识别的耗时,进而提升交互控制的实时性,从而提升整体交互的体验感。造成计算资源的消耗较大且控制指令识别的耗时较长,交互控制的实时性较差,整体交互的体验感较差的原因在于:除眼动追踪技术外,其他技术方案(例如,语音、手势等)均容易被他人感知。在需要保持安静或隐蔽的场合(例如,会议室)下,用户无法使用语音或手势实现自然、流畅的操作,导致用户体验感较差。此外,现有的组合交互控制方案需要同时处理大量多模态数据,导致计算资源的消耗较大,控制指令识别的耗时较长,交互控制的实时性较差。同时,现有方案无法根据不同使用场景动态调整多模态数据的优先级,导致整体交互体验较差。基于此,本公开的一些实施例的基于脑电和眼动融合的多模态智能眼镜交互控制方法,首先,根据所获取的原始脑电信号,生成脑电意图识别结果。由此,可以得到脑电意图的识别结果。然后,根据所获取的眼动图像帧序列,生成眼动意图识别结果。由此,可以得到眼动意图的识别结果。之后,根据所获取的声音信号,生成音频感知特征集。由此,可以得到各个音频感知特征。其次,根据所获取的视频图像信号,生成视觉感知特征集。由此,可以得到各个视觉感知特征。之后,根据上述当前场景类别识别结果,确定目标优先级融合规则。由此,可以确定不同场景下各模态数据的优先级。接着,根据上述目标优先级融合规则、上述脑电意图识别结果、上述眼动意图识别结果、上述语音意图识别结果和上述手势意图识别结果,生成当前执行控制指令。由此,可以得到当前时刻的执行控制指令。最后,根据上述当前执行控制指令,执行对应上述当前执行控制指令的交互操作。也因为不是对所有原始传感器获得的数据进行无差别的融合处理,而是先分别提取脑电意图识别结果、眼动意图识别结果、语音意图识别结果和手势意图识别结果,从而减少了后续融合阶段需要处理的数据量。也不是采用单一固定的优先级规则,而是根据当前场景类别识别结果区分不同场景,在不同场景下采用不同的信号优先级顺序,从而可以兼顾不同场景下的隐蔽性和交互效率。还因为是根据目标优先级融合规则从多个意图识别结果中确定最高优先级的识别结果并生成当前执行控制指令,从而仅需处理优先级最高的单个模态的识别结果即可完成交互控制指令的生成。因此,可以避免同时处理大量多模态数据所带来的计算资源浪费,从而降低计算资源的消耗且缩短控制指令识别的耗时,进而提升交互控制的实时性,同时通过场景自适应的优先级区分策略提升整体交互的体验感。
Smart Images

Figure CN122837628A_ABST
Abstract
Description
Technical Field
[0001] The embodiments disclosed herein relate to the field of human-computer interaction technology, specifically to a multimodal smart glasses interaction control method based on EEG and eye-tracking fusion. Background Technology
[0002] As wearable devices, smart glasses integrate environmental perception, data processing, and information interaction, allowing users to use them in various scenarios such as social interactions and meetings to achieve natural and efficient interactive control. Currently, the common methods for interactive control using smart glasses are: using dynamic eye-tracking technology, or combining eye-tracking with intelligent voice, gesture recognition, head movements, and external sensors to achieve interactive control.
[0003] However, in practice, it has been found that when using smart glasses for interactive control in the above manner, the following technical problems often occur: Aside from eye-tracking technology, other technologies (such as voice and gestures) are easily perceived by others. In quiet or discreet settings (e.g., meeting rooms), users cannot use voice or gestures to achieve natural and smooth operations, resulting in a poor user experience. Furthermore, existing combined interactive control solutions require processing large amounts of multimodal data simultaneously, leading to significant computational resource consumption, long command recognition times, and poor real-time performance of interactive control. Additionally, existing solutions cannot dynamically adjust the priority of multimodal data according to different usage scenarios, resulting in an overall poor interactive experience.
[0004] The information disclosed in this background section is only intended to enhance the understanding of the background of the inventive concept, and therefore may contain information that does not constitute prior art known to those skilled in the art. Summary of the Invention
[0005] The summary portion of this disclosure is intended to provide a brief overview of the concepts, which will be described in detail in the detailed description portion. This summary portion is not intended to identify key or essential features of the claimed technical solutions, nor is it intended to limit the scope of the claimed technical solutions.
[0006] Some embodiments of this disclosure propose a multimodal smart glasses interactive control method, device, electronic device, and computer-readable medium based on EEG and eye-tracking fusion to solve one or more of the technical problems mentioned in the background section above.
[0007] In a first aspect, some embodiments of this disclosure provide a multimodal smart glasses interaction control method based on EEG and eye-tracking fusion. The method includes: generating an EEG intention recognition result based on acquired raw EEG signals; generating an eye-tracking intention recognition result based on acquired eye-tracking image frame sequences; generating an audio perception feature set based on acquired sound signals, wherein the audio perception feature set includes a speech intention recognition result and a first user behavior state signal; generating a visual perception feature set based on acquired video image signals, wherein the visual perception feature set includes a gesture intention recognition result, a current scene category recognition result, and a second user behavior state signal; determining a target priority fusion rule based on the current scene category recognition result; generating a current execution control command based on the target priority fusion rule, the EEG intention recognition result, the eye-tracking intention recognition result, the speech intention recognition result, and the gesture intention recognition result; and executing an interactive operation corresponding to the current execution control command based on the current execution control command.
[0008] Secondly, some embodiments of this disclosure provide a multimodal smart glasses interactive control device based on EEG and eye-tracking fusion, comprising: a first generation unit configured to generate an EEG intention recognition result based on acquired raw EEG signals; a second generation unit configured to generate an eye-tracking intention recognition result based on acquired eye-tracking image frame sequences; a third generation unit configured to generate an audio perception feature set based on acquired sound signals, wherein the audio perception feature set includes a speech intention recognition result and a first user behavior state signal; a fourth generation unit configured to generate a visual perception feature set based on acquired video image signals, wherein the visual perception feature set includes a gesture intention recognition result, a current scene category recognition result, and a second user behavior state signal; a determination unit configured to determine a target priority fusion rule based on the current scene category recognition result; a fifth generation unit configured to generate a current execution control command based on the target priority fusion rule, the EEG intention recognition result, the eye-tracking intention recognition result, the speech intention recognition result, and the gesture intention recognition result; and an execution unit configured to execute an interactive operation corresponding to the current execution control command.
[0009] Thirdly, some embodiments of this disclosure provide an electronic device, including: one or more processors; and a storage device having one or more programs stored thereon, wherein when the one or more programs are executed by the one or more processors, the one or more processors implement the method described in any implementation of the first aspect above.
[0010] Fourthly, some embodiments of this disclosure provide a computer-readable medium having a computer program stored thereon, wherein the computer program, when executed by a processor, implements the method described in any of the implementations of the first or second aspect.
[0011] The various embodiments disclosed herein have the following beneficial effects: Through a multimodal smart glasses interaction control method based on EEG and eye-tracking fusion according to some embodiments of this disclosure, the consumption of computing resources can be reduced and the time for recognizing control commands can be shortened, thereby improving the real-time performance of interactive control and thus enhancing the overall interactive experience. The reason for the high consumption of computing resources, long time for recognizing control commands, poor real-time performance of interactive control, and poor overall interactive experience is that, apart from eye-tracking technology, other technical solutions (e.g., voice, gestures, etc.) are easily perceived by others. In situations requiring quiet or privacy (e.g., a conference room), users cannot use voice or gestures to achieve natural and smooth operation, resulting in a poor user experience. Furthermore, existing combined interactive control solutions require processing a large amount of multimodal data simultaneously, leading to high consumption of computing resources, long time for recognizing control commands, and poor real-time performance of interactive control. At the same time, existing solutions cannot dynamically adjust the priority of multimodal data according to different usage scenarios, resulting in a poor overall interactive experience. Based on this, some embodiments of the multimodal smart glasses interaction control method based on EEG and eye-tracking fusion disclosed herein first generate EEG intention recognition results based on the acquired raw EEG signals. This yields the EEG intention recognition result. Then, based on the acquired eye-tracking image frame sequence, generate eye-tracking intention recognition results. This yields the eye-tracking intention recognition result. Next, based on the acquired sound signals, generate an audio perception feature set. This yields various audio perception features. Secondly, based on the acquired video image signals, generate a visual perception feature set. This yields various visual perception features. Then, based on the aforementioned current scene category recognition results, determine a target priority fusion rule. This determines the priority of each modality of data in different scenes. Next, based on the aforementioned target priority fusion rule, the aforementioned EEG intention recognition results, the aforementioned eye-tracking intention recognition results, the aforementioned voice intention recognition results, and the aforementioned gesture intention recognition results, generate a current execution control command. This yields the execution control command at the current moment. Finally, based on the aforementioned current execution control command, execute the interactive operation corresponding to the aforementioned current execution control command. Because it doesn't indiscriminately fuse all the data from the original sensors, but instead extracts the EEG intention recognition results, eye-tracking intention recognition results, voice intention recognition results, and gesture intention recognition results separately, the amount of data that needs to be processed in the subsequent fusion stage is reduced. It also doesn't use a single fixed priority rule, but distinguishes different scenarios based on the current scene category recognition results, using different signal priority orders in different scenarios, thus balancing concealment and interaction efficiency in different scenarios. Furthermore, because it determines the highest priority recognition result from multiple intention recognition results based on target priority fusion rules and generates the current execution control command, it only needs to process the recognition result of the highest priority single modality to complete the generation of the interaction control command.Therefore, it can avoid the waste of computing resources caused by processing a large amount of multimodal data at the same time, thereby reducing the consumption of computing resources and shortening the time for control command recognition, thus improving the real-time performance of interactive control, while improving the overall interactive experience through scene-adaptive priority differentiation strategy. Attached Figure Description
[0012] The above and other features, advantages, and aspects of the embodiments of this disclosure will become more apparent from the accompanying drawings and the following detailed description. Throughout the drawings, the same or similar reference numerals denote the same or similar elements. It should be understood that the drawings are schematic, and elements are not necessarily drawn to scale.
[0013] Figure 1 This is a flowchart of some embodiments of the multimodal smart glasses interactive control method based on EEG and eye-tracking fusion according to the present disclosure; Figure 2 This is a schematic diagram of the structure of some embodiments of the multimodal smart glasses interactive control device based on EEG and eye-tracking fusion according to the present disclosure; Figure 3 This is a schematic diagram of the structure of an electronic device suitable for implementing some embodiments of the present disclosure. Detailed Implementation
[0014] Embodiments of this disclosure will now be described in more detail with reference to the accompanying drawings. While some embodiments of this disclosure are shown in the drawings, it should be understood that this disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of this disclosure. It should be understood that the accompanying drawings and embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of protection of this disclosure.
[0015] It should also be noted that, for ease of description, only the parts relevant to the invention are shown in the accompanying drawings. Unless otherwise specified, the embodiments and features described in this disclosure can be combined with each other.
[0016] It should be noted that the concepts of "first" and "second" mentioned in this disclosure are used only to distinguish different devices, modules or units, and are not used to limit the order of functions performed by these devices, modules or units or their interdependencies.
[0017] It should be noted that the terms "a" and "a plurality of" used in this disclosure are illustrative rather than restrictive, and those skilled in the art should understand that, unless otherwise expressly indicated in the context, they should be understood as "one or more".
[0018] The names of messages or information exchanged between multiple devices in the embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of such messages or information.
[0019] This disclosure will now be described in detail with reference to the accompanying drawings and embodiments.
[0020] Figure 1 A flowchart 100 is shown illustrating some embodiments of a multimodal smart glasses interaction control method based on EEG and eye-tracking fusion according to this disclosure. This multimodal smart glasses interaction control method based on EEG and eye-tracking fusion includes the following steps: Step 101: Generate EEG intention recognition results based on the acquired raw EEG signals.
[0021] In some embodiments, the aforementioned executing entity can generate a brainwave intention recognition result based on the acquired raw brainwave signals. The raw brainwave signals can characterize brainwaves collected by a sensor. Here, the specific types of the sensor and the brainwaves are not limited; for example, the sensor can be a non-invasive brain-computer interface sensor, and the brainwaves can be event-correlated desynchronization or event-correlated synchronization brainwave signals in the 8-12 Hz Alpha band and 13-30 Hz Beta band of the MI-BCI signal. The brainwave intention recognition result can characterize the control command that the user wants to execute, identified from the raw brainwave signals. The control command can include, but is not limited to: opening an application, closing a window, turning a page, scrolling, selecting, etc. The user can characterize a person wearing the aforementioned smart glasses. For example, the smart glasses can be AR glasses.
[0022] In some optional implementations of certain embodiments, the aforementioned execution entity can generate brainwave intent recognition results based on the acquired raw brainwave signals through the following steps: The first step is to filter the raw EEG signal to obtain a denoised EEG signal. This denoised EEG signal represents the EEG signal obtained after filtering out noise at a preset frequency from the raw EEG signal. The specific value of the preset frequency is not limited; for example, it could be 50 Hz. In practice, the executing entity can use an infinite impulse response digital filter to filter the raw EEG signal in the 8-13 Hz and 13-30 Hz frequency bands, retaining the signal components in the Alpha and Beta bands while filtering out noise signals outside these bands to obtain the denoised EEG signal.
[0023] The second step involves generating an artifact-free EEG signal based on the aforementioned denoised EEG signal. The artifact-free EEG signal characterizes the EEG signal after removing interference from electrooculography (EOG) or electromyography (EMG). The aforementioned EOG signal characterizes the electrophysiological signals generated by eye movements. The aforementioned EMG signal characterizes the electrophysiological signals generated by muscle contractions.
[0024] The third step involves resampling the artifact-free EEG signal to obtain a resampled EEG signal. This resampled EEG signal characterizes the EEG signal obtained after resampling the artifact-free EEG signal. In practice, the executing entity first uses a downsampling algorithm to reduce the sampling frequency of the artifact-free EEG signal to a second preset frequency. Then, the downsampled EEG signal is identified as the resampled EEG signal. For example, the second preset frequency can be 200 Hz.
[0025] The fourth step involves feature extraction from the resampled EEG signals to obtain EEG feature signals. These feature signals can be feature vectors reflecting the differences between different motor imagery tasks. The motor imagery tasks characterize the mental activity of a user imagining body parts moving without actually performing those movements. For example, a left-hand motor imagery task characterizes the mental activity of a user imagining left-hand movement without actually moving their left hand, and a right-hand motor imagery task characterizes the mental activity of a user imagining right-hand movement without actually moving their right hand. In practice, firstly, the executing entity can calculate a set of spatial filters using a common spatial pattern algorithm. This set of spatial filters characterizes the ratio between the signal variance of the left-hand motor imagery task and the signal variance of the right-hand motor imagery task, maximizing or minimizing the ratio. Then, the resampled EEG signals are spatially filtered using this set of spatial filters to obtain filtered EEG signals. Finally, the variance of the filtered EEG signals is determined as the EEG feature signals.
[0026] The fifth step involves generating an EEG intention recognition result based on the aforementioned EEG feature signals. In practice, firstly, the executing entity can classify the EEG feature signals using multi-class linear discriminant analysis to identify the corresponding motor imagery task category. Then, based on a preset mapping relationship, the control command corresponding to the motor imagery task category is determined as the EEG intention recognition result. This preset mapping relationship characterizes the correspondence between the task category and the EEG intention recognition result. For example, the preset mapping relationship could be "Task category: Imagine left hand movement; EEG intention recognition result: Open application". It should be noted that the task category is not limited to binary classification and can be extended to multi-class scenarios. The task category may include, but is not limited to: imagining left hand movement, imagining right hand movement, imagining foot movement, etc.
[0027] In addressing the technical problems mentioned above, and considering the application scenario—specifically, the acquisition of EEG signals for patients with traumatic brain injury undergoing neurorehabilitation training via a motor imagery brain-computer interface in the intensive care unit (ICU)—often presents a second technical problem: patients with traumatic brain injury frequently experience involuntary nystagmus and facial muscle spasms. These pathological physiological activities generate severely distorted waveforms and fluctuating amplitudes in the EEG signals, resulting in artifacts related to electrooculography (EOG) and electromyography (EMG), significantly reducing the effectiveness of the EEG signals. Existing solutions require data acquisition for several times the normal duration and extensive offline post-processing calculations to remove these dynamically changing pathological artifacts, leading to high computational resource consumption and lengthy recognition time for the patient's motor imagery instructions. To meet the following requirements for this application scenario: accurate differentiation between pathological EOG and EMG abnormalities and normal EEG signals, and rapid instruction recognition under limited computational resources, we have decided to adopt the following solution: In some optional implementations of certain embodiments, the aforementioned execution entity may generate an artifact-free EEG signal based on the aforementioned denoised EEG signal through the following steps: The first step involves performing independent component decomposition (ICD) on the denoised EEG signal to obtain multiple independent components as a set of independent components. Each independent component in this set can represent a statistically independent signal source separated from the denoised EEG signal. In practice, the executing entity can use a fast independent principal component analysis (IPA) algorithm to perform ICD on the denoised EEG signal to obtain multiple independent components as a set of independent components.
[0028] The second step is to perform the following steps for each independent component in the above set of independent components: First, a correlation coefficient is generated based on the aforementioned independent components and typical electrooculogram (EOG) waveform characteristics. The typical EOG waveform characteristics characterize the voltage waveform generated by blinking. The correlation coefficient characterizes the Pearson correlation coefficient between the independent components and the typical EOG waveform characteristics. In practice, the executing entity first determines the covariance between the independent components and the typical EOG waveform characteristics as a first value. Then, the product of the standard deviation of the independent components and the standard deviation of the typical EOG waveform characteristics is determined as a second value. Finally, the ratio of the first value to the second value is determined as the correlation coefficient.
[0029] Then, the aforementioned independent components are subjected to frequency domain transformation to obtain the power spectral density curves corresponding to these independent components. In practice, the aforementioned execution entity can first use the Fast Fourier Transform algorithm to transform the aforementioned independent components from the time domain to the frequency domain to obtain the power spectral density curves.
[0030] Subsequently, based on the aforementioned power spectral density curve and the preset frequency band interval, a power spectral density integral value is generated. The preset frequency band interval characterizes the frequency range of the energy distribution of the surface electromyography signal. This preset frequency band interval can be between 20-200 Hz. In practice, the executing entity can determine the power spectral density integral value as the area of the power spectral density curve integrated within the preset frequency band interval.
[0031] The third step is to determine the obtained correlation coefficients as the first set and the obtained power spectral density integral values as the second set.
[0032] The fourth step is to determine the independent components corresponding to each correlation coefficient greater than the first preset threshold in the first set as the set of independent components of the electrooculogram. Here, the specific value of the first preset threshold is not limited; for example, the first preset threshold can be 0.8.
[0033] The fifth step is to determine the independent components corresponding to each power spectral density integral value greater than the second preset threshold in the second set as the electromyographic independent component set. Here, the specific value of the second preset threshold is not limited; for example, the second preset threshold can be 0.5 microvolts squared per hertz.
[0034] Step 6: Based on the aforementioned set of independent electromyographic (EMG) components, the aforementioned set of independent electrooculography (EOG) components, and the aforementioned set of independent components, generate a set of remaining independent components. This set of remaining independent components represents the set of all independent components obtained after removing the aforementioned sets of independent EMG and EOG components from the aforementioned set of independent components. In practice, firstly, the executing entity can delete all independent components included in the aforementioned set of independent EOG components and all independent components included in the aforementioned set of independent EMG components. Then, the remaining independent components are determined as the set of remaining independent components.
[0035] Step 7: Perform linear recombination on the remaining independent component set to obtain the artifact-free EEG signal. In practice, the executing agent can use inverse independent component analysis to perform linear recombination on the remaining independent component set to obtain the artifact-free EEG signal.
[0036] The above-described technical solution, as an inventive point of this disclosure, addresses technical problem two: "leading to excessive computational resource consumption and long time consumption in recognizing the patient's motor imagery intentions." The reasons for this excessive computational resource consumption and long time consumption in recognizing the patient's motor imagery intentions are as follows: Patients with traumatic brain injury often experience involuntary nystagmus and facial muscle spasms. These pathological physiological activities produce severe waveform distortion and fluctuating amplitude artifacts in electrooculography (EEG) signals, significantly reducing the effectiveness of the EEG signals. Existing solutions require collecting data for several times the normal duration and performing extensive offline post-processing calculations to remove these dynamically changing pathological artifacts, resulting in excessive computational resource consumption and long time consumption in recognizing the patient's motor imagery intentions. Solving these factors can reduce computational resource consumption and shorten the time required to recognize the patient's motor imagery intentions. To achieve this effect, the multimodal smart glasses interactive control method based on EEG and eye-tracking fusion disclosed herein decomposes the denoised EEG signal into multiple independent components in real time as an independent component set, which can separate statistically independent EEG artifacts, EMG artifacts, and clean EEG signals. Furthermore, for each independent component in the independent component set, the Pearson correlation coefficient with typical EEG waveform features and the power spectral density integral value within a specific EMG feature frequency band are simultaneously calculated. An automatic screening and removal mechanism for EEG and EMG independent components is used, thus reducing the computational overhead of extensive offline post-processing. In addition, independent component analysis inverse transform is performed only on the remaining independent component set after removing EEG and EMG independent components to reconstruct the signal, without processing all original independent components, thereby significantly reducing computational resource consumption and compressing command recognition time to the millisecond level. This reduces computational resource consumption and shortens the time required to recognize the patient's motor intentions.
[0037] Step 102: Generate eye movement intention recognition results based on the acquired eye movement image frame sequence.
[0038] In some embodiments, the aforementioned execution entity can generate an eye-tracking intent recognition result based on the acquired eye-tracking image frame sequence. Each eye-tracking image frame in the sequence can represent a digital image containing the user's eye region. The eye-tracking intent recognition result can represent a control command determined based on the user's eye-tracking feature parameters. For example, when the gaze point coordinates are located on an icon and the gaze duration exceeds a preset duration threshold (e.g., 3 seconds), the corresponding eye-tracking intent recognition result is a "select" command. The eye-tracking feature parameters can represent parameters describing the user's eye movements. These parameters may include gaze point coordinates, gaze duration, and / or saccade trajectory parameters. The gaze point coordinates can represent the coordinates of the user's gaze at the point of gaze on the screen or in space. The gaze duration can represent the duration the user's gaze remains at any gaze point. The saccade trajectory parameters can represent the path the user's gaze moves between two gaze points.
[0039] In addressing the technical problems mentioned above, and considering the application scenario—eye-tracking intention recognition for myasthenia gravis patients during home rehabilitation training and daily communication and entertainment—the following technical problem arises: Myasthenia gravis patients often experience ptosis, slow and uncoordinated eye movements due to eye muscle weakness, leading to a significant decrease in the accuracy of pupil center and corneal reflex point extraction. Traditional eye-tracking methods cannot adapt to the slow and uncoordinated dynamic changes in eye movements, requiring repeated offline calibration and the collection of large amounts of redundant eye-tracking data for post-processing to obtain stable fixation point coordinates, consuming substantial computational and storage resources. To address the following requirements for this application scenario: adaptability to slow and uncoordinated dynamic changes in eye movements and adaptive processing that does not require frequent offline calibration, we have decided to adopt the following solution: In some optional implementations of certain embodiments, the aforementioned execution entity can generate eye movement intention recognition results based on the acquired eye movement image frame sequence through the following steps: First, for each eye-tracking image frame in the above eye-tracking image frame sequence, perform the following steps: The first sub-step involves converting the aforementioned eye-tracking image frames to obtain binary images corresponding to those frames. In practice, the executing entity can use adaptive threshold binarization technology to convert the eye-tracking image frames into binary images.
[0040] The second sub-step involves feature extraction processing of the aforementioned binary image to obtain the pupil center point and corneal reflection point corresponding to the binary image. The pupil center point represents the coordinates of the geometric center of the user's pupil region. The corneal reflection point represents the center position of the light spot formed by the reflection of a near-infrared light source on the corneal surface. In practice, firstly, the executing entity can use an ellipse fitting algorithm to perform ellipse boundary detection and center localization processing on the pupil region in the aforementioned binary image to obtain the pupil center point. Then, using a speckle detection algorithm, connected component analysis and centroid calculation are performed on the highlighted regions in the aforementioned eye-tracking image frame to obtain the corneal reflection point.
[0041] The third sub-step involves calibrating the aforementioned pupil center point and corneal reflection point based on a preset fixation data set, obtaining calibrated pupil center points and calibrated corneal reflection points. The calibrated pupil center point represents the mapped point obtained after coordinate transformation of the aforementioned pupil center point. The calibrated corneal reflection point represents the mapped point obtained after coordinate transformation of the aforementioned corneal reflection point. The preset fixation data set represents multiple sets of calibrated pupil center points and calibrated corneal reflection points acquired when the user fixates on different calibration points. The calibrated pupil center point represents the original pupil center point acquired when the user fixates on the calibration point. The calibrated corneal reflection point represents the original corneal reflection point acquired when the user fixates on the calibration point. Each set of preset fixation data in the preset fixation data set may include one pupil center point and one corneal reflection point. In practice, firstly, the execution entity calculates the coefficients of the polynomial for each set of preset fixation data in the preset fixation data set using a least-squares polynomial fitting method. Then, using the coefficients of the above polynomials, polynomial calculations are performed on the x-coordinate and y-coordinate of the pupil center point, and the calculated x-coordinate and y-coordinate are used as the calibration pupil center point. Similarly, polynomial calculations are performed on the x-coordinate and y-coordinate of the corneal reflection point, and the calculated x-coordinate and y-coordinate are used as the calibration corneal reflection point.
[0042] The fourth sub-step involves performing vector calculations on the aforementioned calibration pupil center point and calibration corneal reflection point to obtain the pupil-corneal reflection vector. This pupil-corneal reflection vector represents a two-dimensional vector pointing from the calibration corneal reflection point to the calibration pupil center point. In practice, the executing entity first uses vector subtraction to subtract the abscissa of the calibration corneal reflection point from the abscissa of the calibration pupil center point to obtain the first component, and subtracts the ordinate of the calibration corneal reflection point from the ordinate of the calibration pupil center point to obtain the second component. Then, the first component is determined as the abscissa component, and the second component is determined as the ordinate component to obtain the pupil-corneal reflection vector. For example, when the calibration pupil center point is (410, 205) and the calibration corneal reflection point is (415, 200), the pupil-corneal reflection vector is (-5, 5).
[0043] The fifth sub-step involves generating a gaze direction vector based on the aforementioned pupil-corneal reflection vector. This gaze direction vector represents the direction of the user's gaze in three-dimensional space. In practice, firstly, the executing entity can use the arctangent function to convert the horizontal and vertical components of the pupil-corneal reflection vector into horizontal and vertical angles. Then, through trigonometric function operations, these horizontal and vertical angles are converted into direction vectors in three-dimensional space. Finally, this direction vector is determined as the gaze direction vector.
[0044] The second step is to determine the obtained line-of-sight direction vectors as a line-of-sight direction vector sequence.
[0045] The third step involves denoising the aforementioned gaze direction vector sequence to obtain an eye-tracking feature signal sequence. The eye-tracking feature signals in this sequence represent the vectors in three-dimensional space obtained after denoising the gaze direction vectors in the sequence. In practice, firstly, for each gaze direction vector in the sequence, the execution entity uses a median filtering algorithm to perform sliding window processing (e.g., a window size of 3). The gaze direction vectors within the window are sorted according to each component (e.g., x-component, y-component, or z-component), and the median of each component is used as the filtered output value at the current moment, resulting in each denoised gaze direction vector. Then, these denoised gaze direction vectors are used as the eye-tracking feature signal sequence.
[0046] The fourth step involves generating eye movement event detection results based on the aforementioned eye movement feature signal sequence. Each eye movement event detection result can represent the category of eye movement at different times. These results can be either fixation or saccades. In practice, the executing entity can sequentially calculate the angular velocity between each pair of adjacent frames in the aforementioned eye movement feature signal sequence. In response to determining that the angular velocity is below a third preset threshold, the eye movement event detection result of the next frame in the adjacent sequence is determined as fixation; in response to determining that the angular velocity is above a fourth preset threshold, the eye movement event detection result of the next frame in the adjacent sequence is determined as saccades; this process is repeated for all adjacent frames in the aforementioned eye movement feature signal sequence to obtain each eye movement event detection result. Here, the specific values of the third and fourth preset thresholds are not limited; for example, the third preset threshold can be 30 degrees per second, and the fourth preset threshold can be 50 degrees per second. It should be noted that the eye movement category of the first frame in the aforementioned eye movement feature signal sequence is the same as that of the second frame.
[0047] The fifth step involves generating an eye movement event sequence set based on the aforementioned eye movement event detection results. The eye movement event sequences in this set can be either fixation sequences or saccade sequences. A fixation sequence can be a sequence composed of consecutively labeled fixation events from the aforementioned eye movement event detection results. Similarly, a saccade sequence can be a sequence composed of consecutively labeled saccade events from the aforementioned eye movement event detection results. In practice, firstly, the executing entity can determine the consecutively labeled fixation frames from the aforementioned eye movement event detection results as fixation sequences, and the consecutively labeled saccade frames from the aforementioned eye movement event detection results as saccade sequences, thus obtaining each fixation sequence and each saccade sequence. Then, the aforementioned fixation sequences and saccade sequences are arranged sequentially according to their occurrence time to obtain the eye movement event sequence set.
[0048] Step 6: Based on the aforementioned set of eye-tracking event sequences, determine the eye-tracking feature parameters. In practice, firstly, the executing entity can identify the eye-tracking event sequences that meet preset screening conditions from the aforementioned set of eye-tracking event sequences as target eye-tracking event sequences. Then, in response to determining that the target eye-tracking event sequence is a fixation sequence, the average value of each eye-tracking feature signal corresponding to the target eye-tracking event sequence is determined as the fixation point coordinates, and the duration of the target eye-tracking event sequence is determined as the fixation time; in response to determining that the target eye-tracking event sequence is a saccade sequence, the coordinates of the eye-tracking feature signals corresponding to each eye-tracking event detection result in the target eye-tracking event sequence are arranged in chronological order to determine the saccade trajectory. Afterwards, the fixation point coordinates and the fixation time are determined as the first parameter, and the saccade trajectory is used as the second parameter. Finally, the first parameter or the second parameter is determined as the eye-tracking feature parameter. The preset screening condition can be selecting the sequence containing the eye-tracking event detection result with the smallest value at the time of occurrence.
[0049] Step 7: Generate eye movement intention recognition results based on the aforementioned eye movement feature parameters. In practice, in response to determining that the aforementioned eye movement feature parameters are the first parameters, the executing entity can obtain the gaze point coordinates and gaze duration from the aforementioned first parameters, and perform the following steps: in response to determining that the aforementioned gaze point coordinates are located within a preset interaction area and that the aforementioned gaze duration exceeds a fifth preset threshold, the first target eye movement instruction is determined as the eye movement intention recognition result; in response to determining that the aforementioned eye movement feature parameters are the second parameters, the executing entity can obtain the saccade trajectory from the aforementioned second parameters, and perform the following steps: in response to determining that the direction of the aforementioned saccade trajectory is a preset direction, the second target eye movement instruction is determined as the eye movement intention recognition result. The aforementioned preset interaction area can represent the area where operable elements are located on the screen. The aforementioned operable elements can represent interface components that can respond to user interaction; for example, operable elements can be buttons or icons. For example, the fifth preset threshold can be 3 seconds. Both the aforementioned first target eye movement instruction and the aforementioned second target eye movement instruction can represent control instructions generated by user interaction through gaze; for example, the first target eye movement instruction can be selection or confirmation, and the second target eye movement instruction can be page turning or switching. The preset direction can be left or right.
[0050] The above-described technical solution, as an inventive point of this disclosure, addresses technical problem three: "consuming a large amount of computational and storage resources." The reasons for this waste of computational and storage resources are as follows: Myasthenia gravis patients often experience ptosis, slow and uncoordinated eye movements due to weakness of the eye muscles, leading to a significant decrease in the extraction accuracy of the pupil center and corneal reflection point. Traditional eye-tracking methods cannot adapt to the dynamic changes in the slow and uncoordinated eye movements of patients, requiring repeated offline calibration and the collection of a large amount of redundant eye-tracking data for post-processing to obtain stable fixation point coordinates, thus consuming a large amount of computational and storage resources. Solving these factors can reduce the consumption of computational and storage resources. To achieve this effect, the multimodal smart glasses interactive control method based on EEG and eye-tracking fusion disclosed in this disclosure can accurately extract the pupil center and corneal reflection point corresponding to the aforementioned eye-tracking image frames by performing feature recognition processing on the eye-tracking image frames. Furthermore, by calibrating the pupil center and corneal reflection points according to a preset fixation data set, the influence of individual differences on extraction accuracy can be eliminated, eliminating the need for repeated offline calibration. Furthermore, by denoising the aforementioned gaze direction vector sequence to obtain an eye movement feature signal sequence, and generating an eye movement event detection result based on this sequence, an eye movement intention recognition result can be generated. This eliminates the need for collecting a large amount of redundant data for post-processing, thereby significantly reducing the consumption of computational and storage resources. Thus, the computational and storage resources required can be reduced.
[0051] Step 103: Generate an audio perception feature set based on the acquired sound signal.
[0052] In some embodiments, the aforementioned executing entity can generate an audio perception feature set based on the acquired sound signal. The sound signal can represent audio data collected by a microphone. The microphone can be installed in the aforementioned smart glasses. The sound signal can include user voice signals and ambient background sound signals. The user voice signal can represent the sound waveform data emitted by the user when speaking. The ambient background sound signal can represent sound waveform data in the user's environment other than the user voice signal. The audio perception feature set can represent a set of multi-dimensional data features extracted from the sound signal. The audio perception feature set can include a voice intent recognition result, a first user behavior state signal, and a scene category auxiliary determination signal. The voice intent recognition result can represent a control command parsed from the user voice signal. For example, the voice intent recognition result can be a "open application" command. The first user behavior state signal can represent the user's emotion or physiological state inferred from the acoustic features of the user voice signal. Here, the specific content of the emotion and physiological state is not limited; for example, the emotion can be calm, and the physiological state can be fatigue. The scene category auxiliary determination signal can represent the scene type obtained after classifying the ambient background sound signal. For example, the scene category auxiliary determination signal could be a conference room.
[0053] In some optional implementations of certain embodiments, the aforementioned execution entity may generate an audio perception feature set based on the acquired sound signal through the following steps: The first step is to perform semantic analysis on the user's voice signal to obtain the semantic analysis result. This result represents the user's control commands expressed through voice. For example, the semantic analysis result could be "open folder". In practice, the executing entity can use speech recognition technology to convert the user's voice signal into text. Then, the resulting text is used as the semantic analysis result.
[0054] The second step is to generate a first user behavior state signal based on the aforementioned user voice signal. In practice, firstly, the executing entity can extract the acoustic features (e.g., speech rate, pitch, and energy) of the user voice signal using a short-time Fourier transform. Then, the acoustic features are analyzed and processed using a threshold comparison method to determine the user's emotional or physiological state, thus obtaining the first user behavior state signal. For example, when the speech rate is below the first threshold and the energy is below the second threshold, the first user behavior state signal is determined to be "fatigue"; when the speech rate is above the third threshold and the pitch is above the fourth threshold, the first user behavior state signal is determined to be "tension". Here, the specific values of the first, second, third, and fourth thresholds are not limited and can be set according to actual needs.
[0055] The third step involves scene recognition processing of the aforementioned ambient background sound signal to obtain a scene category auxiliary determination signal. In practice, firstly, the executing entity can extract the time-frequency diagram of the ambient background sound signal using a short-time Fourier transform. Then, the time-frequency diagram is compared with a preset scene spectrum template using template matching to obtain various matching degrees. Finally, the scene category corresponding to the highest matching degree among the various matching degrees is determined as the scene category auxiliary determination signal. The preset scene spectrum template can represent the time-frequency diagrams corresponding to various pre-stored scene categories. For example, the preset scene spectrum template may include the time-frequency diagram corresponding to a conference room scene, the time-frequency diagram corresponding to a coffee shop scene, etc.
[0056] The fourth step is to determine the above semantic analysis results, the above first user behavior state signal, and the above scene category auxiliary judgment signal as the audio perception feature set.
[0057] Step 104: Generate a visual perception feature set based on the acquired video image signal.
[0058] In some embodiments, the aforementioned execution entity can generate a visual perception feature set based on the acquired video image signal. The video image signal can represent continuous image frames captured by a camera. The visual perception feature set can represent a collection of multi-dimensional visual features extracted from the video image signal. The visual perception feature set may include a gesture intent recognition result, a current scene category recognition result, and a second user behavior state signal. The gesture intent recognition result can represent the control command corresponding to the user's gesture identified from the video image signal. For example, when the user's gesture is a swipe gesture, the corresponding control command is page turning; when the user's gesture is a fist gesture, the corresponding control command is confirmation. The current scene category recognition result can represent the scene type of the user's environment identified from the video image signal. The second user behavior state signal can represent the user's physical behavior state identified from the video image signal. Here, the specific type of the behavior state is not limited; for example, the behavior state can be walking or sitting still.
[0059] In some optional implementations of certain embodiments, the aforementioned execution entity may generate a visual perception feature set based on the acquired video image signal through the following steps: First, based on the above video image signal, perform the following steps: First, the video image signal is processed for body motion recognition to obtain the body motion recognition result. This result characterizes the type of user gesture detected. For example, the recognition result could be a clenched fist gesture. In practice, the executing entity first extracts the user's hand region from the video image signal using a target detection algorithm (e.g., YOLO). Then, a skeletal keypoint detection algorithm (e.g., MediaPipe algorithm) identifies the positions of keypoints in the hand region. Finally, a support vector machine classifier classifies the positions of these keypoints, outputting the gesture type as the body motion recognition result.
[0060] Then, scene category recognition processing is performed on the aforementioned video image signals to obtain initial scene category recognition results. These initial scene category recognition results characterize the scene type initially identified from the aforementioned video image signals. In practice, the executing entity can use a convolutional neural network to perform feature extraction and classification processing on the aforementioned video image signals, using the output scene category as the initial scene category recognition result. For example, the convolutional neural network can be a ResNet50 network.
[0061] Secondly, based on the initial scene category identification result and the scene category auxiliary determination signal, the current scene category identification result is generated. In practice, the executing entity can compare the initial scene category identification result with the scene category auxiliary determination signal. If the initial scene category identification result and the scene category auxiliary determination signal are the same, the initial scene category identification result is determined as the current scene category identification result. If the initial scene category identification result and the scene category auxiliary determination signal are different, a weighted voting method is used to fuse the initial scene category identification result and the scene category auxiliary determination signal to obtain the current scene category identification result. It should be noted that the weights of the initial scene category identification result and the scene category auxiliary determination signal in the weighted voting are not limited and can be set according to actual needs.
[0062] Finally, behavioral state analysis is performed on the aforementioned video image signals to obtain the second user behavioral state signal. In practice, the executing entity can extract the user's hand and shoulder regions from the video image signals using a target detection algorithm (e.g., YOLO). Then, a skeletal keypoint detection algorithm is used to identify the positions of keypoints in the hand and shoulder regions. The user's behavioral state is determined by combining the changes in the keypoint positions over time; for example, if the shoulder keypoints continuously shift over time, the user's behavioral state is determined to be walking. Finally, the obtained user behavioral state is defined as the second user behavioral state signal.
[0063] The second step is to determine the above-mentioned body movement recognition results, the above-mentioned current scene category recognition results, and the above-mentioned second user behavior state signal as the visual perception feature set.
[0064] Step 105: Determine the target priority fusion rule based on the current scene category recognition result.
[0065] In some embodiments, the aforementioned execution entity can determine a target priority fusion rule based on the current scene category identification result. This target priority fusion rule can characterize the priority ranking rules among various control signals in different scenes. It should be noted that the different scenes mentioned above can include Class I scenes and Class II scenes. The various control signals mentioned above can characterize parameters of the user's intent or interaction identified from signals acquired from different sensors. These various control signals can include EEG intent recognition results, voice intent recognition results, gesture intent recognition results, and eye-tracking intent recognition results.
[0066] In some optional implementations of certain embodiments, the aforementioned execution entity may determine the target priority fusion rule based on the current scene category identification result through the following steps: The first step, in response to determining that the current scene category recognition result is a scene of type one, is to determine the first interaction priority order as the target priority fusion rule. Here, the aforementioned scene of type one can represent a scenario that requires others to be unaware or have weak perception. For example, a scene of type one could be a meeting room or a public space. The aforementioned first interaction priority order could be: EEG intention recognition result priority over eye movement intention recognition result, eye movement intention recognition result priority over voice intention recognition result, and voice intention recognition result priority over gesture intention recognition result.
[0067] The second step, in response to determining that the current scene category recognition result is a Class II scene, establishes the second interaction priority order as the target priority fusion rule. Here, the aforementioned Class II scenes can represent scenarios where the user's behavior does not disturb others. For example, a Class II scene could be a scenario where the user is alone. The aforementioned second interaction priority order could be: voice intent recognition result takes precedence over gesture intent recognition result, gesture intent recognition result takes precedence over EEG intent recognition result, and EEG intent recognition result takes precedence over eye-tracking intent recognition result.
[0068] Step 106: Generate the current execution control command based on the target priority fusion rule, EEG intention recognition result, eye movement intention recognition result, voice intention recognition result, and gesture intention recognition result.
[0069] In some embodiments, the executing entity can generate a current execution control instruction based on the target priority fusion rule, the EEG intention recognition result, the eye-tracking intention recognition result, the voice intention recognition result, and the gesture intention recognition result. This current execution control instruction can characterize the interactive control instructions that the smart glasses need to execute.
[0070] In some optional implementations of certain embodiments, the execution entity may generate the current execution control command by following the steps described above: based on the target priority fusion rule, the EEG intention recognition result, the eye movement intention recognition result, the voice intention recognition result, and the gesture intention recognition result. The first step is to determine the highest priority recognition result based on the aforementioned target priority fusion rules, the aforementioned EEG intention recognition results, the aforementioned eye-tracking intention recognition results, the aforementioned voice intention recognition results, and the aforementioned gesture intention recognition results. The highest priority recognition result can represent the control signal with the highest priority in the current scenario. In practice, the executing entity can prioritize the aforementioned EEG intention recognition results, eye-tracking intention recognition results, voice intention recognition results, and gesture intention recognition results according to the aforementioned target priority fusion rules, and take the control signal with the highest priority as the highest priority recognition result. For example, when the EEG intention recognition result has the highest priority, it is determined as the highest priority recognition result.
[0071] The second step is to determine the highest priority identification result as the current execution control instruction.
[0072] Step 107: Execute the interactive operation corresponding to the current execution control instruction based on the current execution control instruction.
[0073] In some embodiments, the executing entity can perform an interactive operation corresponding to the current execution control instruction. This interactive operation can characterize the action performed by the smart glasses after receiving the current execution control instruction. For example, the interactive operation could be opening a folder or turning a page.
[0074] The various embodiments disclosed herein have the following beneficial effects: Through a multimodal smart glasses interaction control method based on EEG and eye-tracking fusion according to some embodiments of this disclosure, the consumption of computing resources can be reduced and the time for recognizing control commands can be shortened, thereby improving the real-time performance of interactive control and thus enhancing the overall interactive experience. The reason for the high consumption of computing resources, long time for recognizing control commands, poor real-time performance of interactive control, and poor overall interactive experience is that, apart from eye-tracking technology, other technical solutions (e.g., voice, gestures, etc.) are easily perceived by others. In situations requiring quiet or privacy (e.g., a conference room), users cannot use voice or gestures to achieve natural and smooth operation, resulting in a poor user experience. Furthermore, existing combined interactive control solutions require processing a large amount of multimodal data simultaneously, leading to high consumption of computing resources, long time for recognizing control commands, and poor real-time performance of interactive control. At the same time, existing solutions cannot dynamically adjust the priority of multimodal data according to different usage scenarios, resulting in a poor overall interactive experience. Based on this, some embodiments of the multimodal smart glasses interaction control method based on EEG and eye-tracking fusion disclosed herein first generate EEG intention recognition results based on the acquired raw EEG signals. This yields the EEG intention recognition result. Then, based on the acquired eye-tracking image frame sequence, generate eye-tracking intention recognition results. This yields the eye-tracking intention recognition result. Next, based on the acquired sound signals, generate an audio perception feature set. This yields various audio perception features. Secondly, based on the acquired video image signals, generate a visual perception feature set. This yields various visual perception features. Then, based on the aforementioned current scene category recognition results, determine a target priority fusion rule. This determines the priority of each modality of data in different scenes. Next, based on the aforementioned target priority fusion rule, the aforementioned EEG intention recognition results, the aforementioned eye-tracking intention recognition results, the aforementioned voice intention recognition results, and the aforementioned gesture intention recognition results, generate a current execution control command. This yields the execution control command at the current moment. Finally, based on the aforementioned current execution control command, execute the interactive operation corresponding to the aforementioned current execution control command. Because it doesn't indiscriminately fuse all the data from the original sensors, but instead extracts the EEG intention recognition results, eye-tracking intention recognition results, voice intention recognition results, and gesture intention recognition results separately, the amount of data that needs to be processed in the subsequent fusion stage is reduced. It also doesn't use a single fixed priority rule, but distinguishes different scenarios based on the current scene category recognition results, using different signal priority orders in different scenarios, thus balancing concealment and interaction efficiency in different scenarios. Furthermore, because it determines the highest priority recognition result from multiple intention recognition results based on target priority fusion rules and generates the current execution control command, it only needs to process the recognition result of the highest priority single modality to complete the generation of the interaction control command.Therefore, it can avoid the waste of computing resources caused by processing a large amount of multimodal data at the same time, thereby reducing the consumption of computing resources and shortening the time for control command recognition, thus improving the real-time performance of interactive control, while improving the overall interactive experience through scene-adaptive priority differentiation strategy.
[0075] Further reference Figure 2 As an implementation of the methods shown in the above figures, this disclosure provides some embodiments of a multimodal smart glasses interactive control method based on EEG and eye-tracking fusion. These device embodiments are similar to... Figure 1 Corresponding to the method embodiments shown, the device can be specifically applied to various electronic devices.
[0076] like Figure 2 As shown, some embodiments of the multimodal smart glasses interactive control device 200 based on EEG and eye-tracking fusion include: a first generation unit 201, a second generation unit 202, a third generation unit 203, a fourth generation unit 204, a determination unit 205, a fifth generation unit 206, and an execution unit 207. The system comprises the following components: a first generation unit 201, configured to generate a brainwave intention recognition result based on the acquired raw brainwave signals; a second generation unit 202, configured to generate an eye-tracking intention recognition result based on the acquired eye-tracking image frame sequence; a third generation unit 203, configured to generate an audio perception feature set based on the acquired sound signals, wherein the audio perception feature set includes a speech intention recognition result and a first user behavior state signal; a fourth generation unit 204, configured to generate a visual perception feature set based on the acquired video image signals, wherein the visual perception feature set includes a gesture intention recognition result, a current scene category recognition result, and a second user behavior state signal; a determination unit 205, configured to determine a target priority fusion rule based on the current scene category recognition result; a fifth generation unit 206, configured to generate a current execution control command based on the target priority fusion rule, the brainwave intention recognition result, the eye-tracking intention recognition result, the speech intention recognition result, and the gesture intention recognition result; and an execution unit 207, configured to execute an interactive operation corresponding to the current execution control command.
[0077] It is understandable that the units and references described in the multimodal smart glasses interactive control device 200 based on EEG and eye-tracking fusion are... Figure 1 The steps in the described method correspond accordingly. Therefore, the operations, features, and beneficial effects described above for the method also apply to the multimodal smart glasses interactive control device 200 based on EEG and eye-tracking fusion and the units contained therein, and will not be repeated here.
[0078] The following is for reference. Figure 3It shows a schematic diagram of the structure of an electronic device 300 (e.g., a computing device) suitable for implementing some embodiments of the present disclosure. Figure 3 The electronic device shown is merely an example and should not be construed as limiting the functionality and scope of the embodiments of this disclosure.
[0079] like Figure 3 As shown, the electronic device 300 may include a processing unit 301 (e.g., a central processing unit, a graphics processor, etc.), which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 302 or a program loaded from a storage device 308 into a random access memory (RAM) 303. The RAM 303 also stores various programs and data required for the operation of the electronic device 300. The processing unit 301, ROM 302, and RAM 303 are interconnected via a bus 304. An input / output (I / O) interface 305 is also connected to the bus 304.
[0080] Typically, the following devices can be connected to I / O interface 305: input devices 306 including, for example, touchscreens, touchpads, keyboards, mice, cameras, microphones, accelerometers, gyroscopes, etc.; output devices 307 including, for example, liquid crystal displays (LCDs), speakers, vibrators, etc.; storage devices 308 including, for example, magnetic tapes, hard disks, etc.; and communication devices 309. Communication device 309 allows electronic device 300 to communicate wirelessly or wiredly with other devices to exchange data. Although Figure 3 An electronic device 300 with various devices is shown; however, it should be understood that it is not required to implement or possess all of the devices shown. More or fewer devices may be implemented or possessed alternatively. Figure 3 Each box shown can represent a device or multiple devices as needed.
[0081] In particular, according to some embodiments of this disclosure, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, some embodiments of this disclosure include a computer program product comprising a computer program carried on a computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via communication device 309, or installed from storage device 308, or installed from ROM 302. When the computer program is executed by processing device 301, it performs the functions defined in the methods of some embodiments of this disclosure.
[0082] It should be noted that, in some embodiments of this disclosure, the computer-readable medium may be a computer-readable signal medium or a computer-readable storage medium, or any combination thereof. A computer-readable storage medium may be, for example,—but not limited to—an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of a computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In some embodiments of this disclosure, a computer-readable storage medium may be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. In some embodiments of this disclosure, a computer-readable signal medium may include a data signal propagated in baseband or as part of a carrier wave, carrying computer-readable program code. Such propagated data signals may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A computer-readable signal medium can be any computer-readable medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted using any suitable medium, including but not limited to: wires, optical fibers, RF (radio frequency), etc., or any suitable combination thereof.
[0083] In some implementations, clients and servers can communicate using any currently known or future-developed network protocol such as HTTP (Hypertext Transfer Protocol) and can interconnect with digital data communication (e.g., communication networks) of any form or medium. Examples of communication networks include local area networks (“LANs”), wide area networks (“WANs”), the Internet (e.g., the Internet of Things), and peer-to-peer networks (e.g., ad hoc peer-to-peer networks), as well as any currently known or future-developed networks.
[0084] The aforementioned computer-readable medium may be included in the aforementioned electronic device; or it may exist independently and not assembled into the electronic device. The aforementioned computer-readable medium carries one or more programs, which, when executed by the electronic device, cause the electronic device to: generate a brainwave intention recognition result based on the acquired raw brainwave signals; generate an eye-tracking intention recognition result based on the acquired eye-tracking image frame sequence; generate an audio perception feature set based on the acquired sound signals, wherein the audio perception feature set includes a speech intention recognition result and a first user behavior state signal; generate a visual perception feature set based on the acquired video image signals, wherein the visual perception feature set includes a gesture intention recognition result, a current scene category recognition result, and a second user behavior state signal; determine a target priority fusion rule based on the current scene category recognition result; generate a current execution control instruction based on the target priority fusion rule, the brainwave intention recognition result, the eye-tracking intention recognition result, the speech intention recognition result, and the gesture intention recognition result; and execute an interactive operation corresponding to the current execution control instruction based on the current execution control instruction.
[0085] Computer program code for performing operations of some embodiments of this disclosure can be written in one or more programming languages or a combination thereof, including object-oriented programming languages such as Java, Smalltalk, and C++, and conventional procedural programming languages such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).
[0086] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.
[0087] The units described in some embodiments of this disclosure can be implemented in software or hardware. The described units can also be housed in a processor; for example, a processor may be described as including a first generation unit, a second generation unit, a third generation unit, a fourth generation unit, a determining unit, a fifth generation unit, and an execution unit. The names of these units do not necessarily limit the specific unit; for example, the first generation unit may also be described as "a unit that generates brainwave intention recognition results based on acquired raw brainwave signals."
[0088] The functions described above in this document can be performed at least in part by one or more hardware logic components. For example, exemplary types of hardware logic components that can be used, without limitation, include: field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), system-on-a-chip (SoCs), complex programmable logic devices (CPLDs), and so on.
[0089] The above description is merely a selection of preferred embodiments of this disclosure and an explanation of the technical principles employed. Those skilled in the art should understand that the scope of the invention involved in the embodiments of this disclosure is not limited to technical solutions formed by specific combinations of the above-described technical features, but should also cover other technical solutions formed by arbitrary combinations of the above-described technical features or their equivalents without departing from the above-described inventive concept. For example, technical solutions formed by substituting the above-described features with (but not limited to) technical features with similar functions disclosed in the embodiments of this disclosure.
Claims
1. A multimodal smart glasses interactive control method based on EEG and eye-tracking fusion, comprising: Based on the acquired raw EEG signals, generate EEG intention recognition results; Based on the acquired eye movement image frame sequence, generate eye movement intention recognition results; Based on the acquired sound signal, an audio perception feature set is generated, wherein the audio perception feature set includes the voice intent recognition result and the first user behavior state signal; Based on the acquired video image signals, a visual perception feature set is generated, wherein the visual perception feature set includes gesture intention recognition results, current scene category recognition results, and second user behavior state signals; Based on the current scene category recognition results, determine the target priority fusion rule; Based on the target priority fusion rule, the EEG intention recognition result, the eye movement intention recognition result, the voice intention recognition result, and the gesture intention recognition result, a current execution control command is generated; According to the current execution control instruction, perform the interactive operation corresponding to the current execution control instruction.
2. The method according to claim 1, wherein, The step of generating brainwave intention recognition results based on the acquired raw brainwave signals includes: The original EEG signal is filtered to obtain a denoised EEG signal; Based on the denoised EEG signal, generate an artifact-free EEG signal; The artifact-free EEG signal is resampled to obtain a resampled EEG signal; Feature extraction is performed on the resampled EEG signal to obtain EEG feature signals; Based on the EEG characteristic signals, an EEG intention recognition result is generated.
3. The method according to claim 1, wherein, The sound signal includes user voice signal and ambient background sound signal; as well as The step of generating an audio perception feature set based on the acquired sound signal includes: The user's voice signal is subjected to semantic analysis processing to obtain the semantic analysis results; A first user behavior state signal is generated based on the user's voice signal; The ambient background sound signal is processed for scene recognition to obtain a scene category auxiliary determination signal; The semantic analysis results, the first user behavior state signal, and the scene category auxiliary determination signal are determined as the audio perception feature set.
4. The method according to claim 1, wherein, The step of determining the target priority fusion rule based on the current scene category recognition result includes: In response to determining that the current scene category identification result is a scene of a certain type, the first interaction priority order is determined as the target priority fusion rule; In response to determining that the current scene category identification result is a second-class scene, the second interaction priority order is determined as the target priority fusion rule.
5. The method according to claim 1, wherein, The step of generating a current execution control command based on the target priority fusion rule, the EEG intention recognition result, the eye movement intention recognition result, the voice intention recognition result, and the gesture intention recognition result includes: Based on the target priority fusion rule, the EEG intention recognition result, the eye movement intention recognition result, the voice intention recognition result, and the gesture intention recognition result, the highest priority recognition result is determined; The highest priority identification result is determined as the current execution control instruction.
6. The method according to claim 1, wherein, The step of generating a visual perception feature set based on the acquired video image signal includes: Based on the video image signal, perform the following steps: The video image signal is processed for limb movement recognition to obtain limb movement recognition results; The video image signal is subjected to scene category recognition processing to obtain an initial scene category recognition result; Based on the initial scene category recognition result and the scene category auxiliary determination signal, the current scene category recognition result is generated; The video image signal is subjected to behavioral state analysis processing to obtain a second user behavioral state signal; The limb movement recognition result, the current scene category recognition result, and the second user behavior state signal are determined as a visual perception feature set.
7. A multimodal smart glasses interactive control device based on EEG and eye-tracking fusion, comprising: The first generation unit is configured to generate brainwave intention recognition results based on the acquired raw brainwave signals; The second generation unit is configured to generate eye movement intention recognition results based on the acquired eye movement image frame sequence; The third generation unit is configured to generate an audio perception feature set based on the acquired sound signal, wherein the audio perception feature set includes a voice intent recognition result and a first user behavior state signal. The fourth generation unit is configured to generate a visual perception feature set based on the acquired video image signal, wherein the visual perception feature set includes gesture intention recognition result, current scene category recognition result, and second user behavior state signal; The determining unit is configured to determine the target priority fusion rule based on the current scene category recognition result; The fifth generation unit is configured to generate a current execution control command based on the target priority fusion rule, the EEG intention recognition result, the eye movement intention recognition result, the voice intention recognition result, and the gesture intention recognition result; The execution unit is configured to perform an interactive operation corresponding to the current execution control instruction, based on the current execution control instruction.
8. An electronic device, comprising: One or more processors; A storage device on which one or more programs are stored; When the one or more programs are executed by the one or more processors, the one or more processors implement the method as described in any one of claims 1 to 6.
9. A computer-readable medium having a computer program stored thereon, wherein, When the computer program is executed by a processor, it implements the method as described in any one of claims 1 to 6.