Conversation communication method and system based on non-invasive brain-computer interface and multi-modal AI large model
By using a non-invasive brain-computer interface and a multimodal AI large-scale model for dialogue, the problem of insufficient communication links in communication for patients with severe hemiplegia has been solved, achieving efficient, reliable, and natural human-computer dialogue, and improving user autonomy and communication efficiency.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- SHAANXI JIECHUANGRUI INTELLIGENT TECHNOLOGY CO LTD
- Filing Date
- 2025-12-24
- Publication Date
- 2026-05-12
AI Technical Summary
Existing technologies for assisting communication in patients with severe hemiplegia suffer from insufficient communication link bandwidth, poor signal transmission reliability, and low intelligence of the system protocol stack, resulting in poor user autonomy, low interaction efficiency, high learning costs, and a stiff communication experience.
It adopts a dialogue communication method based on non-invasive brain-computer interface and multimodal AI large model. It collects EEG signals through non-invasive brain-computer interface device, and performs semantic understanding and generation in combination with multimodal AI large model. It provides two modes: voice interaction and EEG text input, and achieves seamless switching through visual stimulation interface. It establishes an efficient direct communication link from neural signals to semantics. It integrates signal preprocessing, EEG recognition control, interaction module and speech generation module, and supports edge-cloud collaborative computing and robust algorithm design.
It achieves a paradigm shift from low-dimensional action encoding to high-dimensional neural-semantic direct connection, improving communication efficiency and reliability, ensuring user autonomy and naturalness of communication, reducing learning costs and latency, and providing an efficient, reliable, and flexible human-computer dialogue collaboration system.
Smart Images

Figure CN122018675A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of brain-computer interface and intelligent care technology, and in particular to a dialogue communication method and system based on a non-invasive brain-computer interface and a multimodal AI large model. Background Technology
[0002] For assistive communication in people with severe motor and language impairments (such as patients with severe hemiplegia), existing technologies have significant limitations in terms of communication efficiency, autonomy, and user experience.
[0003] The first type of traditional approach relies on caregivers observing the patient's limited limb movements or simple vocalizations. This communication method essentially reduces the user's rich inner intentions (high-dimensional information) to extremely limited physical signals for transmission. This results in extremely low information entropy, a high error rate, low communication efficiency, and a high risk of misjudging needs.
[0004] The second type of assistive devices, based on specific biosignals, such as eye trackers, while adding a control channel (gaze coordinates), has inherent limitations. From a communication system perspective, eye-tracking signals are susceptible to interference from "channel noise" such as eye fatigue, changes in ambient light, and unconscious saccades, resulting in insufficient stability and low communication link reliability. Furthermore, their functionality is typically limited to simple interface navigation and control, lacking the ability to understand and generate deep semantic intent from users, thus failing to support complete natural dialogue, and exhibiting a low level of communication protocol.
[0005] The third category of existing non-invasive brain-computer interface technologies, particularly those based on steady-state visual evoked potentials (SSVEP) or P300 event-related potentials, largely focuses on medical rehabilitation training or basic spelling applications. Their problems include: 1) Low system integration: These are mostly isolated "stimulus-decoding" modules, not deeply integrated with upstream natural language processing (NLP) engines, forming a "half-baked" communication system. Users need to transmit information through complex encoding (such as character-by-character spelling), resulting in poor communication speed (measured by information transfer rate, ITR) and user experience. 2) Insufficient interference resistance and robustness: In real-world environments, noise such as power frequency interference and electromyographic artifacts severely degrades the signal-to-noise ratio (SNR) of EEG signals. Existing solutions often lack adaptive front-end signal enhancement and robust feature extraction algorithms, leading to a sharp drop in decoding accuracy under non-ideal conditions. 3) Rigid devices and protocols: Devices are often bulky and complex to operate, with fixed communication protocols, unable to dynamically adjust interaction strategies based on user status or environment, lacking personalized adaptability.
[0006] In summary, the common defects of existing technical solutions can be summarized as follows: insufficient communication link bandwidth, poor signal transmission reliability, and low intelligence of system protocol stack, resulting in poor user autonomy, low interaction efficiency, high learning cost, and awkward communication experience. Summary of the Invention
[0007] The purpose of this invention is to provide a dialogue communication method and system based on a non-invasive brain-computer interface and a multimodal AI large model, which can solve existing problems.
[0008] The technical problem solved by this invention is as follows: 1. The interaction link in the existing technology has a "bottleneck" or "dimensionality reduction" problem: Traditional methods rely on residual limb functions or simple vocalizations. Essentially, they reduce the brain's complex intentions (high-dimensional information) to limited physical actions or simple syllables (low-dimensional information) for transmission, resulting in low information entropy, easy distortion, and low efficiency.
[0009] Existing assistive devices, such as eye trackers, although they add information dimensions, do not directly carry semantics in the signal itself (gaze coordinates). They need to be converted into interface instructions, which is a long process and is susceptible to fatigue, ambient light interference, and signal instability.
[0010] 2. Existing technologies are mostly isolated functional modules, lacking organic integration: Limited functionality: Brain-computer interfaces are mostly used for rehabilitation training, and eye trackers are used for basic control, lacking deep integration with natural language processing (NLP).
[0011] Protocol and interface fragmentation: The data transmission protocols between devices are simple and fail to build a coherent data flow from neural signals to natural semantics.
[0012] 3. The user experience bottleneck caused by existing technologies lies in high latency, high learning cost, and limited expression.
[0013] The technical solution of the present invention is as follows: According to one aspect of the present invention, a dialogue communication method based on a non-invasive brain-computer interface and a multimodal AI large model is provided, comprising at least the following steps: S1. The user wears and activates a non-invasive brain-computer interface device, which continuously collects the user's brain signals and establishes and maintains a wireless communication connection with the terminal device. S2. The terminal device launches the application and displays a function selection interface containing at least one function option to the user. S3. By decoding EEG signals, identify the user's intention to select the "Dialogue Communication Scene" function option in the function selection interface; S4. Enter the dialogue and communication scene sub-module. The terminal device presents the user with an interaction mode selection interface that includes at least two modes: "voice interaction" and "electroencephalogram text input". S5. Based on the user's selection in the decoded EEG signal recognition, proceed to the processing flow of the corresponding interaction mode.
[0014] As a further technical solution of the present invention, when the user selects the "voice interaction" mode, the following steps are performed: S6. Collect external voice information through a sound pickup device; S7. Send the external voice information or its converted text to the AI large model processing unit. The AI large model processing unit performs semantic understanding and generates at least one candidate response text. S8. The terminal device displays a visual stimulus interface containing at least one candidate answer text. S9. The user gazes at the visual stimulus area corresponding to the candidate response text of the target, and the brain-computer interface device collects the corresponding electroencephalogram (EEG) signals. S10. The terminal device identifies the user's intention to select the target candidate response text based on the EEG signal. S11. Send the selected candidate answer text to the speech generation module; S12, The speech generation module converts text into speech signals and plays them through an audio output device.
[0015] As a further technical solution of the present invention, when the user selects the "brainwave text input" mode, the following steps are performed: S13. The terminal device displays a character input interface, which contains multiple character selection areas encoded with specific visual stimuli. S14. The user inputs text by sequentially looking at the area corresponding to the target character. The system continuously decodes the EEG signals to identify the text sequence input by the user. S15. Send the text sequence input by the user to the AI large model processing unit for semantic optimization processing to generate optimized text; S16. Send the optimized text to the speech generation module; S17. The speech generation module converts text into speech signals and plays them through an audio output device.
[0016] As a further technical solution of the present invention, the visual stimulus encoding in each interface is implemented based on the steady-state visual evoked potential paradigm. Each selectable option or character area is assigned a unique visual flashing frequency. Furthermore, when the system decodes and confirms the user's intention to gaze at a certain option or area, it immediately switches the visual stimulus of that area from the periodic flashing mode to a static visual confirmation feedback indicator.
[0017] As a further technical solution of the present invention, during a dialogue, it supports seamless switching between "voice interaction" mode and "brainwave text input" mode according to the user's intention; the AI large model processing unit maintains consistent contextual information throughout the dialogue to ensure the continuity of communication.
[0018] As a further technical solution of the present invention, the EEG signal recognition process is implemented based on the event-related potential paradigm; the interaction mode selection interface and the character input interface are constructed as a stimulus matrix, and the system uses the rows or columns of the matrix highlighted by random sequence as rare stimuli, and decodes the user's selection intention by detecting the event-related potential components in the EEG signal that are phase-locked with the rare stimuli.
[0019] As a further technical solution of the present invention, the character input interface is a nine-key virtual keyboard interface, where each number key area corresponds to a group of letters or characters and is encoded by independent visual stimuli; the input process includes the system confirming that the user is looking at the target number key area, and then making a secondary selection in the character group corresponding to the number key to determine the specific input character.
[0020] As a further technical solution of the present invention, the nine-key virtual keyboard interface can be replaced by a twenty-six-key virtual keyboard interface or a handwriting input virtual keyboard interface.
[0021] According to another aspect of the present invention, a dialogue communication system based on a non-invasive brain-computer interface and a multimodal AI large model is also provided, for implementing the above-mentioned dialogue communication method based on a non-invasive brain-computer interface and a multimodal AI large model, the dialogue communication system comprising at least: Portable, non-invasive brain-computer interface device for acquiring and preprocessing users' electroencephalogram (EEG) signals; The terminal device is wirelessly connected to the brain-computer interface device and is used to present a visual stimulus interface, perform real-time intention decoding of brain signals, manage the interaction process, and process audio input and output. In addition, there is an AI large model processing unit that communicates and connects with the terminal device to perform tasks such as speech recognition, semantic understanding, natural language generation, and text optimization.
[0022] As a further technical solution of the present invention, the system integrates a signal preprocessing module, an EEG recognition control module, an interaction module, an AI large model processing module, and a speech generation module; in, The signal preprocessing module, integrated into the brain-computer interface device, is used to filter and reduce noise in the acquired raw EEG signals. The EEG recognition control module and the interaction module are integrated into the terminal device; The EEG recognition control module is used to extract features and decode intent from preprocessed EEG signals to generate control commands; The interaction module is used to respond to control commands, provide and manage a visual interaction interface based on a preset stimulus-response paradigm, and the visual interaction interface is used to at least realize two-way dialogue and communication functions. The AI large model processing module is the core component of the AI large model processing unit. It is deployed in the cloud or locally on the terminal device and communicates with the interaction module and the speech generation module to perform semantic understanding and natural language generation tasks. The speech generation module is used to convert text information into speech signals and output them.
[0023] As a further technical solution of the present invention, the interaction module includes a two-way dialogue communication submodule, which is configured to support at least two interaction channels: The first interaction channel is configured to receive external voice input, call the AI large model processing module to generate at least one candidate response text, provide the user with a selection through the visual interaction interface, and then output it through the voice generation module. And / or, The second interaction channel is configured to provide character input functionality through a visual interaction interface, recognize the character sequence input by the user, call the AI large model processing module to perform semantic optimization on the character sequence, and then output the optimized text through the speech generation module.
[0024] The beneficial effects of this invention are as follows: 1. It fundamentally broadens the communication bandwidth and realizes a paradigm shift from low-dimensional action coding to high-dimensional neural-semantic direct connection.
[0025] Traditional solutions (such as gestures and simple vocalizations) or basic assistive devices (such as eye trackers) essentially transmit a user's rich inner intentions through an extremely limited, low-dimensional, and easily interfered-with physical channel, resulting in low information entropy and high loss. This invention innovatively constructs a direct, high-order communication link from "neural signals to semantic understanding." The core lies in the introduction of a multimodal AI model that acts as an "intelligent semantic modem." It not only decodes brainwave signals into control commands but also deeply understands the user's potential, incompletely expressed intentions and directly generates context-appropriate natural language. This elevates the effectiveness of the communication system from ensuring error-free transmission (correct character spelling) to ensuring accurate semantic delivery and generation, solving the fundamental bottlenecks of dimensionality reduction and distortion in intention transmission in existing technologies.
[0026] 2. An intelligent interaction protocol of "dual-mode flexible input + neural closed-loop confirmation" was established, which improves efficiency while ensuring absolute autonomy and reliability of control.
[0027] Existing technologies suffer from simplistic and rigid interaction modes. This invention innovatively designs a flexible input mechanism that allows for seamless switching between two parallel modes: "voice interaction" and "EEG text input." This allows the system to adaptively select the optimal input strategy based on the dynamic complexity of the communication scenario. More importantly, the system creatively uses EEG signals as a universal control channel for final confirmation, establishing a "neural feedback confirmation loop" for any content generated by AI (candidate answers) or user-defined input. This design achieves two major breakthroughs at the interaction protocol level: first, it firmly hands over final control to the user, ensuring that the user is "in the loop" and avoiding excessive AI autonomy or misjudgment; second, it forms a complete "perception-cognition-decision-execution" intelligent closed loop, fundamentally guaranteeing the authenticity and accuracy of the intent behind each information output.
[0028] 3. Through edge-cloud collaborative computing and robust algorithm design, high availability and low latency experience are achieved in complex scenarios.
[0029] To address the challenges of fragile EEG signals and their susceptibility to environmental noise interference, this invention employs synergistic optimization at the system architecture and algorithm levels. It adopts a heterogeneous collaborative architecture of "real-time local decoding on the terminal + deep AI processing in the cloud": deploying feature extraction and intent matching algorithms requiring millisecond-level response on the terminal ensures real-time control; while placing computationally intensive semantic understanding and generation tasks in the cloud endows the system with powerful intelligence. Simultaneously, by integrating adaptive noise suppression algorithms and supporting multi-paradigm replaceable interaction designs such as SSVEP / P300, the system's anti-interference capability and overall robustness are significantly improved under different user physiological states and diverse environments. This allows the system to move from "laboratory usability" to "real-world usability."
[0030] 4. It provides a fully optimized user experience, significantly improving the naturalness, efficiency, and user dignity of communication. These technical improvements ultimately converge into a qualitative leap in the end-user experience: Increased efficiency: AI completes and optimizes the semantics of vague and short inputs, allowing users to convey complex meanings without having to complete tedious full encoding (such as word-by-word spelling).
[0031] Natural Interaction Flow: Dual-mode switching and context maintenance make the communication process smooth and coherent, closer to the natural rhythm of human conversation.
[0032] Reduced learning and usage costs: Intuitive visual interfaces (such as nine-key keyboards) and intelligent assistance lower the barrier to entry for operation.
[0033] Enhanced sense of control and autonomy: The neural closed-loop confirmation mechanism ensures that users always feel they are in control of the conversation, greatly protecting their autonomy and dignity.
[0034] In summary, this invention is not simply a combination of brain-computer interfaces and AI technology, but rather a systemic upgrade achieved through reconstructing communication links, innovating interaction protocols, and optimizing system architecture. It successfully transforms an originally inefficient, unreliable, and jarring assistive tool into a highly efficient, reliable, flexible, and respectful human-computer dialogue and collaboration system, demonstrating significant progress and outstanding substantive features. Attached Figure Description
[0035] Figure 1 This is a simplified hardware framework diagram of the dialogue communication system based on a non-invasive brain-computer interface and a multimodal AI large model according to the present invention. Figure 2 This is a software system architecture diagram of the dialogue communication system based on a non-invasive brain-computer interface and a multimodal AI large model according to the present invention. Figure 3 This is a simplified operational logic diagram of the dialogue communication method based on non-invasive brain-computer interface and multimodal AI large model of the present invention; Figure 4 One of the interface diagrams displayed on the screen of a terminal device running the dialogue communication method based on a non-invasive brain-computer interface and a multimodal AI large model of the present invention; Figure 5 Another interface diagram is shown on the display screen of a terminal device running the dialogue communication method based on non-invasive brain-computer interface and multimodal AI large model of the present invention.
[0036] Figure 1 The attached figures are labeled as follows: 1-Brain-computer interface device; 2-Terminal device; 3-Cloud server; 11-Head-mounted electrode module; 12-Signal acquisition module; 13-Wireless transmission module; 14-Power supply module; 15-Main control module. Detailed Implementation
[0037] The following specific examples illustrate the implementation of the present invention. Those skilled in the art can easily understand other advantages and effects of the present invention from the content disclosed in this specification. The present invention can also be implemented or applied through other different specific embodiments, and various details in this specification can also be modified or changed based on different viewpoints and applications without departing from the spirit of the present invention. It should be noted that, unless otherwise specified, the following embodiments and features described therein can be combined with each other.
[0038] It should be noted that the illustrations provided in the following embodiments are only schematic representations of the basic concept of the present invention. Therefore, the drawings only show the components related to the present invention and are not drawn according to the actual number, shape and size of the components in the actual implementation. In the actual implementation, the form, quantity and proportion of each component can be arbitrarily changed, and the layout of the components may also be more complex.
[0039] Please see Figures 1-5 The present invention provides a dialogue communication method and system based on non-invasive brain-computer interface and multimodal AI large model according to the following embodiments. It belongs to the field of cross-application of brain-computer interface technology and intelligent care, and is mainly applied to the daily communication of hemiplegic patients, providing hemiplegic patients with autonomous and convenient life and social solutions.
[0040] like Figure 1 As shown, this dialogue communication method based on non-invasive brain-computer interface and multimodal AI large model operates on a dialogue communication system based on non-invasive brain-computer interface and multimodal AI large model. This dialogue communication system includes at least: A portable, non-invasive brain-computer interface device 1 is used to collect and preprocess the user's electroencephalogram (EEG) signals; Terminal device 2 is wirelessly connected to the brain-computer interface device and is used to present a visual stimulus interface, perform real-time intention decoding of brain signals, manage the interaction process, and process audio input and output. In addition, there is an AI large model processing unit, which communicates with the terminal device 2 to perform speech recognition, semantic understanding, natural language generation and text optimization tasks.
[0041] Preferably, the real-time intent decoding task of EEG signals is completed locally on the terminal device, and the AI large model processing unit is deployed on a cloud server 3 or a localized high-performance terminal device (such as a high-end laptop).
[0042] Preferably, the portable non-invasive brain-computer interface device 1 includes at least a main control module 15 (main control chip), a head-mounted electrode module 11, a signal acquisition module 12 (analog front-end chip), a wireless transmission module 13 (Bluetooth module or WiFi module), and a power supply module 14. The head-mounted electrode module 11 uses electrode cotton and saline solution, fitting comfortably against the scalp and supporting quick wear; the signal acquisition module 12 is responsible for acquiring EEG signals and filtering environmental noise; data interaction with the terminal device 2 is achieved through the wireless transmission module 13; the power supply module 14 uses a lithium battery, supporting fast charging and long battery life to meet daily usage needs; the main control module 15 is electrically connected to the signal acquisition module 12, the wireless transmission module 13, and the power supply module 14, used to control each component and perform basic data processing functions. Other specific mechanical structures are described in existing patent literature with publication number CN121129273A.
[0043] like Figure 2 As shown, the software system architecture is as follows: it includes a signal preprocessing module, an EEG recognition and control module, an interaction module, an AI large-scale model processing module, and a speech generation module. These modules work together to achieve the entire workflow. Specifically: The signal preprocessing module, integrated into the brain-computer interface device, is used to filter and reduce noise in the acquired raw EEG signals. The EEG recognition control module and the interaction module are integrated into the terminal device; The EEG recognition control module is used to extract features and decode intent from preprocessed EEG signals to generate control commands; The interaction module is used to respond to control commands, provide and manage a visual interaction interface based on a preset stimulus-response paradigm, and the visual interaction interface is used to at least realize two-way dialogue and communication functions. The AI large model processing module is the core component of the AI large model processing unit. It is deployed in the cloud or locally on the terminal device and communicates with the interaction module and the speech generation module to perform semantic understanding and natural language generation tasks. The speech generation module is used to convert text information into speech signals and output them.
[0044] As a further technical solution of the present invention, the interaction module includes a two-way dialogue communication submodule, which is configured to support at least two interaction channels: The first interaction channel is configured to receive external voice input, call the AI large model processing module to generate at least one candidate response text, provide the user with a selection through the visual interaction interface, and then output it through the voice generation module. And / or, The second interaction channel is configured to provide character input functionality through a visual interaction interface, recognize the character sequence input by the user, call the AI large model processing module to perform semantic optimization on the character sequence, and then output the optimized text through the speech generation module.
[0045] The specific operating principle is as follows: When the system is started or triggered by user operation, the EEG signals collected by the brain-computer interface device are filtered by the signal preprocessing module to remove power frequency interference and electromyographic noise, and then transmitted to the terminal device through the wireless transmission module. The EEG recognition control module of the terminal device extracts the signal frequency features through the feature extraction algorithm and matches them with the preset visual stimulus frequency library to realize intent recognition. The terminal device displays the preset visual stimulus interface (corresponding to different function options). When the user gazes at the visual stimulus area corresponding to the target option, the system analyzes the corresponding generated EEG signals (SSVEP / P300) to identify and determine the user's intent and trigger the corresponding function module.
[0046] like Figure 3 As shown, the dialogue communication method based on non-invasive brain-computer interface and multimodal AI large model of the present invention includes at least the following steps: S1. The user wears and activates the non-invasive brain-computer interface device 1, which continuously collects the user's brain signals and establishes and maintains a wireless communication connection with the terminal device 2. S2. Terminal device 2 starts the application and displays a function selection interface containing at least one function option to the user. S3. By decoding EEG signals, identify the user's intention to select the "Dialogue Communication Scene" function option in the function selection interface; S4. Enter the dialogue and communication scene sub-module. The terminal device 2 presents the user with an interaction mode selection interface that includes at least two modes: "voice interaction" and "electroencephalogram text input". S5. Based on the user's selection in the decoded EEG signal recognition, proceed to the processing flow of the corresponding interaction mode.
[0047] Preferably, when the user selects the "voice interaction" mode, the following steps are performed: S6. Collect external voice information through a sound pickup device (which can be a microphone integrated into the terminal device 2 or an independent microphone device connected to the terminal device 2); S7. Send the external voice information or its converted text to the AI large model processing unit. The AI large model processing unit performs semantic understanding and generates at least one candidate response text. S8. Terminal device 2 displays a visually stimulating interface containing at least one candidate answer text. S9. The user gazes at the visual stimulus area corresponding to the candidate response text of the target, and the brain-computer interface device collects the corresponding electroencephalogram (EEG) signals. S10. The terminal device identifies the user's intention to select the target candidate response text based on the EEG signal. S11. Send the selected candidate answer text to the speech generation module; S12, The speech generation module converts text into speech signals and plays them through an audio output device.
[0048] Preferably, when the user selects the "EEG text input" mode, the following steps are performed: S13. The terminal device displays a character input interface, which contains multiple character selection areas encoded with specific visual stimuli. S14. The user inputs text by sequentially looking at the area corresponding to the target character. The system continuously decodes the EEG signals to identify the text sequence input by the user. S15. Send the text sequence input by the user to the AI large model processing unit for semantic optimization processing to generate optimized text; S16. Send the optimized text to the speech generation module; S17. The speech generation module converts the text into a speech signal and plays it through an audio output device (which can be a speaker integrated into the terminal device 2 or a separate sound playback device).
[0049] Preferably, the character input interface is a nine-key virtual keyboard interface, where each number key area corresponds to a group of letters or characters and is encoded by an independent visual stimulus; the input process includes the system confirming that the user is looking at the target number key area, and then making a secondary selection in the character group corresponding to that number key to determine the specific input character.
[0050] Preferably, the nine-key virtual keyboard interface can be replaced with a twenty-six-key virtual keyboard interface or a handwriting input virtual keyboard interface (achieved through gaze trajectory control and handwriting trajectory recognition) to meet the input habits and preferences of different users.
[0051] like Figure 4-5 As shown, the implementation method of visual stimulus display in brain-computer input methods (taking the nine-key input method as an example) can be as follows: 1. Initial state: "Mosaic masking" of the keyboard area: The system by default covers each key on the nine-key keyboard with a mosaic (checkerboard) mask. The purpose is to make each key flash at a specific frequency by showing and not showing it, thereby stimulating the brain to generate specific electrical signals.
[0052] 2. When the user is looking at the target button: The user's gaze is focused on a particular key (such as the "MNO" area on a nine-key keyboard, e.g.) Figure 5 As shown in the figure, at this time, the visual stimulation (the frequency of mosaic flashing) in this area will be perceived by the brain, generating corresponding visual evoked potentials (electroencephalogram signals).
[0053] The system collects this EEG signal in real time through a brain-computer interface, and then uses an algorithm to decode "which area the user is looking at".
[0054] 3. Triggering the disappearance of mosaic: When the mosaic area is white, the image being stimulated at that moment is not displayed. This is because the display of the mosaic needs to be controlled at a certain frequency to stimulate the brain to generate an electrical signal of the corresponding frequency.
[0055] Once the system confirms that "the user is looking at the 'MNO' area", it will remove the mosaic masking image of that area and replace it with a light blue image with a blue border around it, thus giving the user visual feedback (equivalent to "confirming that you have selected the correct area").
[0056] 4. Subsequent input process (taking a nine-key keyboard as an example): If the user needs to input "A", the selected image will change to a mosaic (checkerboard) image after 1 second. The user then looks at the mosaic flashing area of the "ABC" region. The system will recognize that the user has annotated the target ABC and repeat the process of "gazing → mask flashing → EEG decoding → confirmation of selection".
[0057] Preferably, the visual stimulus encoding in each interface is implemented based on the steady-state visual evoked potential (SSVEP) paradigm, and each selectable option or character area is assigned a unique visual flashing frequency; and when the system decodes and confirms the user's intent to gaze at a certain option or area, it immediately switches the visual stimulus of that area from the periodic flashing mode to a static visual confirmation feedback indicator. Its specific principle is as follows: SSVEP uses flashing stimulation to generate corresponding electrical signals in the brain, flashing at a fixed frequency, for example... Figure 4-5 The keyboard contains 24 targets, each key flashing at a fixed frequency. The flashing frequencies range from 8.9 Hz to 9.8 Hz from top to bottom and left to right, with a frequency difference of 0.3 Hz between each key. The system continuously collects EEG data for calculation. The algorithm determines which frequency a user is looking at, stimulating the brain to produce a corresponding EEG signal. Once the algorithm identifies an EEG frequency that matches a specific frequency, it confirms which key the user is looking at.
[0058] Preferably, during a conversation, it supports seamless switching between "voice interaction" mode and "brainwave text input" mode according to the user's intent; the AI large model processing unit maintains consistent contextual information throughout the conversation to ensure the continuity of communication.
[0059] Preferably, the EEG signal recognition process is implemented based on the event-related potential (e.g., P300) paradigm; the interaction mode selection interface and character input interface are constructed as a stimulus matrix, and the system uses the rows or columns of the matrix highlighted by random sequence as rare stimuli, and decodes the user's selection intention by detecting the event-related potential components in the EEG signal that are phase-locked with the rare stimuli.
[0060] Specifically, to use P300 to replace SSVEP for communication functions in hemiplegic users, the core is to design the interactive interface based on the Oddball Paradigm, decoding user intent by recognizing P300 potentials evoked by "low-probability target stimuli." The following is the specific implementation method: I. Core Principle: P300's "Rare Stimulus-Attention Response" Mechanism: P300 is induced by "interspersing low-probability target stimuli among ordinary stimuli": when a user focuses their attention on a target option, the corresponding stimulus (such as a row / column or a function button on the interface) will appear randomly with a low probability. At this time, the brain will generate a P300 potential. The system can identify the target selected by the user by detecting the temporal characteristics of this potential (the positive peak about 300ms after stimulation).
[0061] II. P300 Alternative Implementation Scheme for Dialogue and Communication Scenarios Interactive Interface Design: Functional Matrix Based on the P300 Speller Paradigm: The communication functions are integrated into a matrix-style visual interface (similar to a nine-key input method interface), with each function / option corresponding to a "target" in the matrix, and the P300 is triggered by "randomly highlighting rows / columns". The system randomly highlights the row or column containing the target option with a low probability (e.g., 10%-20%) (the rest are ordinary stimuli); When a user focuses on a target option, the brain generates a P300 potential when the row / column corresponding to the target is highlighted. The system then decodes this potential to locate the target option.
[0062] The dialogue communication method of this invention, based on a non-invasive brain-computer interface and a multimodal AI large model, enables two-way communication between users and caregivers. Method 1: A terminal device is placed near the user. It picks up the caregiver's speech via microphone, then intelligently provides a response displayed on the stimulation interface. The user confirms the response by gazing at the screen, and the system recognizes the user's chosen answer by identifying intent. Finally, the response is broadcast to the caregiver via voice. (Severe hemiplegia usually involves aphasia, making normal speech communication impossible.) Method 2: Users can also simulate typing, and the input text is recognized by SSVEP EEG signals. The text is then polished by an AI model and converted into speech output for caregivers.
[0063] Method 3: The AI model provides four relevant responses based on the caregiver's words, which the user can choose from. If the user feels that none of them are suitable, they can express their thoughts by brain-computer typing, and then the AI model will refine and improve the speech and broadcast it.
[0064] The following are specific embodiments of the present invention: Example 1: Dialogue and Communication between SSVEP and a Large Cloud Model Reference Figure 1The system hardware consists of a portable non-invasive brain-computer interface device 1 (a head-mounted EEG acquisition device), a terminal device 2 (a tablet computer), and a cloud server 3. The brain-computer interface device 1 uses comfortable dry electrodes and integrates an analog front-end (AFE) chip for amplification and filtering. It transmits the digitized EEG data stream to the terminal device 2 via Bluetooth 5.0.
[0065] Reference Figure 2 The application running on terminal device 2 contains core software modules. When the user launches the application and selects "Conversation," they enter a state similar to... Figure 3 The main flow is shown.
[0066] Example of voice interaction mode: The caregiver asks, "It's a bit chilly today, would you like to add a blanket?" (corresponding to step S6).
[0067] Terminal device 2's microphone picks up sound and uploads the audio to cloud server 3 (corresponding to step S7). A large language model deployed in the cloud (such as the GPT series, Wenxin Yiyan, etc.) recognizes the speech content, understands that it is an inquiry about "cold and hot perception and needs," and, combined with historical dialogues (such as the user being afraid of the cold), generates three candidate answers: A: "Okay, please add another blanket for me, thank you." B: "No need, I feel just right now." C: “I want to close the window a little smaller.” (Corresponding to step S7).
[0068] The candidate answers are displayed on the terminal device 2 as buttons that flash at three different frequencies (e.g., 10Hz, 12Hz, 15Hz) (corresponding to step S8).
[0069] The user felt that answer A best suited their preference, so they stared at the area of button A for about 1 second. Brain-computer interface device 1 collected a strong 10Hz SSVEP response, which was decoded by the CCA algorithm on terminal device 2 to confirm the selection of A (corresponding to steps S9 and S10).
[0070] Terminal device 2 sends text A to the local TTS engine (corresponding to step S11), synthesizes speech and plays it (corresponding to step S12).
[0071] Example of EEG text input mode: After the user selects this mode, terminal device 2 displays as follows: Figure 4 The nine-key input interface shown in the middle area has each numeric keypad covered by a mosaic pattern of corresponding frequency.
[0072] The user wants to express "I want to hear the news". First, look at the third row and fifth column (corresponding to WXYZ). After the system recognizes this, the mosaic on that key area immediately disappears and becomes highlighted in blue (e.g., ...). Figure 5 (As shown in the middle area), the user then looks at the second row, fifth column (corresponding to MNO), and selects when the word "I" appears. This process is repeated to complete the original text input "I want to listen to the news first" (corresponding to steps S13 and S14). This text is sent to the cloud LLM (corresponding to step S15). The LLM optimizes it based on context into grammatically correct and semantically complete "I want to listen to the news" (corresponding to step S16). The optimized text is returned and synthesized into speech for playback (corresponding to step S17).
[0073] Example 2: An Alternative Based on the P300 Paradigm For users sensitive to flickering, the P300 paradigm can be used. The dialogue function or input keyboard is presented in a matrix format. The system randomly and with low probability highlights a specific row or column (rare stimulus). When the user mentally selects a target (e.g., the "drink water" button), each time the target's row or column is highlighted, a positive potential (P300) is induced in the user's EEG approximately 300ms later. By using multiple consecutive stimulus sequences, superimposing the average EEG signal, and detecting the spatiotemporal characteristics of P300, the system can accurately locate the user's selected target. Subsequent interaction with the AI large-scale model is the same as in Example 1.
[0074] Example 3: Localized Deployment and Protocol Replacement For scenarios with poor network conditions or extremely high privacy requirements, a lightweight version of the multimodal large model can be deployed on high-performance terminal devices (such as high-end laptops) to form a local closed-loop system, eliminating network latency and dependency. The wireless transmission module can also be configured for the Wi-Fi protocol according to the actual home environment to achieve a longer and more stable transmission distance and higher data throughput.
[0075] This invention, through the aforementioned hardware and software collaborative communication system architecture and method, provides users with communication barriers with an unprecedented, efficient, intelligent, and natural means of assisting communication, and has significant practical value and social benefits.
[0076] This invention has at least the following advantages over the prior art: I. Breakthroughs in high efficiency and high dimension have been achieved in the human-computer interaction information transmission link, mainly reflected in: 1. Direct Intent Capture and Encoding: Utilizing paradigms such as SSVEP or P300, the user's neural response to specific visual stimuli (frequency, rare events) serves as an information carrier. This is equivalent to establishing a direct, high-bandwidth "neural command channel" between the user's brain and the machine. The user's "gaze selection" intent is encoded in real time into EEG signals with specific time-frequency characteristics for transmission, avoiding the multi-level dimensionality reduction and conversion losses of traditional methods.
[0077] 2. High-efficiency parallel instruction transmission: Through a specially designed visual stimulus matrix (such as a nine-key or function matrix), the system allows users to select from multiple parallel channels (corresponding to different flashing frequencies or stimulus sequences) on a single interface. In the SSVEP paradigm, stimuli of different frequencies can be regarded as multiple parallel frequency division multiplexing (FDM) channels; in the P300 paradigm, randomly highlighted rows / columns can be regarded as interrogation signals in time division multiplexing (TDM). The system decodes these channels in parallel using advanced signal processing algorithms (such as canonical correlation analysis (CCA) for SSVEP, and superposition averaging and classification algorithms for P300), achieving millisecond-level intent recognition and greatly improving the information transmission rate (ITR).
[0078] 3. Robust anti-interference communication design: This invention incorporates a signal preprocessing module to filter power frequency and electromyographic noise, effectively adding adaptive filtering and feature enhancement steps at the receiving end (EEG acquisition device) and front-end processing, thus improving the signal-to-noise ratio (SNR). By placing core AI semantic computation in the cloud and low-latency, high-determinism intent recognition at the terminal, this edge-cloud collaborative computing architecture optimizes the latency and reliability of the overall link.
[0079] II. End-to-end intelligent integration and stability improvement have been achieved in system integration and communication protocols, mainly reflected in: 1. An end-to-end intelligent communication system integrating neural signals and semantic information was constructed: This system deeply integrates brainwave decoding with large language models (LLM). The brain-computer interface, as a high-precision, low-latency command input terminal, is responsible for transmitting discrete commands (selection, confirmation, typing); the multimodal AI large model, as a high-performance, highly intelligent semantic processing and generation core, is responsible for understanding and generating continuous, open-domain semantics. The two exchange data through standardized API interfaces (such as cloud-based RESTful APIs), forming a complete "sensing-decision-execution" communication loop.
[0080] 2. An adaptive hybrid interaction protocol is adopted: The solution provides two modes: "voice interaction (AI-assisted selection)" and "brain-computer text input (AI-assisted polishing)," which can be seamlessly switched. This is equivalent to designing an adaptive hybrid modulation protocol: when the ambient noise is low and the caregiver's voice is clear, the system adopts the efficient mode of "speech recognition + limited option selection"; when the user has complex and personalized expression needs, the system switches to the high-degree-of-freedom mode of "EEG encoding input + semantic optimization." The large AI model acts as a protocol adaptation layer and encoding optimizer, dynamically adjusting the interaction strategy to adapt to different communication scenarios.
[0081] 3. Optimized wireless data transmission link: Bluetooth 5.0 is adopted as the main wireless transmission protocol. Its low power consumption and high connection stability are very suitable for short-range, continuous transmission of EEG signals with moderate data volume but requiring real-time performance. Moreover, Bluetooth mode can be replaced by Wi-Fi, which provides configurable physical layer and data link layer options for different scenarios (such as multi-room coverage in a home), enhancing the system's environmental adaptability and communication reliability.
[0082] Third, a qualitative leap has been achieved in user experience quality (QoE), mainly reflected in: 1. Significantly reduces cognitive and operational load: Users do not need to learn complex encoding rules (such as traditional EEG spelling devices), but only need to "gaze" at the target. The AI big model undertakes the most complex semantic organization work, freeing users from tedious text organization or deep menu navigation, effectively reducing the "brain bandwidth" requirements for communication.
[0083] 2. Provides a low-latency, smooth interactive experience: Local fast EEG decoding (SSVEP / P300 recognition) ensures low-latency response (milliseconds) at the command level, providing users with timely positive feedback (such as mosaic disappearance and color change). Although cloud-based AI semantic generation has some network latency, its processing runs in parallel with the user's intent confirmation process (AI is already calculating when the user makes a selection), and multiple optimized options are provided in advance on the answer selection interface, cleverly masking some of the processing latency, resulting in an overall smooth experience.
[0084] 3. Achieves highly fault-tolerant and personalized natural communication: Multiple candidate answers generated by the AI model provide a semantic-level error correction and reselection mechanism, allowing users to quickly make a second selection even if the initial choice is inaccurate. The AI "polishing" of brainwave typing acts like a powerful post-encoding error correction and enhancement processor, correcting ambiguous intentions in the input and expanding them into complete and appropriate sentences, greatly improving the final quality and personalization of communication.
[0085] In summary, from the perspective of communication algorithms and system integration, this invention is not simply a functional superposition of brain-computer interfaces and large AI models, but rather a deep reconstruction of the "communication protocol stack": At the physical / data link layer, high signal-to-noise ratio EEG features (SSVEP / P300) are used as reliable signal carriers, and wireless transmission is optimized.
[0086] At the network / application layer, an efficient edge-cloud collaborative computing architecture and an adaptive hybrid interaction protocol were designed.
[0087] In the representation layer, a multimodal AI large model is introduced as an intelligent encoder and decoder, which greatly improves the efficiency and richness of semantic information expression.
[0088] Ultimately, this invention constructs a novel human-computer interaction communication system with high bandwidth, low latency, high intelligence, and strong robustness, fundamentally solving the core defects of existing technologies such as insufficient autonomy, low interaction efficiency, poor device compatibility, and unsatisfactory communication experience, providing people with mobility impairments with an efficient, natural, and dignified channel for external communication.
[0089] The above embodiments are merely illustrative of the principles and effects of the present invention and are not intended to limit the invention. Any person skilled in the art can modify or alter the above embodiments without departing from the spirit and scope of the present invention. Therefore, all equivalent modifications or alterations made by those skilled in the art without departing from the spirit and technical concept disclosed in the present invention should still be covered by the claims of the present invention.
Claims
1. A dialogue communication method based on a non-invasive brain-computer interface and a multimodal AI large model, characterized in that, At least the following steps are included: S1. The user wears and activates a non-invasive brain-computer interface device (1), which continuously collects the user's brain signals and establishes and maintains a wireless communication connection with the terminal device (2). S2. The terminal device (2) starts the application and displays a function selection interface containing at least one function option to the user. S3. By decoding the EEG signal, identify the user's intention to select the "Dialogue Communication Scene" function option in the function selection interface; S4. Enter the dialogue and communication scene sub-module. The terminal device (2) presents the user with an interaction mode selection interface that includes at least two modes: "voice interaction" and "electroencephalogram text input". S5. Based on the user's selection in the decoded EEG signal recognition, proceed to the processing flow of the corresponding interaction mode.
2. The dialogue communication method based on non-invasive brain-computer interface and multimodal AI large model according to claim 1, characterized in that, When a user selects the "voice interaction" mode, the following steps are performed: S6. Collect external voice information through a sound pickup device; S7. Send the external voice information or its converted text to the AI large model processing unit. The AI large model processing unit performs semantic understanding and generates at least one candidate response text. S8. The terminal device (2) displays a visual stimulus interface containing at least one candidate answer text; S9. The user gazes at the visual stimulus area corresponding to the candidate answer text of the target, and the brain-computer interface device (1) collects the corresponding EEG signal. S10, Terminal device (2) identifies the user's intention to select target candidate response text based on EEG signals; S11. Send the selected candidate answer text to the speech generation module; S12, The speech generation module converts text into speech signals and plays them through an audio output device.
3. The dialogue communication method based on non-invasive brain-computer interface and multimodal AI large model according to claim 1, characterized in that, When the user selects the "EEG text input" mode, the following steps are performed: S13. The terminal device (2) displays a character input interface, which contains multiple character selection areas encoded with specific visual stimuli. S14. The user inputs text by sequentially looking at the area corresponding to the target character. The system continuously decodes the EEG signals to identify the text sequence input by the user. S15. Send the text sequence input by the user to the AI large model processing unit for semantic optimization processing to generate optimized text; S16. Send the optimized text to the speech generation module; S17. The speech generation module converts text into speech signals and plays them through an audio output device.
4. The dialogue communication method based on non-invasive brain-computer interface and multimodal AI large model according to any one of claims 1-3, characterized in that, The visual stimulus encoding in each interface is implemented based on the steady-state visual evoked potential paradigm. Each selectable option or character area is assigned a unique visual flashing frequency. Furthermore, once the system decodes and confirms the user's intention to gaze at a certain option or area, it immediately switches the visual stimulus of that area from the periodic flashing mode to a static visual confirmation feedback indicator.
5. The dialogue communication method based on non-invasive brain-computer interface and multimodal AI large model according to claim 3, characterized in that, The EEG signal recognition process is based on the event-related potential paradigm. The interaction mode selection interface and character input interface are constructed as a stimulus matrix. The system highlights rows or columns of the matrix as rare stimuli through random sequences and decodes the user's selection intent by detecting the event-related potential components in the EEG signal that are phase-locked with the rare stimuli.
6. The dialogue communication method based on non-invasive brain-computer interface and multimodal AI large model according to claim 3, wherein the character input interface is a nine-key virtual keyboard interface, each number key area corresponds to a group of letters or characters and is encoded by independent visual stimuli; the input process includes the system confirming that the user is looking at the target number key area, and then making a secondary selection in the character group corresponding to the number key to determine the specific input character.
7. In the dialogue communication method based on non-invasive brain-computer interface and multimodal AI large model according to claim 6, the nine-key virtual keyboard interface can be replaced by a twenty-six-key virtual keyboard interface or a handwriting input virtual keyboard interface.
8. A dialogue communication system based on a non-invasive brain-computer interface and a multimodal AI large-scale model, used to implement the dialogue communication method based on a non-invasive brain-computer interface and a multimodal AI large-scale model as described in any one of claims 1-7, characterized in that, This dialogue and communication system includes at least: A portable non-invasive brain-computer interface device (1) for acquiring and preprocessing the user's electroencephalogram (EEG) signals; The terminal device (2) is wirelessly connected to the brain-computer interface device (1) and is used to present a visual stimulus interface, perform real-time intention decoding of EEG signals, manage the interaction process, and process audio input and output. In addition, the AI large model processing unit is connected to the terminal device (2) for performing speech recognition, semantic understanding, natural language generation and text optimization tasks.
9. The dialogue communication system based on non-invasive brain-computer interface and multimodal AI large model according to claim 8, characterized in that, The system integrates a signal preprocessing module, an EEG recognition and control module, an interaction module, an AI large model processing module, and a speech generation module. in, The signal preprocessing module is integrated into the brain-computer interface device (1) and is used to filter and reduce noise in the acquired raw EEG signals. The EEG recognition control module and the interaction module are integrated in the terminal device (2); The EEG recognition control module is used to extract features and decode intent from preprocessed EEG signals to generate control commands. The interaction module is used to respond to the control command, provide and manage a visual interaction interface based on a preset stimulus-response paradigm, and the visual interaction interface is used to at least realize two-way dialogue and communication functions. The AI large model processing module is the core component of the AI large model processing unit. It is deployed in the cloud or on the local terminal device (2) and communicates with the interaction module and the speech generation module to perform semantic understanding and natural language generation tasks. The speech generation module is used to convert text information into speech signals and output them.
10. The dialogue communication system based on a non-invasive brain-computer interface and a multimodal AI large model according to claim 9, characterized in that, The interaction module includes a two-way dialogue communication submodule, which is configured to support at least two interaction channels: The first interaction channel is configured to: receive external voice input, call the AI large model processing module to generate at least one candidate response text, provide the user with a selection through the visual interaction interface, and then output it through the voice generation module. And / or, The second interaction channel is configured to: provide character input function through the visual interaction interface, recognize the character sequence input by the user, call the AI large model processing module to perform semantic optimization on the character sequence, and then output the optimized text through the speech generation module.