Cooperative control method, device and system, storage medium and program product

By employing a collaborative control method combining visual semantic encoding and EEG encoding, the problems of poor signal quality and strong environmental dependence in brain-computer interface technology have been solved, enabling efficient and real-time multi-command control, thereby improving user experience and system adaptability.

CN120973228APending Publication Date: 2025-11-18BEIJING JI MASCH TECH CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202511078895.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-01
Publication Date
2025-11-18

AI Technical Summary

Technical Problem

Existing brain-computer interface technologies suffer from problems such as poor signal quality, high training and calibration costs, high latency, low dimensionality, high cognitive load, strong environmental dependence, poor comfort, and poor mobility, making it difficult to meet the real-time interaction needs of disabled people.

Method used

A collaborative control method is adopted, which uses a preset visual instruction library to encode visual and EEG signals using a visual semantic encoder and an EEG encoder. Combined with contrastive learning, the signals are mapped to a shared semantic space to achieve efficient recognition and control of multiple instructions.

Benefits of technology

It achieves brain-computer interface control with low learning cost, real-time interaction, rich instructions, high comfort, strong environmental adaptability, low cognitive load, low system cost, and high mobility, improving the accuracy and stability of signal recognition.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120973228A_ABST
    Figure CN120973228A_ABST
Patent Text Reader

Abstract

The invention discloses a cooperative control method, device and system, a storage medium and a program product. The method comprises the following steps: presetting a visual instruction library containing a plurality of visual instruction symbols; extracting image features of the visual instruction symbols; inputting the image features into a visual semantic encoder for encoding to obtain n-dimensional visual semantic vectors; acquiring an original EEG signal of a human body; eEG spatial features of the original EEG signals are extracted; inputting the EEG spatial features into an EEG encoder for encoding to obtain an n-dimensional neural response vector; mapping the n-dimensional visual semantic vector and the n-dimensional neural response vector to a shared semantic space; performing comparative learning and target optimization based on the n-dimensional visual semantic vector and the n-dimensional neural response vector to obtain a matched visual-EEG pair vector; and performing instruction classification on the vector based on vision-EEG to obtain instruction probability distribution, and outputting a control instruction for controlling the external equipment based on the instruction probability distribution. The problems of high learning cost, low interaction real-time performance and the like in brain-computer interface control are solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the field of brain-computer interface technology, and in particular to a collaborative control method, device, system, storage medium, and program product that uses visual symbol stimulation to induce brain electrical responses and realize multi-command control. Background Technology

[0002] Brain-computer interface (BCI) is a technology that allows interaction with external devices directly through brain signals, without relying on peripheral nerves and muscle tissue. For patients with ALS, spinal cord injury, amyotrophic lateral sclerosis (ALS), or stroke, traditional medical methods often fail to restore their motor function. BCI technology, however, allows patients to regain mobility by controlling exoskeletons, wheelchairs, or prostheses through their thoughts.

[0003] Traditional control methods are based on motor imagery (MI), which refers to the generation of corresponding electroencephalograms (EEGs) by imagining a specific movement (such as the movement of the left hand, right hand, or foot) in the brain without actually performing the action. These signals can be recognized by the BCI system and converted into control commands, thereby enabling the manipulation of external devices. In practice, the EEG signals generated based on motor imagery are extremely weak (microvolts), easily interfered with by physiological activities such as blinking and electromyography, resulting in a low signal-to-noise ratio. Furthermore, users need to undergo extensive training (5-20 times, each time >30 minutes) to achieve usable control accuracy. Model parameters usually need to be recalibrated before each use, making the process cumbersome. Continuous imagery operations can easily lead to user fatigue, distraction, and increased error rates. Motor imagery-based brain-computer interface control systems have significant response delays (0.5-2 seconds), making it difficult to meet real-time interaction requirements; they can typically only stably recognize 2-3 simple commands (such as imagining the left / right hand), and their performance degrades significantly in complex and noisy environments such as outdoors.

[0004] Based on existing research, it is known that even in disabled individuals such as those with ALS (Amyotrophic Lateral Sclerosis), the eye muscles can still maintain long-term activity despite the loss of function in most of the body's muscles. Therefore, visually evoked EEG signals become a better option. There are already schemes for controlling external devices based on steady-state visual evoked potentials (SSVEPs). Steady-state visual evoked potentials (SSVEPs) are a special type of brain electrophysiological response. When the human eye fixates on a flashing light with a frequency of f (such as an LED or screen flash), the primary visual cortex (V1 area) of the occipital lobe of the brain will generate EEG oscillations synchronized with f and its harmonics (2f, 3f, etc.). For example, when fixating on a 10Hz flashing stimulus, the EEG signal shows energy peaks at frequencies such as 10Hz and 20Hz. The amplitude of SSVEPs usually decreases as the stimulus frequency increases, with the optimal response frequency range being 5-30Hz. By processing SSVEP signals, classification results can be converted into control commands (such as cursor movement, device on / off, etc.) based on the visual target. When using steady-state visual evoked potentials (SSVEP) for external control, users need to stare at high-frequency (especially >30Hz) flickering stimuli for extended periods, which can easily cause visual fatigue, headaches, nausea, and other discomfort. Strong ambient light can overwhelm the stimulus signal, usually requiring additional light shielding or the use of a high-brightness stimulus source. Limited by the resolvable frequency / phase interval, only 4-8 commands can typically be implemented. Users must maintain a high level of concentration while staring at the target stimulus source; distraction leads to a sharp attenuation of the signal. Visual stimulation must be provided by a fixed, dedicated display or LED array, severely limiting its application in mobile scenarios.

[0005] Therefore, there is a need for a more accurate, comfortable, adaptable, and cost-effective visual brain-computer interface control method.

[0006] The above information is presented as background information only to aid in understanding this disclosure. No confirmation or other representation is made regarding whether any of the above constitutes an application of prior art to this disclosure. Summary of the Invention

[0007] This disclosure addresses the problems of poor signal quality, high training and calibration costs, high latency, low dimensionality, high cognitive load, strong environmental dependence, poor comfort, and poor mobility in the aforementioned brain-computer interfaces. It provides a collaborative control method with low learning costs, real-time interaction, rich instructions, high comfort, strong environmental adaptability, low cognitive load, low system cost, and high mobility. It also provides a collaborative control device to address the problem of controlling peripheral devices based on visual signals. Furthermore, it provides a collaborative control system, a non-transitory storage medium, and a computer program product to address the problem of executing the collaborative control method.

[0008] The first aspect of this disclosure provides a collaborative control method, comprising: pre-setting a visual instruction library, the visual instruction library containing multiple visual instruction symbols; performing image preprocessing on the visual instruction symbols to extract image features of the visual instruction symbols; inputting the image features into a visual semantic encoder for encoding to obtain an n-dimensional visual semantic vector; simultaneously acquiring raw EEG signals of the human body while the human visually fixates on a visual instruction symbol in the visual instruction library; performing data preprocessing on the raw EEG signals to extract EEG spatial features of the raw EEG signals; inputting the EEG spatial features into an EEG encoder for encoding to obtain an n-dimensional neural response vector; mapping the n-dimensional visual semantic vector and the n-dimensional neural response vector to a shared semantic space; performing comparative learning and target optimization based on the n-dimensional visual semantic vector and the n-dimensional neural response vector to obtain a matching visual-EEG pair vector; classifying instructions based on the visual-EEG pair vector to obtain an instruction probability distribution; and outputting control instructions for controlling external devices based on the instruction probability distribution.

[0009] For example, in at least one embodiment, the EEG spatial features include multiple EEG signal tensor dimensions, wherein the tensor dimensions include timing length, number of channels, and signal features.

[0010] For example, in at least one embodiment, obtaining the n-dimensional neural response vector includes the step of processing the EEG spatial features using a Transformer encoder.

[0011] For example, in at least one embodiment, the contrastive learning, based on a loss function, causes similar n-dimensional visual semantic vectors and n-dimensional neural response vectors to move closer together in the shared semantic space, while dissimilar n-dimensional visual semantic vectors and n-dimensional neural response vectors move further apart in the shared semantic space.

[0012] For example, in at least one embodiment, the loss function is defined as:

[0013] L = -log[exp(sim(v)] i ,e i +) / τ) / ∑ j K exp(sim(v i ,e j -) / τ)],where,

[0014] L is the loss function, v i Let e ​​be the response vector of the i-th neural network. i + represents the positive sample of the i-th visual semantic vector, e j- represents the j-th negative sample of the visual semantic vector, and K represents the number of negative samples of the visual semantic vector.

[0015] For example, in at least one embodiment, the contrastive learning is based on multiple pre-trainings of a deep learning model to classify matching visual-EEG vector pairs.

[0016] A second aspect of this disclosure provides a collaborative control device, comprising: a visual instruction library, an acquisition module, and a deep learning module; the visual instruction library contains multiple visual instruction symbols; the acquisition module is configured to acquire raw EEG signals of the human body while the human visually fixates on a visual instruction signal in the visual instruction library; the deep learning module includes an image extraction layer, a visual semantic encoder, a spatial feature extraction layer, an EEG encoder, a cross-modal projection layer, a contrast learning layer, and an output layer; the image extraction layer is configured to perform image preprocessing on the visual instruction symbols to extract image features of the visual instruction symbols; the visual semantic encoder is configured to encode the image features to obtain an n-dimensional visual signal. The system comprises: a visual semantic vector; a spatial feature extraction layer configured to preprocess the original EEG signal and extract its EEG spatial features; an EEG encoder configured to encode the EEG spatial features to obtain an n-dimensional neural response vector; a cross-modal projection layer mapping the n-dimensional visual semantic vector and the n-dimensional neural response vector to a shared semantic space; a contrastive learning layer performing contrastive learning and target optimization based on the n-dimensional visual semantic vector and the n-dimensional neural response vector to obtain a matching visual-EEG pair vector; and an output layer performing instruction classification based on the visual-EEG pair vector to obtain an instruction probability distribution, and outputting control instructions based on the instruction probability distribution.

[0017] For example, in at least one embodiment, the EEG spatial features include multiple EEG signal tensor dimensions, wherein the tensor dimensions include temporal length, number of channels, and signal features; the visual semantic encoder and the EEG encoder are Transformer encoders; the contrastive learning, based on a loss function, causes similar n-dimensional visual semantic vectors and n-dimensional neural response vectors to move closer together in the shared semantic space, and dissimilar n-dimensional visual semantic vectors and n-dimensional neural response vectors to move further apart in the shared semantic space; the loss function is defined as:

[0018] L = -log[exp(sim(v)] i ,e i +) / τ) / ∑ j K exp(sim(v i ,e j -) / τ)],

[0019] Where L is the loss function, v i Let e ​​be the response vector of the i-th neural network. i + represents the positive sample of the i-th visual semantic vector, e j - represents the j-th visual semantic vector negative sample, K represents the number of visual semantic vector negative samples, and τ represents a hyperparameter; the contrastive learning layer is configured to classify matching visual-EEG pairs based on multiple prior pre-training; it also includes an external device configured to perform control based on the logically optimized control instructions.

[0020] A third aspect of this disclosure provides a cooperative control system, comprising: a memory for non-temporarily storing computer-executable instructions; and a processor for running the computer-executable instructions, characterized in that the computer-executable instructions, when run by the processor, execute the cooperative control method described in any of the preceding claims.

[0021] The fourth aspect of this disclosure provides a non-transitory storage medium for non-transitory storage of computer-executable instructions, characterized in that, when the computer-executable instructions are executed by a computer, the cooperative control method described in any of the preceding claims is executed.

[0022] The fifth aspect of this disclosure provides a computer program product, including a computer program, characterized in that, when executed by a processor, the computer program implements the cooperative control method described in any of the preceding claims. Attached Figure Description

[0023] To more clearly illustrate the technical solutions of the embodiments of this disclosure, the accompanying drawings of the embodiments will be briefly described below. Obviously, the drawings described below only relate to some embodiments of this disclosure, and are not intended to limit this disclosure.

[0024] Figure 1 This is a flowchart of a collaborative control method according to an embodiment of the present disclosure;

[0025] Figure 2 This is a schematic diagram of a visual instruction library and signal board according to embodiments of the present disclosure;

[0026] Figure 3 This is a schematic diagram of a collaborative control device according to an embodiment of the present disclosure;

[0027] Figure 4 This is a schematic diagram of a deep learning module according to an embodiment of the present disclosure;

[0028] Figure 5 This is a schematic diagram of a second type of cooperative control device according to an embodiment of the present disclosure. Detailed Implementation

[0029] To make the objectives, technical solutions, and advantages of the embodiments of this disclosure clearer, the technical solutions of the embodiments of this disclosure will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this disclosure. All other embodiments obtained by those skilled in the art based on the described embodiments of this disclosure without creative effort are within the scope of protection of this disclosure.

[0030] Unless otherwise defined, the technical or scientific terms used in this disclosure shall have the ordinary meaning understood by one of ordinary skill in the art to which this disclosure pertains. The terms “first,” “second,” and similar terms used in this disclosure do not indicate any order, quantity, or importance, but are merely used to distinguish different components. Terms such as “comprising” or “including” mean that the element or object preceding the word encompasses the elements or objects listed following the word and their equivalents, without excluding other elements or objects. Terms such as “connected” or “linked” are not limited to physical or mechanical connections, but can include electrical connections, whether direct or indirect. Terms such as “upper,” “lower,” “left,” and “right” are used only to indicate relative positional relationships, and these relative positional relationships may change accordingly when the absolute position of the described objects changes. In this disclosure, “multiple” means two or more.

[0031] According to embodiments of this disclosure, Figure 1 A collaborative control method is disclosed, comprising the following steps:

[0032] S1011: Preset visual instruction library, which contains multiple visual instruction symbols. For example... Figure 2 The visual instruction library shown includes six visual instruction symbols: front 21, back 22, left 23, right 24, stop 25, and jump 26. Figure 2 The command signals in the library are primarily designed for controlling mobile devices and game interfaces. Understandably, the visual command library can include various types of command symbols; for example, a visual command library might contain seven visual command symbols, each representing a syllable, used to control musical performance. Setting visual command symbols according to requirements improves the specificity of control while conserving computational resources.

[0033] S1012: Perform image preprocessing on the visual instruction symbols to extract their image features. Image preprocessing includes removing irrelevant information interference, such as removing noise from the image; enhancing effective features, such as strengthening edges; and standardizing data specifications, such as normalizing the size of multiple visual instruction symbols. The extracted image features of the instruction symbols include the image's size, shape, color, texture, brightness, etc., for example... Figure 2In the visual instruction symbols shown, the first 21 is set to a gradient green, the last 22 to a gradient blue, the left 23 to a gradient yellow, the right 24 to a gradient purple, the stop 25 to a gradient red, and the jump 26 to a gradient cyan. The use of multiple image features helps improve symbol recognition and enhances the discriminative power of EEG signals.

[0034] S1013: Image features are input into the visual semantic encoder for encoding, obtaining an n-dimensional visual semantic vector. The visual semantic encoder adopts a Transformer encoder architecture, consisting of multiple identical encoder layers stacked together. Each layer contains two main modules: a multi-head self-attention mechanism and a feed-forward neural network. The multi-head attention mechanism calculates attention separately after segmenting the input and then concatenates the results. This is beneficial for capturing features in different subspaces. The output of each encoder layer passes through the feed-forward neural network, which transforms the representation at each position, facilitating further feature extraction.

[0035] S1021: While the human body is visually fixating on a visual instruction symbol in the visual instruction library, the raw EEG signal of the human body is acquired. Different visual instruction symbols will produce different raw EEG signals. For example, when the human body visually fixates on... Figure 2 The first 21 in the sequence will generate corresponding EEG signals; when the human eye focuses on... Figure 2 The EEG signal corresponding to the last 22 in the original EEG signal will be generated; the EEG signal corresponding to the first 21 is different from the EEG signal corresponding to the last 22. Therefore, by identifying the visual instruction symbol corresponding to the original EEG signal, control can be achieved by focusing on a specific visual instruction symbol.

[0036] S1022: Perform data preprocessing on the raw EEG signal to extract its EEG spatial features. Data preprocessing of the raw EEG signal includes signal filtering, amplification, and other processing. Extracting the EEG spatial features of the raw EEG signal includes extracting features such as the time sequence length, the corresponding number of channels, and the feature dimension.

[0037] S1023: The EEG spatial features are input into the EEG encoder for encoding to obtain an n-dimensional neural response vector. The EEG encoder adopts a Transformer encoder architecture, consisting of multiple identical encoder layers stacked together. Each layer contains two main modules: a multi-head self-attention mechanism and a feed-forward neural network. The multi-head attention mechanism calculates attention separately after input segmentation and then concatenates the results. This is beneficial for capturing features in different subspaces. The output of each encoder layer passes through the feed-forward neural network, which transforms the representation at each position, facilitating further feature extraction. Using the Transformer encoder to process EEG temporal signals overcomes the shortcomings of traditional methods in capturing temporal features.

[0038] S103: Map the n-dimensional visual semantic vector and the n-dimensional neural response vector to a shared semantic space. Drawing inspiration from the image-text alignment concept of CLIP, a shared semantic space is constructed for visual semantics and neural responses. In some embodiments, the n-dimensional visual semantic vector and the n-dimensional neural response vector are low-dimensional dense vectors representing their respective objects. In the shared semantic space, the distance between the n-dimensional visual semantic vector and the n-dimensional neural response vector reflects their similarity.

[0039] S104: Contrastive learning and target optimization are performed based on n-dimensional visual semantic vectors and n-dimensional neural response vectors to obtain matching visual-EEG pairs. Contrastive learning, based on a loss function, encourages similar n-dimensional visual semantic vectors and n-dimensional neural response vectors to move closer together in the shared semantic space, while dissimilar n-dimensional visual semantic vectors and n-dimensional neural response vectors move further apart in the shared semantic space. In some embodiments, cosine similarity, Euclidean distance, etc., are used to measure the similarity between samples. In a batch of data, for each sample, one positive sample and multiple negative samples are constructed. For example, for a gaze... Figure 2 For the n-dimensional neural response vectors corresponding to the first 21 visual instruction symbols in the shared semantic space, the positive samples constructed are the n-dimensional visual semantic vectors corresponding to the first 21 visual instruction symbols, and the five negative samples are the n-dimensional visual semantic vectors corresponding to the last 22, left 23, right 24, stop 25, and jump 26. Target optimization aims to maximize the similarity of matched visual-EEG pairs and minimize the similarity of unmatched pairs.

[0040] S105: Based on the visual-EEG, perform instruction classification on the vector to obtain the instruction probability distribution. Based on the instruction probability distribution, output control instructions for controlling external devices. For example... Figure 2The visual command symbols shown are of 6 types, and the corresponding visual-EEG pair vectors are also of 6 types. After comparative learning and target optimization, the probability distribution of the visual command symbols corresponding to the visual-EEG pair vectors can be obtained. If the probability of the top 21 visual command symbols corresponding to the visual-EEG pair vector is significantly greater than that of the other 5 visual command symbols, it can be determined that the visual command symbol corresponding to the visual-EEG pair vector is one of the top 21. At this time, a forward control command is output to control the external device to move forward.

[0041] Preferably, in at least one embodiment, the loss function is defined as:

[0042] L = -log[exp(sim(vi,ei+) / τ) / ∑jKexp(sim(vi,ej-) / τ)], where,

[0043] L is the loss function, vi is the i-th neural response vector, ei+ is the i-th positive sample of the visual semantic vector, ej- is the j-th negative sample of the visual semantic vector, K is the number of negative samples of the visual semantic vector, and τ is the hyperparameter.

[0044] Preferably, in at least one embodiment, contrastive learning, based on multiple pre-training iterations of a deep learning model, can classify matching visual-EEG vector pairs. During training, hyperparameter values ​​and the number of negative samples are dynamically adjusted. At the beginning of training, larger hyperparameter values ​​are set to smoother the similarity distribution, which helps the model quickly learn initial features. In subsequent stages, the hyperparameter values ​​are decreased gradient-wise, which helps the model distinguish samples more precisely. The number of negative samples can be dynamically increased or decreased according to the model's training state and data characteristics, which helps balance training efficiency and model performance.

[0045] According to embodiments of this disclosure, Figure 3 A collaborative control device 1 is disclosed, comprising: a visual instruction library 2, an acquisition module 3, and a deep learning module 4. The visual instruction library 2 contains multiple visual instruction symbols 21-25. The images of the visual instruction symbols 21-25 in the visual instruction library 2 are input into the deep learning module 4.

[0046] The acquisition module 3 is installed in the human brain. For example, in some embodiments, the acquisition module 3 is a head-mounted non-invasive brain-computer interface signal acquisition device. For example, in some embodiments, the acquisition module 3 is a multi-channel EEG device that acquires and transmits the user's EEG signals. For example, in some embodiments, the acquisition module 3 has a sampling rate of 500 Hz, 8-16 channels, a common-mode rejection ratio >110 dB, electrode points that are distributed as widely as possible, and real-time transmission.

[0047] A signal board 31 with visual instruction symbols 21-25 is set up in the human visual environment. When the human visually fixates on a visual instruction signal on the signal board 31, such as visual instruction signal 21, the acquisition module 3 acquires the raw EEG signal of the human body. The acquisition module 3 can start acquiring raw EEG signals when the human body begins to fixate, which helps to save acquisition resources. The acquisition module 3 can also continuously acquire raw EEG signals of the human body. The subsequent deep learning module 4 extracts the spatial features of the corresponding visual instruction signal from the continuously acquired raw EEG signals, which helps to enhance the continuity and integrity of the acquisition. The raw EEG signals acquired by the acquisition module 3 are input into the deep learning module 4.

[0048] like Figure 4 As shown, deep learning module 4 includes an image extraction layer 41, a visual semantic encoder 42, a spatial feature extraction layer 43, an EEG encoder 44, a cross-modal projection layer 45, a contrastive learning layer 46, and an output layer 47. The execution of deep learning module 4 can be implemented using any hardware system capable of performing deep learning, such as a computer, microprocessor, etc.

[0049] Image extraction layer 41 is configured to perform image preprocessing on visual instruction symbols 21-25 respectively, extracting image features of the visual instruction symbols. Image preprocessing includes removing irrelevant information interference, such as removing noise from the image; enhancing effective features, such as strengthening edges; and standardizing data specifications, such as normalizing the size of multiple visual instruction symbols. The extracted image features of the instruction symbols include image size, shape, texture, color, brightness, etc., for example... Figure 3 In the visual instruction symbols shown, the first 21 is set to a gradient green, the last 22 to a gradient blue, the left 23 to a gradient yellow, the right 24 to a gradient purple, and the stop 25 to a gradient red. The use of multiple image features helps improve symbol recognition and enhances the discriminative power of EEG signals.

[0050] Preferably, visual instruction symbols 21-25 are dynamic images or dynamic effects, and the corresponding image preprocessing is video preprocessing for dynamic images or dynamic effects.

[0051] The visual semantic encoder 44 encodes the image features of the visual instruction symbols 21-25 respectively, obtaining an n-dimensional visual semantic vector. The visual semantic encoder 44 adopts a Transformer encoder architecture, consisting of multiple identical encoder layers stacked together. Each layer contains two main modules: a multi-head self-attention mechanism and a feed-forward neural network. The multi-head attention mechanism calculates attention separately after input segmentation and then concatenates the results. This is beneficial for capturing features in different subspaces. The output of each encoder layer passes through the feed-forward neural network, which transforms the representation at each position, facilitating further feature extraction.

[0052] The spatial feature extraction layer 43 is configured to preprocess the raw EEG signal, extracting its spatial features. When a human gazes at a visual instruction symbol on the signal panel 31, such as visual instruction symbol 21, the human body will generate a raw EEG signal with specific features. The raw EEG signal generated will differ depending on the visual instruction symbol the human body gazes at. Extracting the spatial features of the raw EEG signal includes extracting features such as the temporal length, the corresponding number of channels, and the feature dimension.

[0053] The EEG encoder 44 is configured to encode EEG spatial features to obtain an n-dimensional neural response vector. The EEG encoder employs a Transformer encoder architecture, consisting of multiple identical encoder layers stacked together. Each layer contains two main modules: a multi-head self-attention mechanism and a feed-forward neural network. The multi-head attention mechanism calculates attention separately for each input segment and then concatenates the results. This is beneficial for capturing features in different subspaces. The output of each encoder layer passes through the feed-forward neural network, which transforms the representation at each position, facilitating further feature extraction. Using the Transformer encoder to process EEG temporal signals overcomes the shortcomings of traditional methods in capturing temporal features.

[0054] A cross-modal projection layer 45 maps the n-dimensional visual semantic vector and the n-dimensional neural response vector to a shared semantic space. Drawing inspiration from the image-text alignment concept of CLIP, a shared semantic space for visual semantics and neural responses is constructed. In some embodiments, the n-dimensional visual semantic vector and the n-dimensional neural response vector are low-dimensional dense vectors representing their respective objects. In the shared semantic space, the distance between the n-dimensional visual semantic vector and the n-dimensional neural response vector reflects their similarity.

[0055] Contrastive learning layer 46 performs contrastive learning and target optimization based on n-dimensional visual semantic vectors and n-dimensional neural response vectors to obtain matching visual-EEG pairs. Contrastive learning uses a loss function to encourage similar samples to move closer together in the shared semantic space, and dissimilar samples to move further apart. Cosine similarity and Euclidean distance are typically used to measure the similarity between samples. In a batch of data, for each sample, one positive sample and multiple negative samples are constructed. For example, for a gaze... Figure 2 For the n-dimensional neural response vectors corresponding to the first 21 visual instruction symbols in the shared semantic space, the positive samples constructed are the n-dimensional visual semantic vectors corresponding to the first 21 visual instruction symbols, and the five negative samples are the n-dimensional visual semantic vectors corresponding to the last 22, left 23, right 24, stop 25, and jump 26. Target optimization aims to maximize the similarity of matched visual-EEG pairs and minimize the similarity of unmatched pairs.

[0056] Output layer 47 performs instruction classification on the vectors based on visual-EEG, obtains the instruction probability distribution, and outputs control instructions based on the instruction probability distribution. For example... Figure 3 The visual command symbols shown are of 5 types, and the corresponding visual-EEG pair vectors are also of 5 types. After comparative learning and target optimization, the probability distribution of the visual command symbols corresponding to the visual-EEG pair vectors can be obtained. If the probability of the top 21 visual command symbols corresponding to the visual-EEG pair vector is significantly greater than that of the other 5 visual command symbols, it can be determined that the visual command symbol corresponding to the visual-EEG pair vector is one of the top 21. At this time, a forward control command is output to control the external device 5 to move forward.

[0057] Preferred, in Figure 3 In this embodiment, the external device 5 is a smart wheelchair. Visual command symbols on the human gaze signal panel 31, for example, upon gazing at a visual command symbol 22, generate EEG signals with specific characteristics. The acquisition module 3 acquires these EEG signals with specific characteristics to obtain raw EEG signals. The deep learning module 4 performs deep learning and outputs a "backward" control command based on this raw data. Therefore, the human body moves backward by gazing... Figure 3 The five visual command symbols can control the smart wheelchair to move forward, backward, turn left, turn right, and stop.

[0058] Preferred, in Figure 5In this embodiment, the visual instruction library 2 includes six categories of visual instruction symbols: forward (21), backward (22), left (23), right (24), stop (25), and jump (26). Correspondingly, there are also six categories of visual-EEG pair vectors. The external device 5 is the game interface. Visual instruction symbols on the human gaze signal board 32, such as the gaze-on visual instruction symbol backward (22), generate EEG signals with specific characteristics. The acquisition module 3 acquires these EEG signals with specific characteristics to obtain raw EEG signals. The deep learning module 4 performs deep learning and outputs a "backward" control command based on this raw data. Therefore, the human body, through gaze... Figure 3 The five visual command symbols in the game can control the forward, backward, left, right, stop, and jump movements of the controlled object in the game interface.

[0059] Preferably, when setting specific control logic, the control logic is optimized according to the controlled object. Figure 3 In the illustrated embodiment, when the duration of continuous gaze at a visual instruction symbol reaches a time threshold (e.g., 3 seconds), the external device 5, the smart wheelchair, executes the control command corresponding to the visual instruction symbol. This logical adjustment helps avoid misoperation and ensures the safety of the smart wheelchair. Figure 5 In the illustrated embodiment, control commands are acquired in real time, and external device 5 executes the control commands corresponding to the visual command symbols in real time. This logical adjustment helps to improve the gaming experience.

[0060] Preferably, the loss function of the contrastive learning layer 46 is defined as:

[0061] L = -log[exp(sim(v)] i ,e i +) / τ) / ∑ j K exp(sim(v i ,e j -) / τ)],where,

[0062] L is the loss function, v i Let e ​​be the response vector of the i-th neural network. i + represents the positive sample of the i-th visual semantic vector, e j - represents the j-th negative sample of the visual semantic vector, K represents the number of negative samples of the visual semantic vector, and τ is a hyperparameter.

[0063] Preferably, in at least one embodiment, the contrast learning layer 46, based on multiple pre-training iterations of the deep learning model 4, is able to classify matching visual-EEG vector pairs. During training, the hyperparameter values ​​and the number of negative samples are dynamically adjusted. At the beginning of training, larger hyperparameter values ​​are set to smoother the similarity distribution, which helps the model quickly learn initial features. In subsequent stages, the hyperparameter values ​​are decreased gradient-wise, which helps the model distinguish samples more precisely. The number of negative samples can be dynamically increased or decreased according to the model's training state and data characteristics, which helps balance training efficiency and model performance.

[0064] According to an embodiment of this disclosure, a cooperative control system is also provided, comprising: a memory for non-temporarily storing computer-executable instructions; and a processor for running the computer-executable instructions, wherein the computer-executable instructions are executed by the processor to perform the cooperative control method described in any of the above embodiments.

[0065] According to embodiments of this disclosure, a non-transitory storage medium is also provided for storing computer-executable instructions non-transitory, wherein when the computer-executable instructions are executed by a computer, the cooperative control method described in any of the above embodiments is executed.

[0066] According to embodiments of this disclosure, a computer program product is also provided, including a computer program, characterized in that, when the computer program is executed by a processor, it implements the collaborative control method described in any of the above embodiments.

[0067] The following points also need to be explained:

[0068] (1) The accompanying drawings of the embodiments of this disclosure only involve the structures involved in the embodiments of this disclosure, and other structures can be referred to the general design.

[0069] (2) For clarity, the thickness of devices, layers, or regions is enlarged or reduced in the drawings used to describe embodiments of the present disclosure, i.e., these drawings are not drawn to scale. It will be understood that when an element such as a layer, film, region, or substrate is referred to as being “above” or “below” another element, the element may be “directly” located “above” or “below” the other element or there may be intermediate elements.

[0070] (3) Where there is no conflict, the embodiments of this disclosure and the features in the embodiments can be combined with each other to obtain new embodiments.

[0071] The above are merely specific embodiments of this disclosure, but the scope of protection of this disclosure is not limited thereto. The scope of protection of this disclosure shall be determined by the scope of the claims.

Claims

1. A collaborative control method, comprising: A preset visual instruction library is provided, which contains multiple visual instruction symbols. The visual instruction symbols are preprocessed to extract their image features; The image features are input into a visual semantic encoder for encoding to obtain an n-dimensional visual semantic vector. While the human body is visually focusing on one of the visual instruction symbols in the visual instruction library, the raw EEG signal of the human body is collected. The raw EEG signal is preprocessed to extract its EEG spatial features; The EEG spatial features are input into an EEG encoder for encoding to obtain an n-dimensional neural response vector. Map the n-dimensional visual semantic vector and the n-dimensional neural response vector to a shared semantic space; Based on the n-dimensional visual semantic vector and the n-dimensional neural response vector, comparative learning and target optimization are performed to obtain the matching visual-EEG pair vector; Based on the visual-EEG, the vector is classified into commands to obtain the command probability distribution. Based on the command probability distribution, control commands for controlling external devices are output.

2. The collaborative control method according to claim 1, characterized in that, The EEG spatial features include multiple EEG signal tensor dimensions, and the tensor dimensions include time length, number of channels, and signal features.

3. The cooperative control method according to claim 2, characterized in that, Obtaining the n-dimensional neural response vector includes the step of processing the EEG spatial features using a Transformer encoder.

4. The collaborative control method according to claim 1, characterized in that, The contrastive learning, based on a loss function, causes similar n-dimensional visual semantic vectors and n-dimensional neural response vectors to move closer together in the shared semantic space, while dissimilar n-dimensional visual semantic vectors and n-dimensional neural response vectors move further apart in the shared semantic space.

5. The cooperative control method according to claim 4, characterized in that, The loss function is defined as: L = -log[exp(sim(v i ,e i +) / τ) / ∑ j K exp(sim(v i ,e j -) / τ)],where, L is the loss function, v i Let e ​​be the response vector of the i-th neural network. i + represents the positive sample of the i-th visual semantic vector, e j - represents the j-th negative sample of the visual semantic vector, and K represents the number of negative samples of the visual semantic vector.

6. The cooperative control method according to claim 1, characterized in that, The contrastive learning, based on multiple pre-training of a deep learning model, can classify matching visual-EEG vector pairs.

7. A cooperative control device, comprising: Visual instruction library, acquisition module, and deep learning module; The visual instruction library contains multiple visual instruction symbols; The acquisition module is configured to acquire the raw EEG signal of the human body while the human visually fixates on one of the visual instruction signals in the visual instruction library. The deep learning module includes an image extraction layer, a visual semantic encoder, a spatial feature extraction layer, an EEG encoder, a cross-modal projection layer, a contrastive learning layer, and an output layer. The image extraction layer is configured to perform image preprocessing on the visual instruction symbols and extract the image features of the visual instruction symbols; The visual semantic encoder is configured to encode the image features to obtain an n-dimensional visual semantic vector; The spatial feature extraction layer is configured to perform data preprocessing on the original EEG signal and extract the EEG spatial features of the original EEG signal; The EEG encoder is configured to encode EEG spatial features to obtain an n-dimensional neural response vector; The cross-modal projection layer maps the n-dimensional visual semantic vector and the n-dimensional neural response vector to a shared semantic space; The contrastive learning layer performs contrastive learning and target optimization based on the n-dimensional visual semantic vector and the n-dimensional neural response vector to obtain a matching visual-EEG pair vector; The output layer performs instruction classification on the vector based on the visual-EEG, obtains the instruction probability distribution, and outputs control instructions based on the instruction probability distribution.

8. The cooperative control device according to claim 7, characterized in that, The EEG spatial features include multiple EEG signal tensor dimensions, and the tensor dimensions include time length, number of channels, and signal features; The visual semantic encoder and the EEG encoder are Transformer encoders; The contrastive learning, based on a loss function, causes similar n-dimensional visual semantic vectors and n-dimensional neural response vectors to move closer together in the shared semantic space, while dissimilar n-dimensional visual semantic vectors and n-dimensional neural response vectors move further apart in the shared semantic space. The loss function is defined as: L = -log[exp(sim(v i ,e i +) / τ) / ∑ j K exp(sim(v i ,e j -) / τ)],where L is the loss function, v i Let e ​​be the response vector of the i-th neural network. i + represents the positive sample of the i-th visual semantic vector, e j - represents the negative sample of the j-th visual semantic vector, K is the number of negative samples of the visual semantic vector, and τ is a hyperparameter; The contrastive learning layer is configured to classify matching visual-EEG pairs based on prior pre-training. It also includes external devices configured to perform control based on the logic-optimized control commands.

9. A cooperative control system, comprising: Memory is used to store non-temporary executable instructions for a computer. And a processor for running the computer-executable instructions, characterized in that the computer-executable instructions are executed by the processor at runtime according to any one of claims 1-6.

10. A non-transitory storage medium for non-transitory storage of computer-executable instructions, characterized in that, When the computer-executable instructions are executed by a computer, the collaborative control method according to any one of claims 1-6 is performed.

11. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by the processor, it implements the cooperative control method according to any one of claims 1-6.

Citation Information

Patent Citations

  • Visual question-answering method, device and equipment based on relation alignment, medium and product

    CN118211657A

  • Electroencephalogram signal decoding and visualization method and device based on pre-training diffusion model, equipment, medium and product

    CN120234713A

  • System and method for generating visual identity and category reconstruction from electroencephalography (EEG) signals

    US20190357797A1

  • Artificial Intelligence-Assisted Virtual Object Builder

    US20230260208A1