Method for imagined speech fragment identification and semantic reconstruction

The method addresses the challenges of BCI by using a controlled identification window and word-specific models within a database of predetermined commands to enhance accuracy and efficiency in speech recognition, ensuring robustness and reliability across diverse speech patterns and accents, making the technology more accessible and user-friendly.

WO2026013188A1PCT designated stage Publication Date: 2026-01-15MINDSPELLER BCI BV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
PCT/EP2025/069715
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-07-10
Filing Date
2025-07-10
Publication Date
2026-01-15

AI Technical Summary

Technical Problem

Existing brain-computer interface (BCI) technologies face challenges in accurately identifying intended speech from complex and noisy brainwave data, require substantial computational resources, and struggle with efficient processing and management of neural data, particularly for real-time applications, leading to inaccuracies and inefficiencies in recognizing and reconstructing compound commands.

Method used

A method involving a controlled identification window and word-specific identification models, coupled with a database of predetermined commands, optimizes the process by triggering data recording only upon speech input and deleting it after command identification, using word-specific models to enhance accuracy and efficiency, and incorporating dynamic adjustments for various speech patterns and contexts.

Benefits of technology

The method improves accuracy and efficiency in speech recognition from brain data, optimizes memory utilization, and enhances flexibility in recognizing complex commands, ensuring robustness and reliability across diverse speech patterns and accents, making the technology more accessible and user-friendly.

✦ Generated by Eureka AI based on patent content.
Patent Text Reader

Abstract

The current invention relates to an imagined, performed, mouthed or attempted speech fragment identification and semantic reconstruction. The method comprising the steps of: receiving, via a brain data capturing element, a brain data fragment containing speech; and searching for a command comprising at least one word, contained in said brain data fragment, said search being carried out based on a database comprising data related to at least one predetermined command, after the step of receiving a brain data fragment containing speech, an identification window having a starting-point, a duration and an end-point is selected in the brain data fragment, and within which identification window a word-specific word identification model is used in order to identify at least one word contained in at least one of the predetermined commands. In a second aspect, the invention relates to a system for controlling an electromechanical device.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] METHOD FOR IMAGINED SPEECH FRAGMENT IDENTIFICATION AND SEMANTIC RECONSTRUCTION

[0002] FIELD OF THE INVENTION

[0003] The present invention pertains to the technical field of human-machine interaction. More in particular, the invention pertains to the field of imagined speech recognition and brain-computer interfacing.

[0004] BACKGROUND

[0005] Brain-computer interfaces (BCIs) have emerged as a transformative technology, enabling direct communication between the human brain and computers. BCIs hold significant promise for a wide range of applications, including aiding individuals with speech or motor impairments, enhancing communication systems, and providing new modalities for human-computer interaction. A critical application of BCIs involves interpreting brainwave data to identify and reconstruct intended speech, thereby allowing users to communicate through thought alone.

[0006] Existing technologies in the realm of BCIs and neural signal processing face several notable challenges. One of the primary issues is the accurate identification of intended speech from complex and often noisy brainwave data. Brain signals corresponding to speech are intricate and can be affected by various external and internal factors, leading to potential inaccuracies in interpreting the intended commands.

[0007] Another significant challenge lies in the efficient processing and management of neural data. Continuous monitoring and analysis of brainwave signals require substantial computational resources and memory, particularly for real-time applications. This can result in considerable computational overhead and inefficiencies, hindering the practical deployment of such systems.

[0008] Furthermore, the complexity of recognizing and reconstructing compound or sequential commands from brainwave data adds an additional layer of difficulty. Accurate identification of successive syllables, words or phrases intended by the user demands sophisticated algorithms and models, which can be computationally intensive and prone to errors.

[0009] The present invention aims to resolve at least some of the problems and disadvantages mentioned above. SUMMARY OF THE INVENTION

[0010] The present invention and embodiments thereof serve to provide a solution to one or more of above-mentioned disadvantages. To this end, the present invention relates to a method according to claim 1. Preferred embodiments of the method are shown in any of the claims 2 to 14.

[0011] The present invention pertains to a computer-implemented method for speech (whether imagined, attempted, performed or mouthed) fragment identification and semantic reconstruction. The method involves receiving a fragment of brain data pertaining to speech and searching for a command within the fragment based on a database of predetermined commands. The method uses an identification window and a word-specific identification model to identify words within the commands. The model for each word includes a standard duration, and the database can contain models for successive words of a command. The duration of the identification window is carefully controlled, and the recording of the sound fragment (which may be virtual, for instance in the case of imagined speech, where the sound fragment is an intended sound fragment determined from brain activity) is triggered by receiving speech (or an equivalent fragment of brain data) and deleted once a command is identified. The method also includes detecting and counting the duration of each word in a brain data fragment, moving the identification window as needed, and using word identification models in a specific order. The invention offers several advantages, such as enhanced accuracy and efficiency in speech recognition from brain data, optimal utilization of memory resources, increased flexibility in recognizing complex commands, and improved overall efficiency.

[0012] A goal of the invention is to facilitate semantic reconstruction.

[0013] A goal of the invention is to optimize the process of command identification within fragments of brain data pertaining to speech.

[0014] A goal of the invention is to adapt to various speech patterns (or corresponding brain data representing speech or sounds) and command structures.

[0015] A goal of the invention is to ensure robustness and reliability in command identification.

[0016] A goal of the invention is to promote accessibility and inclusivity by providing a user- friendly interface for individuals with diverse speech patterns and accents, accommodating a wide range of users and enhancing usability in multicultural environments.

[0017] A goal of the invention is to improve the accuracy of identifying and reconstructing speech commands from brain data by using word-specific identification models and a carefully controlled identification window.

[0018] A goal of the invention is to optimize memory utilization by triggering the recording of brain data fragments only upon receiving speech and deleting these fragments once a command is identified.

[0019] A goal of the invention is to increase flexibility in recognizing complex commands by employing a database that contains models for successive words, allowing for more accurate context-based recognition.

[0020] A goal of the invention is to improve processing efficiency by detecting and counting the duration of each word by moving the identification window as needed to ensure timely and accurate word identification.

[0021] In a second aspect, the present invention relates to a system for controlling an electromechanical device according to claim 15. More particular, the system is directed at the execution of the method according to any of the claims 1 to 14.

[0022] DETAILED DESCRIPTION OF THE INVENTION

[0023] Unless otherwise defined, all terms used in disclosing the invention, including technical and scientific terms, have the meaning as commonly understood by one of ordinary skill in the art to which this invention belongs. By means of further guidance, term definitions are included to better appreciate the teaching of the present invention.

[0024] As used herein, the following terms have the following meanings:

[0025] "A", "an", and "the" as used herein refers to both singular and plural referents unless the context clearly dictates otherwise. By way of example, "a compartment" refers to one or more than one compartment.

[0026] "Comprise", "comprising", and "comprises" and "comprised of" as used herein are synonymous with "include", "including", "includes" or "contain", "containing", "contains" and are inclusive or open-ended terms that specifies the presence of what follows e.g. component and do not exclude or preclude the presence of additional, non-recited components, features, element, members, steps, known in the art or disclosed therein.

[0027] Furthermore, the terms first, second, third and the like in the description and in the claims, are used for distinguishing between similar elements and not necessarily for describing a sequential or chronological order, unless specified. It is to be understood that the terms so used are interchangeable under appropriate circumstances and that the embodiments of the invention described herein are capable of operation in other sequences than described or illustrated herein.

[0028] The recitation of numerical ranges by endpoints includes all numbers and fractions subsumed within that range, as well as the recited endpoints.

[0029] Whereas the terms "one or more" or "at least one", such as one or more or at least one member(s) of a group of members, is clear per se, by means of further exemplification, the term encompasses inter alia a reference to any one of said members, or to any two or more of said members, such as, e.g., any >3, >4, >5, >6 or >7 etc. of said members, and up to all said members.

[0030] Reference throughout this specification to "one embodiment" or "an embodiment" means that a particular feature, structure or characteristic described in connection with the embodiment is included in at least one embodiment of the present invention. Thus, appearances of the phrases "in one embodiment" or "in an embodiment" in various places throughout this specification are not necessarily all referring to the same embodiment but may. Furthermore, the particular features, structures or characteristics may be combined in any suitable manner, as would be apparent to a person skilled in the art from this disclosure, in one or more embodiments. Furthermore, while some embodiments described herein include some, but not other features included in other embodiments, combinations of features of different embodiments are meant to be within the scope of the invention, and form different embodiments, as would be understood by those in the art. For example, in the following claims, any of the claimed embodiments can be used in any combination.

[0031] The term "semantic reconstruction" refers in the present invention to the process of determining the intended meaning or semantic content from fragments of speech, which involves understanding both the context and the significance of the words a user imagines or attempts speaking. The term "brain data fragment comprising speech" or "fragment of brain data" should be understood as a fragment of "brainwave data", "brainwaves" or "brain data" containing speech-related neural activity. Brainwaves containing speech refer to the specific neural signals or electrical activity generated by the brain when a person imagines or attempts speaking, is thinking about or processing speech. These brainwaves or brain data can be detected and recorded using various neuroimaging techniques or brain data capturing devices, such as electroencephalography (EEG). The brainwaves contain patterns that correspond to the cognitive processes involved in speech production or recognition. Furthermore, the term "brainwave fragment containing speech" or "speech fragment" or "imaged speech fragment" or "mimed or mouthed speech fragment" are synonyms and in the context of this invention refers to a fragment of brain data or brainwaves comprising speech-related neural activity, which may comprise one or multiple words, that a system attempts to recognize and interpret.

[0032] The term "command" refers to a specific set of words that trigger a certain action in the computer implemented system.

[0033] The term "word" may refer to words, but also includes, unless explicitly indicated otherwise, syllables, phonemes as well as words.

[0034] The term "database” refers to a structured set of data. In this context, it contains data related to at least one predetermined command.

[0035] The term "identification window" should be understood as a selected part of the brain data fragment within which the system tries to identify words.

[0036] The term "word identification model" refers to a programmed model used to identify specific words or sequences of words. Such a model may be a general model that attempts to identify words from a speech / sound fragment (or brain data representing a speech fragment), and presenting for instance a list of scores for potentially detected words (presumably, with cutoff values). However, it may be a collection of individual word identification models that individually determine a likelihood of 'their' associated word being identified (or inversely, not being identified) within the speech / sound fragment, with the (one or more) models presenting the highest score(s) providing the end result, namely the most likely word(s). The term "standard duration" refers in the present invention to the typical length of time required to articulate a specific word, representing an average across varied speech patterns.

[0037] The terms "starting-point", "duration", and "end-point" of the identification window are the time at which the window starts, the length of the window, and the time at which the window ends, respectively.

[0038] The term "memory unit" refers to a part of the system where data can be stored and retrieved, such as a hard drive or RAM.

[0039] The term "pause" refers to a period of silence in the fragment of brain data pertaining to imagined, performed and attempted speech.

[0040] The term "processing unit" refers to the part of the system that carries out operations, such as a CPU.

[0041] The term "non-volatile memory" should be understood as a type of memory that can retain the stored information even when not powered.

[0042] The term "pre-process" refers in the present invention to performing a set of operations on data before it undergoes the main processing stage, such as noise reduction or amplification for speech, to improve the clarity of the input data.

[0043] The term "brain fragment capturing element" refers in the present invention to a device or component that captures brain data, for example brain waves.

[0044] The term "database" refers in the present invention to a structured collection of data that stores information related to predetermined commands, against which the system searches to match input speech fragments.

[0045] The term "fuzzy logic algorithms" refers in the present invention to a type of algorithm used to handle uncertain or imprecise information, allowing the system to make decisions based on ambiguous or vague input, such as speech with distortions or various accents.

[0046] The term "identification window" refers in the present invention to a specified segment of the audio signal identified for closer analysis, with a defined starting point, duration, and endpoint, focusing the system's analysis to identify words or commands. The term "word-specific word identification model" refers in the present invention to a computational model designed to recognize specific words within speech, tailored to the acoustic and linguistic characteristics of individual words.

[0047] The term "statistical distribution" refers in the present invention to a mathematical function describing all possible values and likelihoods a random variable can take within a range, used to model variations in word pronunciation durations among a population.

[0048] The term "machine learning model" refers in the present invention to a type of computational model that learns patterns from data, used to improve the system's ability to recognize and understand speech based on previous examples.

[0049] The term "dynamic updating mechanism" refers in the present invention to a feature allowing the system to update its database or models based on new data, helping it adapt to new commands or variations in speech over time.

[0050] The term "mouthed speech" or "mimed speech" refers to articulation without sound, and are used throughout this document as interchangeable. Both terms include each other, unless explicitly indicated otherwise.

[0051] The term "electromechanical device" refers in the present invention to a device combining electrical and mechanical processes and components, which can be controlled by commands through the described system.

[0052] In a first aspect, the invention relates to a computer implemented method for speech fragment identification (which includes identification of phonemes, syllables, words, etc.) and semantic reconstruction. The method comprises a step of receiving a brain data fragment comprising imagined, performed, mouthed (or mimed) and attempted speech. This step serves as the initial point of interaction with the system. By receiving brain data fragment comprising imagined, performed and attempted speech, the system obtains the input necessary for processing and interpreting user commands. Optionally, the system could pre-process the brain data fragment to enhance clarity, employing noise reduction or amplification techniques, which would improve recognition accuracy in varied auditory environments. The "input" or "command" can be understood as the imagined, performed, mouthed (or mimed) or attempted speech of the user.

[0053] Preferably, the brain data fragment comprising speech data is received via a brain fragment capturing element. The brain fragment capturing element acts as the primary data source, capturing input from users. This input serves as the raw material for speech recognition and semantic reconstruction processes. Brain fragment capturing elements are devices or components capable of capturing brain data. Various brain fragment capturing elements can be used for the purpose of this invention. For example, but not limited to, brain fragment capturing elements include electroencephalography (EEG) - or even stereoelectroencephalography (sEEG), Magnetoencephalography (MEG), Functional Near-Infrared Spectroscopy (fNIRS), Electrocorticography (ECoG), Local Field Potential (LFP), Positron Emission Tomography (PET), Functional Magnetic Resonance Imaging (fMRI), Single Photon Emission Computed Tomography (SPECT), functional ultrasound (fUS), optically- pumped magnetometer MEG (OPM-MEG). Furthermore the brain data capturing device can also include surface and dry / wet electrodes (in case of an EEG), amplifiers, analog-to-digital converters (ADC), signal processing units, data acquisition systems, wireless transmitters, headsets or caps, ground and reference electrodes, software platforms or combinations thereof.

[0054] Preferably, the method comprises the step of searching for a command comprising at least one word, phoneme or syllable contained in said brain data fragment, said search being carried out in a database comprising data related to at least one predetermined command. To increase adaptability, the system could employ a dynamic updating mechanism for the database, allowing it to learn and adapt to new commands over time based on user interactions. By searching within a database of predetermined commands, words or phrases can be matched to predefined commands, facilitating accurate interpretation of intentions of a user. Searching for commands in a predefined database allows for understanding the semantic meaning of the speech fragment. By associating words with specific commands, it can infer the user's intended actions or requests, leading to more meaningful interactions. Searching for commands within a database of predetermined commands improves the efficiency and performance of the speech recognition process. Furthermore, instead of analyzing the entire speech fragment for every possible command, the focus could be on a predefined set of commands, reducing computational complexity and processing time. The database can be tailored to include commands relevant to the intended functionality, ensuring accurate recognition and response to user input. As it evolves and new commands are introduced, the database can be updated and expanded to accommodate additional commands. This scalability ensures that it can adapt to changing user needs and support future enhancements without requiring significant modifications to the underlying speech recognition algorithms. This process is crucial for translating brain data relating to imagined, performed, mouthed or attempted speech into actions that can be executed or words and / or phrases by devices such as computers, communication aids, or robots. Using a database of predetermined commands helps quickly and accurately identify intended actions because the database serves as a reference point, enabling the matching of input speech against a curated list of commands it is programmed to recognize. This specificity significantly reduces the scope of search and comparison, enhancing both the speed and accuracy of command identification. Having a database allows for scalability. Predetermined commands can be easily updated, modified, or expanded without altering the fundamental operation. Furthermore, this approach directly impacts the user experience by ensuring that commands are recognized and acted upon swiftly and accurately. It minimizes errors in command interpretation, making the technology more reliable and user-friendly. Users can interact using (brain data corresponding to) natural language, making the technology more accessible and convenient. By structuring the command search around a database, there's room for implementing algorithms that understand context or variations in how commands are phrased. This allows for recognizing commands not just by direct matches, but also by understanding variations in language use, accents, or even colloquialisms, as long as they're represented within the database. Additionally, the database could utilize fuzzy logic algorithms to enhance the system's ability to recognize commands with varied accents, dialects, or colloquial expressions, broadening the system's accessibility.

[0055] Preferably, after the step of receiving a brain data fragment comprising imagined, performed, mouthed or attempted speech, an identification window having a starting-point, a duration and an end-point is selected. Moreover, the identification window selection process could be optimized using machine learning algorithms to automatically adjust its parameters based on the characteristics of the incoming brain data related to speech, such as speed and pauses, for improved word identification accuracy. By selecting an identification window within the received brain data fragment, speech is effectively segmented into smaller, manageable segments. This segmentation facilitates the analysis of individual words or phrases within the larger speech fragment, improving the accuracy of word identification. The identification window allows focusing attention on specific portions of the brain data fragment containing speech where relevant commands or words are likely to be found. By narrowing down the search space, efficiency in allocating computational resources and reducing processing time can be achieved. The starting-point, duration, and end-point of the identification window can be dynamically adjusted based on factors such as the characteristics of the speech fragment and the performance of the word identification model. This flexibility enables adaptation of the search strategy to various speech patterns and environmental conditions, enhancing overall robustness and reliability. The method could also incorporate real-time feedback mechanisms to dynamically adjust the identification window in response to user corrections, further enhancing recognition accuracy through interactive learning.

[0056] Preferably, within the identification window a word-specific word identification model is used to identify at least one word contained in the at least one of the predetermined commands. This approach enables leveraging specialized models tailored to the characteristics of each word, thereby enhancing the accuracy of word recognition and reducing the risk of false positives or misinterpretations. The identification window serves as the context within which the search for commands within the predetermined database is conducted. By focusing the search within this window, it ensures that only relevant portions of the speech fragment are considered during command identification, thereby further enhancing efficiency and accuracy. Thus, the receiving of the brain data fragment triggers a step of selecting an identification window. At least one routine or script is stored in a device memory, which routine or script is directed at selecting said identification window. A final substep in the routine or script for selecting an identification window, triggers the selection of said at least one word-specific word identification model stored in memory. In addition to word-specific models, context-aware algorithms that consider the syntactic and semantic context surrounding the identified words could be incorporated, improving the accuracy of command identification by leveraging surrounding linguistic cues.

[0057] According to a further or alternative embodiment, the model of each word comprises a standard duration of its associated word, said model being configured to detect its corresponding word by comparing the duration of the identification window with said duration of said word. Within the same language, brainwaves or brain data containing imagined, performed, mouthed or attempted speech or imagined, performed, mouthed or attempted speech patterns can vary significantly from one person to another, these speech patterns include the pronunciation of sentences, words, or even individual letters, phonemes, syllables, etc. Such manifestations of these alterations are known from regional accents, differences in physiognomy among individuals, speech impediments, or other. In order to overcome such differences in pronunciation, a word identification model comprises a standard duration for the word related to said model, which duration is an estimation based an average pronunciation time of said word among individuals of a predetermined population. The model further includes a statistical distribution related to said average, preferably said distribution is a normal distribution. It follows that a word identification model is configured to calculate a score, which score reflects the degree of certainty the part of the brain data fragment within an identification window corresponds to the word associated to said model. By assigning a standard duration to each word within the word identification model, consistency in the detection process is ensured. This consistency helps maintain reliability and accuracy in identifying words across different speech fragments and variations in pronunciation or enunciation. The standard duration of each word serves as a reference point for comparing against the duration of the identification window. By comparing the duration of the identification window with the standard duration of the word, the ability to assess the likelihood of a match and determine the presence of the word within the speech fragment is enhanced. The comparison between the duration of the identification window and the standard duration of the word establishes a threshold for word detection. In some embodiments, a (dynamic) time warping operation can be performed during speech decoding, to counteract potential influence on actual length due to emphasis or accent and ensure that all variations result in a same length. If the duration of the identification window closely matches the standard duration of the word, it increases the confidence level that the word is present in the speech fragment, leading to more accurate detection and interpretation. While the word identification model may include a standard duration for each word, it can also account for variations in speech tempo, rhythm, and pronunciation. The model can dynamically adjust its detection criteria based on contextual factors, allowing for robust performance across different speaking styles and environments. By leveraging the standard duration of each word for detection, the optimization of resource utilization and processing efficiency is achieved. Rather than exhaustively analyzing the entire speech fragment, the focus can be on specific time intervals within the identification window that align with the expected durations of individual words, reducing computational overhead. To accommodate users with speech variations, the standard duration could be adjusted based on user-specific speech profiles, creating a personalized recognition experience that caters to the individual's speech patterns.

[0058] According to a further or alternative embodiment, said database further comprises word identification models corresponding to at least two successive words which together form a command, expression or sentence. Preferably for each command, said database comprises any 1 to n-number of successive words of a command, with n being the number of words of a command. In this way, the database comprises all possible word identification models with which to identify any of the possible commands, even if an identification window only contains part of one command. By including word identification models for successive words of a command, a deeper understanding of the context in which individual words are used is gained. This can be supported even further by knowledge of the more general context in which the command is used. This contextual information improves the accuracy of command interpretation by considering the relationship (syntax, amongst others) between consecutive words and the overall structure of the command. Having word identification models for multiple successive words allows for the recognition of complete commands more effectively. By considering the sequential arrangement of words within a command, it is possible to identify and interpret complex command structures with greater accuracy and reliability. Including word identification models for all possible combinations of successive words within a command achieves comprehensive coverage of potential commands. This comprehensive coverage ensures that accurate identification and interpretation of any command can occur, even if the identification window only contains part of a command. Incorporating word identification models for successive words enables handling cases where the identification window contains only part of a command. Additionally, the database could implement predictive modeling techniques to anticipate the next word in a command sequence based on the identification of preceding words, enhancing the fluidity and naturalness of interactions.

[0059] According to a further or alternative embodiment, the maximum duration between the starting-point and the end-point of the identification window does not exceed the duration of the word identification model with the longest duration present in said database. In this case, the longest duration of an identification window will be the duration of the longest command in the database. This permits reducing the computational complexity associated with the method, by preventing a continuous and excessive expansion of the identification window beyond of diminished or even no returns. Furthermore, limiting the duration of the identification window to the longest word model ensures that excessive time is not spent analyzing segments of audio that are unlikely to contain meaningful speech. This optimization prevents the waste of computational resources on overly long fragments where no commands are likely to be found, thereby speeding up the processing time. This constraint also helps in maintaining a high level of accuracy in word identification. By ensuring that the identification window is no longer than necessary, the risk of including speech that could confuse the identification process is minimized. It focuses the analysis of fragments that are just the right length to contain potential commands, reducing the likelihood of false positives or negatives. By aligning the identification window's duration with the durations of word models, potentially applying (dynamic) time warping, the ability to more accurately adapt to natural speech patterns is enhanced. This is especially important for recognizing commands within continuous speech, where pauses, speed, and emphasis can vary widely between users and contexts. The approach ensures that flexibility and responsiveness to these variations are maintained. For real-time or near-real-time applications, such as brainwave or brain data -activated controls or interactive brain data response systems, maintaining a maximum duration for the identification window ensures prompt responsiveness to user commands. This responsiveness is crucial for user satisfaction and usability, making the technology more practical and appealing for everyday use.

[0060] According to a further or alternative embodiment, the minimum duration between the starting-point and the end-point of the identification window is at least the same as the duration of the word detection model of the word with the smallest phoneme duration. By establishing a minimum duration that matches the shortest word model, the system and / or method guarantees that even the briefest commands or words can be captured and analyzed effectively. This ensures that no potential command is too short to be detected, addressing a common challenge in speech recognition where very brief fragments might otherwise be overlooked or dismissed as noise. This choice directly contributes to the accuracy of the speech recognition process. It ensures that enough time is allocated to accurately identify short words, which are often critical for understanding commands or queries. By not cutting off the analysis too early, the risk of misinterpretation or missing commands altogether is reduced, brain data comprising speech patterns vary significantly among individuals, including the duration of words and pauses. Setting a minimum duration based on the shortest word model accommodates these variations, ensuring flexibility enough to handle different speaking rates and styles. This adaptability is crucial for intended use by a broad user base. While it's important to not excessively analyze long brain data fragments, it's equally important to ensure that the analysis window is not so short that it misses potential commands. By setting a sensible minimum duration, the optimization of computational resources is achieved, focusing on segments of audio that are just long enough to contain meaningful speech without wasting resources on fragments that are too short to be useful. According to a further or alternative embodiment, the identification window is larger than the time necessary to pronounce and conjunction words. In this way, the identification window is kept advantageously too large to detect only conjunction words, which conjunction words do not provide any substantial contribution to the identification of a command. The detection window is kept large enough for the identification of words with the highest relevance for the identification of a command.

[0061] According to a further or alternative embodiment, the step of receiving a brain data fragment containing speech, triggers the recording of said brain data fragment in a memory unit, said recording being deleted from said memory unit when a command is identified. This ensures the efficient use of the memory unit by only temporarily storing brain data fragments until they are processed, and a command is identified. By deleting the recording after command identification, memory resources are conservatively managed, preventing unnecessary accumulation of data and maintaining optimal performance over time. Automatically deleting recordings after processing addresses potential privacy concerns. It ensures that brain data of the user is not permanently stored without necessity, aligning operations with privacy best practices and possible regulatory requirements regarding data retention. Triggering recording upon receiving speech and deleting it post-command identification minimizes processing overhead. It focuses computational resources on currently relevant data, enhancing responsiveness and reducing the time between speech input and action execution. This approach allows for focusing on identifying and executing commands effectively. Once a command is recognized and presumably acted upon, the associated brain data fragment no longer serves a purpose within the context of command execution, thus its deletion helps maintain a lean processing pipeline. Deleting recordings after command identification also readies for new inputs without the risk of confusion with previously stored data. This ensures continuous processing of incoming speech commands without degradation in performance due to old data accumulation.

[0062] According to a further or alternative embodiment, once a command is identified, a further step of repeating the command to the user is carried out in order to receive a confirmation that the identified command is the command the user intended to communicate. Once confirmed, a further step of carrying out an action associated with said command is triggered, the conclusion of each action triggers the deletion of said identified command from at least a volatile memory. Preferably, said identified and executed command is moved from said volatile memory, to which command a time stamp is then associated. In this way, the volatile / working memory of the device executing the method is advantageously cleared while a historical record of the identified and executed commands is maintained, by preference, for each user.

[0063] According to a further or alternative embodiment, said brain data fragment recorded in said memory unit is deleted from said memory unit after failing to detect a number of words exceeding the number of words of the predetermined command having the largest number of words. This significantly contributes to efficient use of memory, ensuring that resources are not wasted on storing and processing irrelevant audio data. It optimizes performance, allowing for quicker, more accurate responses to user commands by focusing on potentially relevant data. Moreover, by dynamically adjusting to the complexity of commands and promptly removing unnecessary data, this method also addresses user privacy and experience considerations. This approach ensures that the process remains both efficient and responsive, maintaining a balance between performance and user-centric design.

[0064] According to a further or alternative embodiment, a new brain data fragment starts to be recorded and saved to said memory unit after the previous brain data fragment is deleted. This ensures a continuous operation and responsiveness of the voice recognition functionality without unnecessary delays or interruptions. By immediately starting to record a new brain data fragment after the previous one is deleted, the invention ensures that no potential voice commands are missed, maintaining a seamless interaction experience for the user. This approach reflects a proactive system design that prioritizes operational efficiency and user engagement. By keeping the recording process in a ready state, the invention can capture and process voice commands in real-time, facilitating a dynamic and interactive environment. This continuous recording cycle is crucial for applications requiring constant vigilance and immediate response to voice inputs, embodying a design that is both user-centric and performance-oriented.

[0065] According to a further or alternative embodiment, a threshold for deletion of the brain data fragment containing speech is chosen. Preferably, a threshold for the deletion is based on the detection of failure relative to the longest command, this avoids unnecessary retention of audio (brain) data that does not result in recognized commands. This helps in preventing memory overflow and maintains processing speed by not clogging the memory with unprocessable data. Limiting data retention to the context of successful command recognition focuses the processing efforts on potentially actionable brain data fragments. Once it's determined that a fragment will not yield command recognition, for example, because the number of attempted detections exceeds the length of any known command, it's possible to discard this data to make room for new inputs. This ensures that processing power is allocated to new, potentially recognizable commands rather than being wasted on fragments that have already been deemed irrelevant. In instances where speech is misinterpreted or background noise triggers a recording, this mechanism serves as an error-handling procedure, allowing for recovery and preparation for new input. By clearing the memory of these unproductive brain data fragment s after a certain threshold of detection failure, it's possible to reset its state, reducing the chance of error accumulation and lag. Automatically deleting unproductive fragments enhances user privacy by ensuring that only relevant data, i.e., that leads to a recognized command, is potentially retained for any length of time.

[0066] According to a further or alternative embodiment, the method further comprises a step of detecting and counting the duration of each word in a brain data fragment, the duration of each word being counted by counting the elapsed time between each consecutive pause having a duration of at least 0.2s, preferably at least 0.3s, more preferably at least 0.4s, more preferably at least 0.5s, more preferably at least Is, more preferably at least 1.2s, more preferably at least 1.4s, more preferably at least 1.6s, more preferably at least 1.8s, more preferably at least 2s, more preferably at least 3s, more preferably at least 4s, more preferably at least 5, even more preferably at least 10s. This approach allows for more precise segmentation of speech into individual words or commands. By identifying pauses of a specific minimum duration, it is possible to accurately distinguish between separate words. Another option, which counters issues with parsing when insufficient pause is present between distinct words, is to have roaming, gliding word identification windows, which move over the brain data in order to parse. Another option to address this is to focus on detection of phonemes and / or syllables. This segmentation is crucial for processing complex commands or sentences where the distinction between words is essential for understanding the intended command. Setting this optimal threshold for pauses helps differentiate between meaningful pauses in speech and incidental background noise or brief hesitations. This is key in environments with variable noise levels or when the user's imagined, performed or attempted speech pattern includes natural hesitations, ensuring consistent performance across different speaking styles and ambient conditions. The ability to measure word durations (independently from the identification of the word) contributes to understanding natural speech patterns, which can be used to improve recognition algorithms, making them more adaptable to the natural variances in speech speed, rhythm, and cadence across different users. It can also aid in distinguishing between words that sound similar but have different lengths. Detecting the duration of words provides additional temporal context that can be invaluable for semantic analysis. Understanding the timing between words can offer clues about the speech's structure and meaning, enabling more sophisticated interpretations of user commands, especially in languages where timing and pauses significantly impact meaning. By efficiently segmenting speech into discrete words based on pause duration, it is possible to optimize processing resources, focusing computational effort on segments of audio likely to contain meaningful information, and reducing the processing load by ignoring segments that don't meet the criteria for meaningful pauses.

[0067] According to a further or alternative embodiment, the end-point of the identification window is moved to the end of each new detected word. Adjusting the end-point of the identification window to the end of each newly detected word optimizes responsiveness and accuracy in processing commands, ensuring continuous alignment with the flow of speech. This dynamic adjustment enhances the ability to adapt in real-time to variations in speech length and pacing, focusing analysis on the most relevant audio segments for command identification. It leads to quicker recognition of commands, reducing the lag between speech input and response, critical for applications requiring immediate action based on commands. Furthermore, this approach facilitates more natural interactions, mirroring how listeners focus on recent speech parts in conversation, and accounts for dynamic changes in speech, such as pauses, speed, and emphasis, ensuring agility and capability to adjust to diverse speaking styles and environments.

[0068] According to a further or alternative embodiment, if the duration of the identification window reaches its maximum, the starting point of the identification window is moved to the beginning of the next word nearest to the current position of said starting-point. Shifting the starting point of the identification window to the beginning of the next word when the identification window reaches its maximum duration ensures continuous and efficient speech processing. This prevents the window from becoming stagnant or overly focused on a segment that has already been analyzed or is too long, which could hinder the recognition of subsequent words or commands. By dynamically adjusting the starting point in response to reaching the window's duration limit, the process stays aligned with the ongoing flow of speech, maintaining the relevance of the analysis. This adjustment allows for the efficient handling of long speech segments and ensures that the platform remains responsive to new input, facilitating accurate and timely command identification without unnecessary delays or processing of irrelevant data.

[0069] According to a further or alternative embodiment, the identification window is moved to the next word in the speech fragment after successful identification of the current word. Moving the identification window to the next word in the speech fragment after successfully identifying the current word ensures that the processing flow remains efficient and focused on new, unprocessed speech data. This approach allows for sequential processing of speech, where each word is analyzed in turn, enhancing the ability to continuously interpret and respond to commands without re-analyzing previously identified portions of speech. It aligns the processing with the linear nature of speech, ensuring that as soon as a word is recognized, the platform immediately shifts its attention to the next potential command or word, thereby maintaining an effective and streamlined recognition process. This method optimizes the responsiveness and accuracy, enabling it to keep pace with the user's speech and facilitating a more natural and intuitive interaction.

[0070] According to a further or alternative embodiment, the identification window is moved to the next word in the speech fragment after failing identification of the current word. This ensures that the method and / or system does not become bogged down by attempting to re-analyze a segment of speech that has already proven difficult to interpret. Instead, it progresses to the next segment in the hope of more successful identification, thereby preventing wasted computational resources on likely unrecognizable inputs. This method significantly enhances a system's overall throughput and responsiveness by ensuring continuous progression through the speech input, even in the face of challenges with certain words or phrases. It reflects an understanding that perfect recognition of every word is not always possible or necessary for effective command identification or for understanding the intended meaning of the user's speech.

[0071] According to a further or alternative embodiment, word identification models are used following an order going from the shortest (word identification) model duration to longest (word identification) model duration. Prioritizing shorter models accelerates the identification process for common, shorter words, which are frequently used in speech. This ensures swift parsing and recognition of these words, expediting the overall comprehension of the user's commands or queries. By giving precedence to shorter models, the platform reduces the chance of incorrectly identifying longer phrases or words when a shorter word is commanded. This approach maintains high accuracy levels by systematically adjusting the complexity of the models based on the word's length being identified. Additionally, it enables incremental processing of speech, allowing the platform to progressively comprehend more complex phrases and sentences. This approach signifies a layered strategy to speech recognition, dynamically adapting the processing strategy according to the speech input's complexity. Furthermore, shorter words often constitute fundamental language elements (such as prepositions, articles, and conjunctions) crucial for grammatical structure and meaning. Prioritizing their identification assists in structuring subsequent analysis of longer words or phrases, enhancing the recognition process's logical and contextual awareness.

[0072] According to a further or alternative embodiment, the method could include a feedback loop where the system provides users with the option to correct or confirm the identification of commands. This user feedback could be used to refine and update the word identification models and the database of commands, ensuring continuous improvement of the system's accuracy and adaptability to user preferences.

[0073] In a most preferred embodiment, the invention relates to a computer implemented method for speech fragment identification and semantic reconstruction, the method comprising the steps of: receiving, via a brain fragment capturing element, a brain data fragment containing speech; and searching for a command comprising at least one word, contained in said brain data fragment, said search being carried out based on a database comprising data related to at least one predetermined command, wherein after the step of receiving a brain data fragment containing speech, an identification window having a starting-point, a duration and an end-point is selected in the brain data fragment, and within which identification window a word-specific word identification model is used in order to identify at least one word contained in at least one of the predetermined commands. By utilizing word-specific word identification models and fine-tuning the identification window based on word durations, the present method achieves higher accuracy in recognizing speech fragments compared to traditional methods. This increased accuracy can lead to more reliable interpretation of user commands, enhancing the overall user experience. The use of predetermined commands and word identification models allows for efficient searching and identification of commands within speech fragments. This efficiency can lead to faster response times and smoother operation of the platform, improving user satisfaction and productivity. The recording and deleting of brain data fragments from memory, as well as its management of identification windows, optimizes memory usage by only storing relevant data and minimizing unnecessary storage overhead. This efficient memory management can lead to cost savings and improved performance. Furthermore, detecting and counting the duration of each word, adjusting the identification window dynamically, and using word identification models in a specific order based on model durations demonstrates flexibility and adaptability to various speech patterns and command structures. This adaptability enhances the versatility of the platform across different user scenarios and environments.

[0074] In a second aspect, the invention relates to a system for controlling an electromechanical device. The system could also be adapted to interface with multiple devices simultaneously, allowing users to control various aspects of their environment through a single, centralized command system.

[0075] Preferably, the system has a brain data capturing element for capturing a brain data fragment comprising speech from a user. Brain data commands provide a handsfree, intuitive way for users to interact with the system, making it accessible to a wider range of users, including those with physical disabilities or situations where manual control is inconvenient or unsafe (e.g., driving). Incorporating brain data control expands accessibility, allowing users to operate the device without needing to physically touch it or navigate through complex menus. This can be particularly beneficial in environments where users' hands are occupied or when users are not in close proximity to the device. Additionally, a user who can no longer speak can still communicate through a device, such as a robot, using their brain data or brain waves to express their own language and accent. Furthermore, commands derived from brain data can often accomplish tasks more quickly than manual controls, especially for complex or multi-step operations. This can significantly enhance the efficiency and overall user experience of the system. Additionally, it can enhance the efficiency with which the user can handle the system, as they improve during (long-term) use thereof. This can be further supplemented by providing online feedback via (auditory and / or other) signals.

[0076] Preferably, the system has a non-volatile memory element for storing at least one word identification model, at least one predetermined command and a number of predetermined responses. Non-volatile memory retains information even when the system is powered off or reset. This ensures that critical data, such as word identification models and commands, are permanently stored and readily available upon system startup, eliminating the need for reconfiguration or relearning. Storing information in non-volatile memory allows the system to quickly access and process commands without significant delays. It enables efficient retrieval of word identification models and responses, leading to faster processing of user commands and a smoother user experience. The use of non-volatile memory furthermore enhances the reliability of the system. Since the essential components for voice recognition and response are securely stored, the system can consistently recognize and respond to user commands, even after periods of inactivity or through power cycles. A non-volatile memory also facilitates the easy updating and expansion of word identification models, commands, and responses. As the system's capabilities grow or as updates are required to improve performance or add functionalities, the stored data can be updated without hardware modifications.

[0077] Preferably, the system has a processing unit in communication with the brain data capturing element, and preferably said non-volatile memory. The processing unit is configured to execute said at least one word identification model. The processing unit acts as the central command center of the system, coordinating the flow of data between the data capturing element and the memory. This setup ensures that the brain data input is efficiently processed and matched against stored word identification models to execute commands accurately. Equipping the system with a dedicated processing unit enables real-time or near-real-time processing of commands. This is essential for a responsive user experience, where delays between command issuance and system response are minimized. The processing unit is responsible for executing sophisticated word identification models that may involve complex algorithms, including those based on artificial intelligence and machine learning. These models require substantial computational resources and specialized processing capabilities to analyze speech patterns accurately.

[0078] Preferably, the system has a sound emitting element configured to communicate with the user. The sound emitting element is in communication with said processing unit. The sound emitting element provides users with immediate auditory feedback in response to their brain data commands. This feedback can confirm that a command has been received and processed, inform the user of the outcome of their command, or notify them of errors or the need for additional input, thereby facilitating a two-way interaction that mirrors natural human communication patterns. Furthermore, auditory feedback typically generates brain signals similar to those of performed speech, resulting in the user 'hearing' their own speech. Additionally, such a feedback is extremely relevant to allow communication based on attempted or imagined speech. Auditory responses can make the system more intuitive and user-friendly. By receiving vocal confirmations, prompts, or instructions, users can navigate the system's features more easily and make corrections or adjustments to their commands as needed, enhancing the overall usability of the system. For users with visual impairments or in situations where visual feedback is impractical (e.g., when the user's attention is focused elsewhere), audio feedback ensures that the system remains accessible to all users, broadening its applicability and inclusiveness. The sound emitting element can provide essential operational confirmations and alerts, informing users about the system's status, completed actions, or warnings about errors or malfunctions. This capability is crucial for ensuring the safety and reliability of the system, particularly in applications where timely notifications are critical. Integrating a sound emitting element allows for a range of auditory responses, from simple beeps or chimes to complex verbal instructions or confirmations. This flexibility supports customization to suit different applications, user preferences, or branding requirements, enhancing the system's appeal and effectiveness.

[0079] The non-volatile memory further comprises instructions to carry out the method according to any of the embodiments of the previous aspect. It is evident to a person skilled in the art that the non-volatile memory further comprises instructions to execute the method as described in any of the embodiments outlined in the previous aspect. It is understood by those skilled in the art that the advantages highlighted for the method equally apply to the system implementing said method.

[0080] Additionally, the system could incorporate a feedback mechanism where it verbally communicates with the user to clarify or confirm commands, thus reducing errors and enhancing user confidence in the system's ability to understand and execute commands accurately.

[0081] The present invention will be now described in more details, referring to examples that are not limitative.

[0082] EXAMPLES

[0083] The present invention will now be further exemplified with reference to the following examples. The present invention is in no way limited to the given examples.

[0084] Example 1. A user is interacting with a home automation system using brainwave commands. The user imagines or thinks about the command 'Turn on the living room lights'. The system receives this brainwave fragment and begins the identification process. Initially, it uses a model for single-word recognition. However, due to the complexity of the command, the single-word model fails to recognize the command accurately. The system then switches to a model capable of recognizing two-word phrases, and then to a three-word model, and so on, until the entire command is recognized accurately. This demonstrates the adaptive learning and flexible command recognition capabilities of the system.

[0085] Example 2. In a different scenario, the user tries to pronounce the command 'Play Mozart'. The system receives this brain data fragment and uses a single-word model to identify the command. The command is recognized accurately, and the system executes it by playing Mozart's music. Once the command is executed, the recorded brain data fragment is deleted from the memory unit, demonstrating efficient memory usage.

[0086] Example 3. Consider a situation where a user imagines a complex command in French. The system receives the brain data fragment and begins the identification process. Since the command is complex and in a different language, the initial models fail to recognize it accurately. The system then switches to models capable of recognizing larger word lengths, thereby improving the overall recognition, and understanding performance, even in complex language environments.

[0087] Example 4. In a scenario where a user has lost the ability to speak, they can still communicate by imagined, performed or attempted speaking the words they want to say. The system captures their brain data and processes them to generate speech through a device, such as a robot. Remarkably, the generated speech retains the user's unique language and accent, allowing them to continue speaking as they did before losing their voice. This seamless integration ensures that their personal way of speaking is preserved, providing a familiar and comfortable means of communication.

[0088] Example 5. Consider a situation where a user gives a command, but the system fails to recognize it due to noise, such as from the EEG. The system then moves the identification window to the next word in the speech fragment after failing identification of the current word, thereby improving the resilience of the speech recognition system.

[0089] Example 6. Consider a system that comprises:

[0090] - A brain data capturing element for receiving commands from a user.

[0091] - A non-volatile memory element that stores word identification models for individual words and sequences of words, predetermined commands, and a number of predetermined responses. - A processing unit in communication with the brain data capturing element and the non-volatile memory, configured to execute the word identification models.

[0092] - A sound emitting element for communicating responses or feedback to the user, in connection with the processing unit.

[0093] Upon a user's command, whether by thinking, imagining, mouthing or attempting to pronounce a command, the brain data capturing element receives a brain data fragment containing speech. This fragment is temporarily stored in a memory unit. An identification window is then selected within the brain data fragment, determined by a starting point, a duration, and an endpoint. This window is designed to encapsulate the potential location of at least one word contained in the predetermined commands. Utilizing a word-specific identification model, the system processes the brain data fragment within the identification window to identify any word or sequence of words matching those within the database of predetermined commands.

[0094] The model of each word includes a standard duration, enabling the system to detect words by comparing the duration of the identification window against the model's duration. This comparison helps in accurately identifying words and reducing false positives. Based on the outcome of the word identification step, the system dynamically adjusts the identification window.

[0095] If a word is identified, the endpoint of the identification window is moved to the end of the newly detected word.

[0096] If the identification window's duration reaches its maximum without successful word identification, the starting point shifts to the beginning of the next word nearest to the current starting point.

[0097] Once a command is identified within the brain data fragment, the processing unit triggers a predetermined response, which is communicated to the user via the sound emitting element. If a command is not detected within the brain data fragment, or if the number of unsuccessfully detected words exceeds the largest number of words in the predetermined commands, the brain data fragment is deleted from memory, and the system is ready to receive a new brain data fragment.

[0098] Consider a scenario where a user imagines the command "turn on the light". The brain data capturing element captures this command as a brain data fragment, which is then processed following the steps described. The system efficiently identifies the command through the specified identification window and word identification models, despite potential noise or variations in speech patterns. Consequently, the command "turn on the light" is recognized, and the corresponding action is either spoken aloud by a device, such as a computer, or executed directly, with a verbal confirmation provided to the user.

[0099] It is supposed that the present invention is not restricted to any form of realization described previously and that some modifications can be added to the presented example of fabrication without reappraisal of the appended claims. It is clear that the invention can be applied to various systems such as home automation systems, car voice control systems, voice-controlled television systems, or even voice- controlled robotic systems for instance. The present invention is in no way limited to the embodiments described in the examples. On the contrary, methods according to the present invention may be realized in many different ways without departing from the scope of the invention.

Claims

CLAIMS1. A computer implemented method for imagined, performed, mouthed or attempted speech identification and semantic reconstruction, the method comprising the steps of:- receiving, via a brain fragment capturing element, a fragment of brain data comprising imagined, performed, mouthed or attempted speech; and- searching for a command comprising at least one word, contained in said brain data fragment, said search being carried out based on a database comprising data related to at least one predetermined command, wherein, after the step of receiving a brain data fragment comprising imagined, performed, mouthed or attempted speech, an identification window having a starting-point, a duration and an end-point is selected in the brain data fragment, and within which identification window a wordspecific word identification model is used in order to identify at least one word contained in at least one of the predetermined commands.

2. The method according to any of the previous claims, characterized in that, the word identification model of each word comprises a standard duration of its associated word, said word identification model being configured to detect its corresponding word by comparing the duration of the identification window with said duration of said word.

3. The method according to any of the previous claims, characterized in that, said database further comprises word identification models corresponding to at least two successive words of a command.

4. The method according to any of the previous claims, characterized in that, the maximum duration between the starting-point and the end-point of the identification window does not exceed the duration of the word identification model with the longest duration present in said database.

5. The method according to any of the previous claims, characterized in that, the minimum duration between the starting-point and the end-point of the identification window is at least the same as the duration of the word detection model of the word with the smallest sound duration.

6. The method according to any of the previous claims, characterized in that, the step of receiving a brain data fragment triggers the recording of saidbrain data fragment in a memory unit, said recording being deleted from said memory unit when a command is identified.

7. The method according to claim 6, characterized in that, said brain data fragment recorded in said memory unit is deleted from said memory unit after failing to detect a number of words exceeding the number of words of the predetermined command having the largest number of words.

8. The method according to any of the previous claims 6-7, characterized in that, a new brain data fragment starts to be recorded and saved to said memory unit after the previous brain data fragment is deleted.

9. The method according to any of the previous claims 3-8, characterized in that, the method further comprises a step of detecting and counting the duration of each word in a brain data fragment, the duration of each word being counted by counting the elapsed time between each consecutive pause having a duration of at least 0.2s.

10. The method according to claim 9, characterized in that, the end-point of the identification window is moved to the end of each new detected word.

11. The method according to any of the previous claims 9-10, characterized in that, if the duration of the duration of the identification window reaches its maximum, the starting point of the identification window is moved to the beginning of the next word nearest to the current position of said starting- point.

12. The method according to claim 9, characterized in that, the identification window is moved to the next word in the speech fragment after successful identification of the current word.

13. The method according to any of the previous claims 9 and 12, characterized in that, the identification window is moved to the next word in the speech fragment after failing identification of the current word.

14. The method according to any previous claim 2-13, characterized in that, word identification models are used following an order going from the shortest model duration to longest model duration.

15. System for controlling an electromechanical device, the system comprising:- a brain fragment capturing element for capturing a brain data fragment comprising imagined, performed, mouthed or attempted speech from a user;- a non-volatile memory element for storing at least one word identification model, at least one predetermined command and a number of predetermined responses;- a processing unit in communication with said brain fragment capturing element, and said non-volatile memory, said processing unit being configured to execute said at least one word identification model; and - a sound emitting element configured to communicate with the user, said sound emitting element being in communication with said processing unit; characterized in that, the non-volatile memory further comprises instructions to carry out the method according to any of the claims 1-14.