A language decoding method and device, electronic equipment and storage medium
By combining EEG decoding scores and linguistic probability scores in EEG signal decoding, and dynamically planning candidate sentence sequences, the problem of insufficient accuracy and naturalness in language decoding in existing technologies is solved, and more natural and accurate language output is achieved.
Patent Information
- Application Number
- CN202610673217.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2026-05-15
- Publication Date
- 2026-08-25
- Estimated Expiration
- 2046-05-15
AI Technical Summary
Existing technologies for language decoding based on electroencephalogram (EEG) signals are insufficient in terms of accuracy, comprehensibility, and naturalness of interaction in achieving continuous, fluent, and natural language output.
By acquiring the EEG decoding results at the current time step, and based on the EEG decoding scores and linguistic probability scores of preset syllables, a candidate sentence sequence is dynamically planned. Combined with sentence end detection, a language decoding process is constructed, and sentences are generated based on multi-source information.
It improves the accuracy, comprehensibility, and naturalness of language decoding results, making the language output closer to natural language communication.
Smart Images

Figure CN122220477B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of brain-computer interface and information processing technology, and in particular to a language decoding method, device, electronic device and storage medium. Background Technology
[0002] Brain-computer interface (BCI) technology aims to establish a direct communication pathway between the brain and external devices. One important research direction in BCI is decoding the user's intended meaning based on electroencephalogram (EEG) signals and translating it into specific language content. This has significant application value for patients who have lost motor or language abilities.
[0003] However, the language decoding method based on EEG signals in related technologies still faces significant challenges in achieving continuous, fluent and natural language content output. The accuracy, comprehensibility (understandability) and naturalness of the language decoding results still need to be improved. Summary of the Invention
[0004] To address the problems of existing technologies, embodiments of this application provide a language decoding method, apparatus, electronic device, and storage medium. The technical solution is as follows: On the one hand, a language decoding method is provided, the method comprising: Obtain the EEG decoding result at the current time step; the EEG decoding result is obtained based on the EEG signal to be processed of the target object at the current time step, and the EEG decoding result includes multiple preset syllables and the EEG decoding score corresponding to each preset syllable; For each preset syllable, based on the candidate characters associated with the preset syllable, the historical candidate sentences of the previous time step are expanded to obtain the set of expanded sentences corresponding to the preset syllable; For each extended sentence corresponding to the preset syllable, the comprehensive language decoding score of the extended sentence is determined based on the historical comprehensive language decoding score of the previous time step, the first linguistic probability score corresponding to the extended sentence, and the EEG decoding score corresponding to the preset syllable; the first linguistic probability score represents the co-occurrence probability of the historical candidate sentences and candidate characters used for extension. Based on the comprehensive language decoding score of the extended sentences corresponding to each preset syllable, the candidate sentence sequence for the current time step is determined; Sentence end detection is performed based on the candidate sentence sequence at the current time step, and the language content to be output at the current time step is determined based on the detection result of the sentence end detection.
[0005] On the other hand, a language decoding apparatus is provided, the apparatus comprising: The EEG decoding result acquisition unit is used to acquire the EEG decoding result at the current time step; the EEG decoding result is obtained based on the EEG signal to be processed of the target object at the current time step, and the EEG decoding result includes multiple preset syllables and the EEG decoding score corresponding to each preset syllable; The sentence expansion unit is used to expand the historical candidate sentences of the previous time step for each preset syllable based on the candidate characters associated with the preset syllable, so as to obtain the expanded sentence set corresponding to each preset syllable; The integrated language decoding score determination unit is used to determine the integrated language decoding score of each extended sentence corresponding to the preset syllable based on the historical integrated language decoding score of the previous time step, the first linguistic probability score corresponding to the extended sentence, and the EEG decoding score corresponding to the preset syllable; the first linguistic probability score represents the co-occurrence probability of the historical candidate sentences and candidate characters used for expansion; The candidate sentence determination unit is used to determine the candidate sentence sequence at the current time step based on the comprehensive language decoding score of the extended sentences corresponding to each preset syllable; The language content to be output unit is used to perform sentence end detection based on the candidate sentence sequence at the current time step, and determine the language content to be output at the current time step based on the detection result of the sentence end detection.
[0006] In some implementations, the integrated language decoding score determination unit includes: The weight coefficient acquisition unit is used to acquire the first weight coefficient, the second weight coefficient, and the third weight coefficient; The weighted summation unit is used to perform a weighted summation of the historical comprehensive language decoding score of the previous time step, the first linguistic probability score corresponding to the extended sentence, and the EEG decoding score corresponding to the preset syllable based on the first weight coefficient, the second weight coefficient, and the third weight coefficient, to obtain the comprehensive language decoding score of the extended sentence.
[0007] In some implementations, the weight coefficient acquisition unit is specifically used to: determine the first weight coefficient, the second weight coefficient, and the third weight coefficient based on the current total number of characters in the expanded sentence; wherein the first weight coefficient and the second weight coefficient are both positively correlated with the current total number of characters, and the third weight coefficient is negatively correlated with the current total number of characters.
[0008] In some implementations, the output language content determination unit, when performing sentence end detection based on the candidate sentence sequence at the current time step, specifically performs the following: For each candidate sentence at the current time step, based on the historical comprehensive language decoding score and the second linguistic probability score of the previous time step, determines the sentence completeness score corresponding to the candidate sentence, where the second linguistic probability score represents the linguistic probability score of the historical candidate sentence corresponding to the candidate sentence as a complete sentence; based on the third linguistic probability score and the EEG decoding score of the preset syllable corresponding to the candidate sentence, determines the language decoding score when the preset syllable is the beginning of a sentence, where the third linguistic probability score represents the linguistic probability score of the candidate character of the preset syllable corresponding to the candidate sentence as the beginning of a sentence at the current time step; based on the sentence completeness score of the candidate sentence and the language decoding score when the corresponding preset syllable is the beginning of a sentence, determines the detection score corresponding to the candidate sentence; and based on the detection scores and comprehensive language decoding scores corresponding to each candidate sentence, determines the detection result of the sentence end detection.
[0009] In some implementations, the output language content determination unit, when determining the detection result of sentence end detection based on the detection score and comprehensive language decoding score corresponding to each candidate sentence, is specifically configured to: determine the highest score based on the detection score and comprehensive language decoding score corresponding to each candidate sentence; if the highest score is any of the detection scores, then determine that the detection result of sentence end detection is that the sentence is complete; if the highest score is any of the comprehensive language decoding scores, then determine that the detection result of sentence end detection is that the sentence is incomplete.
[0010] In some implementations, when the output language content determination unit determines the output language content of the current time step based on the detection result of the sentence end detection, it is specifically used to: if the detection result of the sentence end detection is that the sentence is not completed, then take the candidate sentence corresponding to the highest score as the output language content of the current time step; if the detection result of the sentence end detection is that the sentence is completed, then determine the output language content of the current time step based on the language decoding score when each preset syllable is the beginning of the sentence, and update the candidate sentence sequence of the current time step.
[0011] In some implementations, the candidate sentence determination unit is specifically used to: select a preset number of extended sentences with the highest integrated language decoding scores as the candidate sentence sequence for the current time step.
[0012] On the other hand, an electronic device is provided, including a processor and a memory, wherein the memory stores at least one instruction or at least one program, the at least one instruction or the at least one program being loaded and executed by the processor to implement the language decoding method of any of the above aspects.
[0013] On the other hand, a computer-readable storage medium is provided, wherein at least one instruction or at least one program is stored therein, the at least one instruction or the at least one program being loaded and executed by a processor to implement the language decoding method as described above.
[0014] On the other hand, a computer program product is provided, which includes a computer program that, when executed by a processor, implements the language decoding method of any of the above aspects.
[0015] This application embodiment targets the EEG signal of the target object. It acquires the EEG decoding result at the current time step, which includes multiple preset syllables and the decoding confidence score (i.e., EEG decoding score) for each preset syllable. Then, for each preset syllable, it expands the historical candidate sentences from the previous time step based on its associated candidate characters. For each expanded sentence, it determines the comprehensive language decoding score based on the historical comprehensive language decoding score from the previous time step, the first linguistic probability score of the expanded sentence, and the EEG decoding score of its corresponding preset syllable. Based on this, a candidate sentence sequence for the current time step is formed. Furthermore, by detecting sentence endings based on the candidate sentence sequence at the current time step, the language content to be output is determined. This constructs language decoding as a sequence search process based on dynamic programming of multi-source information. This makes the language decoding process not only dependent on the credibility of the current neural signal, but also continuously constrained and guided by the rules of language structure and semantic boundaries. As a result, the final language decoding result is more in line with human language expression and understanding habits, which greatly improves the accuracy, comprehensibility (intelligibility), and naturalness of the output language content in real communication scenarios. This makes the language output process based on EEG signals closer to the natural language communication process. Attached Figure Description
[0016] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0017] Figure 1 This is a schematic diagram of an implementation environment provided in an embodiment of this application; Figure 2 This is a flowchart illustrating a language decoding method provided in an embodiment of this application; Figure 3 This is a flowchart illustrating another language decoding method provided in an embodiment of this application; Figure 4 This is an example of the language decoding process provided in the embodiments of this application; Figure 5 This is an overall schematic diagram of the language decoding process based on electroencephalogram (EEG) signals provided in the embodiments of this application; Figure 6 This is a structural block diagram of a language decoding device provided in an embodiment of this application; Figure 7 This is a hardware structure block diagram of an electronic device provided in an embodiment of this application. Detailed Implementation
[0018] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of this application.
[0019] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this application described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or server that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or devices.
[0020] In the embodiments of this application, the terms "module" or "unit" refer to a computer program or part of a computer program that has a predetermined function and works with other related parts to achieve a predetermined goal, and can be implemented wholly or partially using software, hardware (such as processing circuitry or memory), or a combination thereof. Similarly, a processor (or multiple processors or memory) can be used to implement one or more modules or units. Furthermore, each module or unit can be part of an overall module or unit that includes the functionality of that module or unit.
[0021] It is understood that in the specific embodiments of this application, data such as user information are involved. When the above embodiments of this application are applied to specific products or technologies, user permission or consent is required, and the collection, use and processing of related data must comply with the relevant laws, regulations and standards of the relevant countries and regions.
[0022] Please see Figure 1 The diagram shown is a schematic diagram of an implementation environment provided in this application embodiment. The implementation environment includes an EEG acquisition device 110, an audio acquisition device 120, and a computer device 130.
[0023] The EEG acquisition device 110 can acquire EEG signals, such as electrocorticography (ECoG) signals, from a user or subject using a non-invasive brain-computer interface. The audio acquisition device 120 can acquire the user's or subject's speech audio; for example, the audio acquisition device can be a microphone. The computer device 130 can connect and communicate with both the EEG acquisition device 110 and the audio acquisition device 120, enabling the audio acquisition device 120 to send real-time acquired speech audio to the computer device 130, and the real-time EEG signals acquired by the EEG acquisition device 110 to the computer device 130.
[0024] The computer device 130 can determine the language content to be output based on the acquired EEG signals of the user or subject using the language decoding method of this application embodiment, thereby realizing language content decoding based on EEG signals and outputting the speech intent expressed by the user or subject. The computer device 130 can store the articulation boundary model and EEG decoding model of the user or subject, which are obtained through machine learning based on the user's or subject's EEG-syllable corpus.
[0025] In this embodiment of the application, the computer device 130 may be a terminal and / or a server. It may be a single computer device or a system composed of multiple computer devices. When the computer device 130 includes multiple computer devices, the multiple computer devices can communicate with each other through wired or wireless network connections.
[0026] In a specific application scenario, computer device 130 may include a first computer device 131, a second computer device 132, and a third computer device 133. The first computer device 131 can connect and communicate with the EEG acquisition device 110 and the audio acquisition device 120 to generate training corpus for the target object to construct an EEG-syllable corpus for that target object. The second computer device 132 is used to train a model based on the target object's EEG-syllable corpus to obtain an EEG decoding model for the target object. It can also train other necessary models, such as articulation boundary models. The third computer device 133 can connect and communicate with the EEG acquisition device 110 to call the models trained by the second computer device 132, such as the EEG decoding model and the articulation boundary model, to decode the target object's EEG signals to determine the corresponding language content.
[0027] It should be noted that the terminals involved in the embodiments of this application may include, but are not limited to, mobile phones, computers, smart voice interaction devices, smart home appliances, vehicle terminals, aircraft, etc. The servers involved in the embodiments of this application may be independent physical servers, server clusters or distributed systems composed of multiple physical servers, or cloud servers providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, CDN (Content Delivery Network), and big data and artificial intelligence platforms.
[0028] Please see Figure 2 The diagram shown is a flowchart illustrating a language decoding method provided in an embodiment of this application. This method can be applied to... Figure 1 The computer device 130 is described. It should be noted that this specification provides the operational steps of the methods described in the embodiments or flowcharts, but based on conventional or non-inventive labor, more or fewer operational steps may be included. The order of steps listed in the embodiments is merely one possible execution order among many and does not represent the only execution order. In actual system or product execution, the methods shown in the embodiments or drawings can be executed sequentially or in parallel (e.g., in a parallel processor or multi-threaded processing environment). Specifically, as shown... Figure 2 As shown, the method may include: S201, Obtain the EEG decoding result at the current time step.
[0029] The EEG decoding result is obtained by decoding the EEG signal of the target object at the current time step. The EEG decoding result includes multiple preset syllables and the EEG decoding score corresponding to each preset syllable.
[0030] It should be noted that the electroencephalogram (EEG) decoding scores corresponding to each preset syllable in the EEG decoding results of the embodiments of the present application can also be referred to as decoding confidence or confidence, which represents the probability that the to-be-processed EEG signal belongs to the corresponding preset syllable.
[0031] The method for obtaining the EEG decoding result of the current time step will be described in detail in the subsequent content of the embodiments of the present application.
[0032] S203. For each of the preset syllables, based on the candidate characters associated with the preset syllable, expand the historical candidate sentence of the previous time step to obtain an extended sentence set corresponding to the preset syllable.
[0033] In the embodiments of the present application, a syllable-candidate character mapping relationship library can be pre-constructed or loaded. This mapping relationship library is constructed based on a general Chinese dictionary or a dictionary in a specific field, covering all syllables and common Chinese characters required for the language decoding task in the embodiments of the present application. For example, it includes 401 to 418 syllables and multiple candidate Chinese characters corresponding to each syllable. For example, the candidate characters associated with the syllable "yi" include: "一", "以", "意", "义", "已", and the candidate characters associated with the syllable "shi" can include: "是", "时", "事", "市", "实".
[0034] Then, the candidate characters associated with each preset syllable in the EEG decoding result can be determined based on this syllable-candidate character mapping relationship library. It can be understood that the candidate characters associated with each preset syllable can be one or more.
[0035] Among them, the historical candidate sentence of the previous time step refers to the set of candidate sentences screened at the previous moment (such as the t-1 moment) in the language decoding process. Usually, it is an ordered sequence, and each historical candidate sentence has its comprehensive language decoding score in the language decoding process at the t-1 time step, hereinafter referred to as the historical comprehensive language decoding score. And the language content output by the language decoding at the t-1 time step is the candidate sentence selected from this set of candidate sentences. In specific implementation, a candidate sentence sequence can be maintained for the language decoding process. The candidate sentences in this candidate sentence sequence are updated as the language decoding process progresses. That is, when performing language decoding at the t time step, the candidate sentence sequence stores the candidate sentences screened at the t-1 time step. When the language decoding at the t time step ends and enters the language decoding at the t + 1 time step, the candidate sentence sequence is updated to the candidate sentences screened at the t time step.
[0036] When performing the sentence expansion operation, assume that the historical candidate sentence of the previous time step is expressed as For a preset syllable syl_k decoded at the current time step t, we can first query all candidate characters corresponding to the preset syllable syl_k according to the syllable-candidate character mapping database, for example, denoted as set . Furthermore, for each historical candidate sentence and each candidate character Performing string concatenation involves appending candidate words to the end of historical candidate sentences to generate an expanded sentence. traversal All historical candidate sentences and All candidate characters in the set will generate all extended sentences that constitute the set of extended sentences corresponding to the preset syllable syl_k.
[0037] Understandably, the number of extended sentences in the extended sentence set corresponding to a preset syllable is equal to the product of the number of historical candidate sentences and the number of candidate characters corresponding to the preset syllable. Therefore, in order to save computing resources and improve the speed of language decoding, in some examples, for each preset syllable, only the most commonly used first few (such as the first 3 or the first 5) candidate characters can be selected to participate in the extension.
[0038] By performing the above expansion process on each preset syllable in the EEG decoding results, an expanded sentence set corresponding to each preset syllable can be obtained. The expanded sentence set of all preset syllables constitutes the expanded sentence space of the current time step.
[0039] S205, for each extended sentence corresponding to the preset syllable, the comprehensive language decoding score of the extended sentence is determined based on the historical comprehensive language decoding score of the previous time step, the first linguistic probability score corresponding to the extended sentence, and the EEG decoding score corresponding to the preset syllable.
[0040] The first linguistic probability score represents the co-occurrence probability of the historical candidate sentence and candidate characters used in the expansion. Specifically, the first linguistic probability score can be represented by the conditional probability of the corresponding candidate character appearing subsequently, given the historical candidate sentence used in the expansion. .
[0041] In some examples, the first linguistic probability score of an expanded sentence can be obtained based on a statistical linguistic model such as an N-gram model. Specifically, when it is necessary to calculate the candidate words... Add to history candidate sentences When calculating the first linguistic probability score of the extended sentence at the end, the N-gram model can extract a string of a certain length from the end of the historical candidate sentence as contextual information for the current prediction, and query the N-gram model for candidate words that appear given that context. The conditional probability is used to obtain the first linguistic probability score of the expanded sentence. In other examples, the first linguistic probability score of the expanded sentence can also be achieved through a linguistic probability model built using deep learning. This model can predict the probability that the next character is a corresponding candidate character based on the input historical candidate sentences. The higher the probability, the more the expansion conforms to human language rules. The first linguistic probability score of the corresponding expanded sentence can be obtained based on the probability output by this linguistic probability model. The linguistic probability model built using deep learning can include recurrent neural networks, long short-term memory networks, etc.
[0042] Specifically, each extended sentence obtained in step S203 has the following correspondence: extended sentence - preset syllable - EEG decoding score - historical candidate sentence - historical comprehensive language decoding score - candidate character. The EEG decoding score in this correspondence is the EEG decoding score of the preset syllable corresponding to the candidate character at the end of the corresponding extended sentence in the EEG decoding result. Therefore, after determining the first linguistic probability score corresponding to each extended sentence, for each extended sentence, its corresponding comprehensive language decoding score can be calculated based on its first linguistic probability score, EEG decoding score, and historical comprehensive language decoding score.
[0043] In some examples, when calculating the comprehensive language decoding score of an expanded sentence, the historical comprehensive language decoding score used can be the comprehensive language decoding score calculated in the previous time step of the historical candidate sentence upon which the expanded sentence is based. In other examples, when calculating the comprehensive language decoding score of an expanded sentence, the historical comprehensive language decoding score used can also be the comprehensive language decoding score corresponding to the candidate sentence output in the previous time step. In this case, the historical comprehensive language decoding score used when calculating the comprehensive language decoding score for all expanded sentences in the current time step is the same value, which is beneficial to improving language decoding speed. Understandably, when the current time step is at the beginning of a sentence, the historical comprehensive language decoding score is recorded as 0.
[0044] In practice, the comprehensive language decoding score of an extended sentence can be the sum of its corresponding first linguistic probability score, EEG decoding score, and historical comprehensive language decoding score.
[0045] In some implementations, step S205 above, when determining the comprehensive language decoding score of the extended sentence based on the historical comprehensive language decoding score of the previous time step, the first linguistic probability score corresponding to the extended sentence, and the EEG decoding score corresponding to the preset syllable, may include: Obtain the first weight coefficient, the second weight coefficient, and the third weight coefficient; Based on the first weighting coefficient, the second weighting coefficient, and the third weighting coefficient, the historical comprehensive language decoding score of the previous time step, the first linguistic probability score corresponding to the extended sentence, and the EEG decoding score corresponding to the preset syllable are weighted and summed to obtain the comprehensive language decoding score of the extended sentence.
[0046] Among them, the first weight coefficient corresponds to the historical comprehensive language decoding score item, the second weight coefficient corresponds to the linguistic probability score item, and the third weight coefficient corresponds to the EEG decoding score item, thereby expanding the comprehensive language decoding score of the sentence. It can be expressed as the following formula (1): (1) in, This represents the first weighting coefficient. This represents the second weighting coefficient. Indicates the third weighting coefficient; This indicates the historical comprehensive language decoding score. This represents the first linguistic probability score. This indicates the brainwave decoding score.
[0047] In some examples, the first weighting coefficient, the second weighting coefficient, and the third weighting coefficient can be fixed values set based on experience.
[0048] In some implementations, obtaining the first weight coefficient, the second weight coefficient, and the third weight coefficient may include: determining the first weight coefficient, the second weight coefficient, and the third weight coefficient based on the current total number of characters in the expanded sentence; wherein the first weight coefficient and the second weight coefficient are both positively correlated with the current total number of characters, and the third weight coefficient is negatively correlated with the current total number of characters.
[0049] Specifically, the first, second, and third weighting coefficients can dynamically change based on the total number of characters in the expanded sentence. The first and second weighting coefficients increase with the total number of characters, meaning that the longer the sentence, the greater the contribution of historical decisions and linguistic patterns to the comprehensive language decoding score at the current time step. For example, the first and second weighting coefficients increase with the total number of characters and approach 1. The third weighting coefficient decreases with the total number of characters, meaning that the longer the sentence, the smaller the contribution of the EEG decoding result to the comprehensive language decoding score at the current time step. For example, the third weighting coefficient decreases with the total number of characters and approaches 0.5.
[0050] In specific implementation, the comprehensive language decoding score of the extended sentence under this implementation method It can be expressed as the following formula (2): (2) in, Indicates the total number of characters in the expanded sentence.
[0051] In the above implementation, by introducing a dynamic weight coefficient that adapts to sentence length, the language decoding process can intelligently and dynamically adjust the contribution of different information sources according to different stages of the decoding process. This is more in line with the laws of cognition and signal processing, improves the accuracy of language decoding, and significantly enhances the stability of long sentence decoding.
[0052] S207, Based on the comprehensive language decoding score of the extended sentences corresponding to each preset syllable, determine the candidate sentence sequence for the current time step.
[0053] Specifically, for K preset syllables, each preset syllable corresponds to a set of extended sentences, and each extended sentence can be used to calculate its corresponding comprehensive language decoding score. Merging the extended sentence sets of all preset syllables forms a global candidate pool containing M extended sentences. It can be understood that as the number of decoded syllables increases, if the sentence does not end, the number of obtained extended sentences grows in a tree-like manner. Therefore, in order to reduce computational complexity and improve the efficiency and accuracy of language decoding while ensuring the quality of language decoding, it is necessary to select a controllable and optimal subset of sentences from all possible extended sentence sets generated at the current time step as the "candidate sentence sequence for the current time step," which is used for the final output decision of the current time step and the decoding iteration of the next time step.
[0054] In some implementations, a predetermined number of sentences with the highest comprehensive language decoding scores in the global candidate pool can be selected as the candidate sentence sequence for the current time step. The predetermined number can be set based on practical experience; for example, it can be the top N extended sentences sorted by their comprehensive language decoding scores from highest to lowest as the candidate sentence sequence for the current time step, where N can be any number between 20 and 30.
[0055] S209, Sentence end detection is performed based on the candidate sentence sequence at the current time step, and the language content to be output at the current time step is determined according to the detection result of the sentence end detection.
[0056] Specifically, sentence end detection mainly determines whether a complete sentence has been formed before the current time step, and then outputs the language content based on the result of sentence end detection.
[0057] In some implementations, such as Figure 3 As shown, step S209 above, when performing sentence end detection based on the candidate sentence sequence at the current time step, may include: S301, for each candidate sentence in the current time step, the sentence completeness score corresponding to the candidate sentence is determined based on the historical comprehensive language decoding score and the second linguistic probability score of the previous time step.
[0058] The second linguistic probability score represents the linguistic probability score of the historical candidate sentence corresponding to the candidate sentence as a complete sentence. In specific implementation, the second linguistic probability score can be represented by the conditional probability of the occurrence of an end marker (such as E) indicating the end of a sentence, given a historical candidate sentence. The calculation method of the second linguistic probability score can refer to the aforementioned calculation method of the first linguistic probability score, and will not be repeated here.
[0059] Specifically, the historical comprehensive language decoding score and the second linguistic probability score of the previous time step can be summed to obtain the sentence completeness score of the candidate sentence. The evaluation is "regardless of what you want to say now, does the part that has just been said itself resemble a complete sentence?"
[0060] In practice, the sentence completeness score corresponding to the candidate sentence can be calculated using the following formula (3): (3) in, This indicates the score for sentence completeness; This indicates the historical comprehensive language decoding score of the previous time step; The second linguistic probability score represents the linguistic probability score of the corresponding historical candidate sentence as a complete sentence at the current time step, without adding any new words. and This refers to the aforementioned first and second weighting coefficients. It is understandable that... The higher the score, the more complete the sentence.
[0061] S303, based on the third linguistic probability score and the EEG decoding score of the corresponding preset syllable, determine the language decoding score when the preset syllable is the beginning of a sentence.
[0062] The third linguistic probability score represents the linguistic probability score of the candidate character corresponding to the preset syllable of the candidate sentence as the beginning of the sentence at the current time step. In specific implementation, the third linguistic probability score can be represented by the conditional probability of a sentence-initial identifier (such as B) appearing before the corresponding candidate character. The calculation method of the third linguistic probability score can refer to the aforementioned calculation method of the first linguistic probability score, and will not be repeated here.
[0063] Specifically, the third linguistic probability score and the corresponding preset syllable's EEG decoding score can be summed to obtain the language decoding score when the preset syllable is used as the beginning of a sentence. The evaluation is "How good would it be if we started a new sentence now, beginning with this preset syllable?"
[0064] In specific implementation, the language decoding score when the preset syllable is the beginning of a sentence can be calculated using the following formula (4). : (4) in, The third linguistic probability score represents the linguistic probability score of the candidate character corresponding to the preset syllable of the candidate sentence as the beginning of the sentence at the current time step; This represents the EEG decoding score of the preset syllable in the EEG decoding results at the current time step; This refers to the aforementioned second weighting coefficient. This represents the aforementioned third weighting coefficient.
[0065] S305, based on the sentence completeness score of the candidate sentence and the language decoding score when the corresponding preset syllable is the beginning of the sentence, determine the detection score corresponding to the candidate sentence.
[0066] Specifically, the sentence completeness score of the candidate sentence and the language decoding score when the corresponding preset syllable is the beginning of the sentence can be summed to obtain the detection score corresponding to the candidate sentence. This score represents the comprehensive score of the event that "the historical candidate sentence ends here and the corresponding preset syllable is the beginning of a new sentence".
[0067] In specific implementation, the detection score corresponding to the candidate sentence can be calculated using the following formula (5). : (5) S307, Based on the detection score and comprehensive language decoding score corresponding to each candidate sentence, determine the detection result of sentence end detection.
[0068] For example, the detection result for sentence end detection, based on the detection score and comprehensive language decoding score corresponding to each candidate sentence, may include: The highest score is determined based on the detection score and comprehensive language decoding score corresponding to each candidate sentence; If the highest score is any of the detection scores, then the detection result of the sentence end detection is determined to be that the sentence is complete; If the highest score is any of the comprehensive language decoding scores, then the detection result of the sentence end detection is determined to be that the sentence is incomplete.
[0069] In the specific implementation, after obtaining the detection scores of each candidate sentence, all detection scores and the comprehensive language decoding scores of all candidate sentences in the candidate sentence sequence are sorted in descending order. If a detection score is ranked first, the detection result of the sentence end detection is determined to be that the sentence is complete. That is, the historical candidate sentence in the candidate sentence corresponding to the detection score ranked first is determined to be the previous sentence that has ended, and the preset syllable corresponding to the candidate sentence is determined to be the start of the next sentence. Conversely, if the comprehensive language decoding score is ranked first, the detection result of the sentence end detection is determined to be that the sentence is not complete.
[0070] The above implementation does not rely on a single threshold or fixed rules, but is based on the relative confidence of all predicted hypotheses, which improves the stability of sentence end detection. Moreover, whether a sentence ends depends closely on the integrity of the historical sentence itself and the possibility of the current syllable as a new beginning, which fully considers the context and participates in the global sorting of the current time step, thus improving the accuracy of sentence end detection.
[0071] In some examples, to save computational resources and reduce unnecessary computational resource consumption to improve the efficiency of language decoding, the process before step S303 may include: sorting the sentence completeness scores and the comprehensive language decoding scores corresponding to each candidate sentence in descending order; if a preset number (e.g., the first 20) of the sorted sentences include one or more sentence completeness scores, then for the candidate sentences corresponding to those one or more sentence completeness scores, steps S303 to S307 are executed. This ensures that only high-quality potentially complete sentences can participate in subsequent detection steps, reducing unnecessary computation, saving computational resources, and thus improving the efficiency and accuracy of language decoding.
[0072] In some implementations, determining the language content to be output at the current time step based on the sentence end detection result may include: If the sentence end detection result is that the sentence is incomplete, then the candidate sentence corresponding to the highest score is taken as the language content to be output at the current time step. If the sentence end detection result indicates that the sentence is complete, then the language content to be output at the current time step is determined based on the language decoding score when each preset syllable is the beginning of the sentence, and the candidate sentence sequence at the current time step is updated.
[0073] Specifically, if the sentence end detection result is that the sentence is incomplete, then the candidate sentence with the highest comprehensive language decoding score can be used as the language content to be output at the current time step and can be output.
[0074] If the sentence end detection result indicates that the sentence is complete, then the language content to be output at the current time step is determined based on the language decoding scores when each of the aforementioned preset syllables is the sentence beginning, and the candidate sentence sequence at the current time step is updated. Specifically, this can be based on the multiple calculated values mentioned above. Sort in descending order and then select the first one. The corresponding sentence is taken as the language content to be output at the current time step, and the number of sentences in the candidate sentence sequence is taken into account. The corresponding sentence is updated in the candidate sentence sequence of the current time step obtained in the aforementioned step S207.
[0075] To facilitate understanding of the language decoding process in the embodiments of this application, the following is combined with... Figure 4 The example shown illustrates this.
[0076] like Figure 4 As shown, in the initial state (time step 0), the candidate sentence sequence is B (sentence start marker), and the historical comprehensive language decoding score is 0. It should be noted that this example uses three preset syllables "ni, hao, wo" from the EEG decoding results to illustrate the sentence search process. In practical applications, the EEG decoding results can include more preset syllables, and this example only shows the top 3 results; in actual applications, there could be more, such as the top 20.
[0077] Time step 1: EEG decoding results: ni 0.4 hao 0.3 wo 0.3 ……; where the EEG decoding score corresponding to the preset syllable ni is 0.4, the EEG decoding score corresponding to the preset syllable hao is 0.3, and the EEG decoding score corresponding to the preset syllable wo is 0.3.
[0078] The candidate sentence sequence for time step 1 is: {Byou1.2;Bme1.0;Bhao0.3}; where the comprehensive language decoding score of the candidate sentence "Byou" is 1.2, the comprehensive language decoding score of the candidate sentence "Bme" is 1.0, and the comprehensive language decoding score of the candidate sentence "Bhao" is 0.3, all of which are calculated based on the aforementioned formula (1). The value is 0. Therefore, the language content to be output at time step 1 is "you", and "you" is output.
[0079] Time step 2: EEG decoding results: {ni 0.1 hao 0.4 wo 0.5 ……}; Candidate sentence sequence at time step 2: {B Hello 2; B You and me 1.7; B I'm fine 1.5}; This candidate sentence sequence is obtained by performing the foregoing steps S203 to S209 of the embodiment of the present application based on the historical comprehensive language decoding score of 1.2 at time step 1 and the historical candidate sentences {B You; B Me; B Fine}. The sentence end detection indicates that the sentence is not completed. Thus, the language content to be output at this time step 2 is "Hello", and "Hello" is output.
[0080] Time step 3: Brain decoding result: {ni 0.4 hao 0.3 wo 0.3 ……}; Candidate sentence sequence at time step 3: {B Hello E 2.4; B Hello you 2.4; B You me fine 2.0}, where adding E indicates the end of the sentence; This candidate sentence sequence is obtained by performing the foregoing steps S203 to S209 of the embodiment of the present application based on the historical comprehensive language decoding score of 2 at time step 2 and the historical candidate sentences {B Hello; B You me; B I'm fine}. The sentence end detection indicates that the sentence is completed. Thus, a new sentence starts. At this time, according to multiple Perform a descending sort, and take the highest corresponding sentence as the beginning of the new sentence to output. For example, Figure 4 the second sentence beginning "I" in At the same time, use the top 3 corresponding sentences (i.e., the new sequence) to update the candidate sentence sequence. Thus, the updated candidate sentence sequence is similar to the foregoing time step 1, consisting of the sentence start marker B and the subsequent candidate characters.
[0081] The technical solution of the embodiment of the present application constructs language decoding as a sequence search process based on multi-source information dynamic programming, making the language decoding process not only depend on the credibility of the current neural signal, but also be continuously constrained and guided by language structure rules and semantic boundaries. Thus, the final language decoding result is more in line with human language expression and understanding habits, greatly improving the accuracy, intelligibility (understandability), and naturalness of interaction of the output language content in real communication scenarios, making the language output process based on brain electrical signals closer to the natural language communication process.
[0082] Next, an explanation is given for Figure 5 the technical solution shown in Figure 5 This is the overall schematic diagram of the language decoding process based on brain electrical signals provided by the embodiment of the present application.
[0083] As [[ID=29]] Figure 5 shown, it includes constructing an electroencephalogram-syllable corpus of the target object and a model training stage, and a subsequent real-time decoding application stage for the electroencephalogram signal of the target object.
[0084] First, we will introduce the construction process of the target object's EEG-syllable corpus, which includes the following steps (1) to (5): Step (1): Obtain the audio signal and EEG signal of the target object; the EEG signal is the synchronous EEG signal of the target object when it outputs the audio signal based on the training text.
[0085] The target subject can be any brain-computer interface subject or user. Audio signals can be recorded while the target subject reads training text, and the target subject's electroencephalogram (EEG) signals are simultaneously acquired, resulting in the target subject's audio signal and corresponding EEG signal. Understandably, the target subject's audio signal and EEG signal are time-aligned. The target subject's audio signal includes at least one syllable. A syllable is the smallest unit of speech formed by the combination of a single vowel phoneme and a consonant phoneme. A single vowel phoneme can also form a syllable on its own; that is, a syllable includes at least one phoneme.
[0086] In specific implementation, in order to improve the quality of training corpus, the training text needs to meet at least one of the following requirements: (1) The training text includes all Chinese syllables, usually 402 to 418 syllables, so as to ensure that the training corpus can fully cover the syllables in Chinese, which is conducive to improving the completeness and stability of the EEG decoding model trained based on the training corpus. (2) Each syllable in the training text appears at least 60 times to ensure that each syllable has enough training corpus and avoid the problem of insufficient and unstable model learning in the later stage due to too few samples of a single syllable. For example, each syllable appears 60 to 120 times in the training text. (3) The syllable correlation degree in the training text conforms to the distribution of modern Chinese, where the syllable correlation degree can be obtained by conditional probability P( | ) representation, that is, when the previous syllable When it appears, the next syllable is The probability of this. Thus, when subjects read texts that conform to the rules of natural language, the neural processes of their speech planning and articulation are fluent and natural. This helps to produce purer and stronger EEG signals that are highly correlated with natural speech, thereby improving the signal-to-noise ratio and specificity of EEG signals, and thus improving the quality of training data.
[0087] Step (2): Based on the time-domain acoustic intensity envelope of the audio signal, determine the candidate audio segments in the audio signal that meet the preset acoustic intensity conditions.
[0088] The temporal intensity envelope of the audio signal is the sound intensity change curve of the audio signal over time. Candidate audio segments that meet the preset intensity conditions in the audio signal can include one or more. Each candidate audio segment corresponds to a start time and an end time in the audio signal. Therefore, detecting one or more candidate audio segments through the temporal intensity envelope of the audio signal yields one or more time boundaries for the candidate audio segments. It can be understood that the time boundaries of these one or more candidate audio segments include the time boundaries of syllables, i.e., the start and end times of the syllables in the audio signal. Therefore, step (2) can also be performed as follows: Figure 5 The diagram shows what is called syllable boundary detection, which may also include the time boundaries of non-syllable sounds such as coughs.
[0089] The preset sound intensity conditions may include a sound intensity exceeding a sound intensity threshold and a duration exceeding a time threshold. The sound intensity threshold and time threshold may be fixed preset values or dynamically calculated changing values.
[0090] In some implementations, to improve the accuracy of audio time boundary detection, determining candidate audio segments in the audio signal that meet preset sound intensity conditions based on the time-domain sound intensity envelope of the audio signal may include: for the current audio frame of the audio signal, determining the original sound intensity of the current audio frame based on the instantaneous energy of a first preset number of frequency points within the current audio frame; averaging the original sound intensity of the current audio frame and the original sound intensities of a second preset number of audio frames preceding the current audio frame to obtain the target sound intensity of the current audio frame; obtaining the time-domain sound intensity envelope of the audio signal based on the target sound intensity of each audio frame in the audio signal; and determining audio segments whose target sound intensity exceeds a preset sound intensity threshold and whose duration exceeds a preset duration as candidate audio segments based on the time-domain sound intensity envelope. Specifically, the audio signal can be divided into multiple audio frames of equal length. Then, each audio frame is traversed sequentially. For the current audio frame, the original sound intensity is determined based on the instantaneous energy of a first preset number of frequency points within that current audio frame. This original sound intensity is then averaged with the original sound intensities of the previous second preset number of audio frames. This average sound intensity is used as the target sound intensity of the current audio frame. Based on the target sound intensities of each audio frame, the temporal sound intensity envelope of the audio signal is obtained. This temporal sound intensity envelope can then be scanned. Whenever it rises above a preset sound intensity threshold and lasts for more than a preset duration (e.g., 100ms), the audio segment between the start and end times (usually when it falls back below the preset sound intensity threshold) is marked as a candidate audio segment. The first and second preset numbers can both be set based on practical experience.
[0091] In practical applications, the target sound intensity of an audio frame can be calculated using the following formula:
[0092] in, A signal representing an audio frame; The number of frequency points is the first preset number; Instantaneous energy at a frequency point can also be called instantaneous power; The reference power can be set based on the acoustic hearing threshold, for example, it can be... ; -1 indicates the second preset number of audio frames preceding this audio frame; Indicates the first The target sound intensity of the audio frame; in, This indicates the original sound intensity of the audio frame, measured in decibels.
[0093] The above implementation determines the original sound intensity of the current audio frame and filters out rapid, meaningless energy fluctuations (such as consonant plosives and slight inhalation sounds) by averaging M frames. This makes the temporal sound intensity envelope of the obtained audio signal more clearly show the stable intensity of the syllable. Furthermore, by combining a preset sound intensity threshold and a preset duration, one or more stable and meaningful candidate audio segments can be accurately and reliably detected to obtain accurate boundary information.
[0094] In some possible implementations, considering that different people have different voice intensities, and the same person's voice intensity may also change at different times and in different environments, in order to improve the accuracy and stability of candidate audio segment detection, a preset sound intensity threshold is a preset multiple of the target statistical value, which is the statistical value of the target sound intensity within a previous preset time period.
[0095] For example, the target statistical value can be the standard deviation of the target sound intensity within a previous preset time period. The length of this preset time period can be set based on practical experience, such as 10 seconds. The preset multiplier can also be set based on practical experience, such as between 1.2 and 2.0 times, preferably 1.5, thus the preset sound intensity threshold is 1.5 times the standard deviation of the target sound intensity within the previous 10 seconds.
[0096] In specific implementation, during the scanning of the temporal sound intensity envelope, when the current scanning duration is less than the length of a preset time period (e.g., 10 seconds), candidate audio segments can be detected based on the default preset sound intensity threshold. When the current scanning duration is greater than or equal to the length of the preset time period (e.g., 10 seconds), the preset sound intensity threshold is dynamically calculated based on 1.5 times the standard deviation of the target sound intensity within the previous preset time period (e.g., 10 seconds).
[0097] In other possible implementations, in order to improve the detection accuracy of syllable time boundaries by ensuring that the detected candidate audio segments are as close as possible to the audio segments corresponding to syllables, the following formula can be used when determining the candidate audio segments in the audio signal that meet the preset sound intensity conditions based on the time-domain sound intensity envelope of the audio signal:
[0098] in, Indicates duration (unit: seconds); This represents the average sound intensity of the current audio segment, which is the average value of the target sound intensity within the current audio segment; The standard deviation of background noise intensity; This represents the mean value of the background noise intensity. Represents a minimal constant, such as 1e 8. Numerical stability terms to prevent the denominator from being zero; This represents the time threshold (unit: seconds), which can be set based on practical experience, such as 0.3 seconds.
[0099] Specifically, based on the temporal sound intensity envelope, audio segments whose target sound intensity exceeds a fixed sound intensity threshold (which can be set based on practical experience) are detected. The duration T of the audio segment is obtained based on its start and end times. If the duration T is less than 0.1 seconds, the audio segment is determined to be a non-candidate audio segment; conversely, if the duration T is greater than or equal to 0.1 seconds, the average sound intensity I is calculated based on the target sound intensity within the audio segment. This, in turn, determines the degree to which the average sound intensity I deviates from the noise baseline, i.e., calculates... Then multiply this by T. If the resulting product is greater than the time threshold... If the audio segment is selected as a candidate audio segment, then the product result is determined to be no greater than the time threshold. Then the audio segment is determined to be a non-candidate audio segment.
[0100] This implementation can effectively distinguish speech from transient noise, adapt to the volume differences of different speakers, and maintain stable detection performance in different noise environments, thus providing reliable speech segments for the determination of syllable boundaries.
[0101] Step (3): Input the audio signal into the syllable detection model to perform syllable detection and obtain the syllable detection result; the syllable detection result indicates the syllable audio segment in the audio signal and the preset syllable and confidence level corresponding to the syllable audio segment.
[0102] Here, a syllable audio segment refers to the time interval corresponding to a syllable in an audio signal. The syllable detection model is a pre-trained neural network model that can perform syllable detection on the input audio signal and output a syllable detection result. This result indicates the syllable audio segment in the audio signal, along with the corresponding preset syllable and a confidence level. The confidence level represents the probability that the syllable audio segment belongs to the preset syllable. For example, the syllable detection result can be represented as {( , ),syl,z}, where, ( , ) represents the start and end times of the syllable audio segment, syl represents the preset syllable, and z represents the confidence level.
[0103] Based on this, embodiments of this application may further include a step of training the syllable detection model, specifically including: Obtain the sample audio signal and the corresponding label. The label indicates the reference syllable audio segment in the sample audio signal, the reference preset syllable corresponding to the reference syllable audio segment, and the probability of being the reference preset syllable. It can be understood that the probability of being the reference preset syllable is 0 (i.e. not the reference preset syllable) or 1 (i.e. being the reference preset syllable). The sample audio signal is input into the syllable detection model to be trained to detect syllables and obtain the predicted syllable detection result; the predicted syllable detection result indicates the predicted syllable audio segment in the sample audio signal, as well as the predicted preset syllable and confidence level corresponding to the predicted syllable audio segment; Based on the difference between the predicted syllable detection result and the labeled label corresponding to the sample audio signal, the model parameters of the syllable detection model to be trained are adjusted, and the training continues iteratively until the training termination condition is met, thus obtaining the trained syllable detection model.
[0104] Among them, the reference preset syllable is any one of multiple preset syllables, which can cover all Chinese syllables, such as 402 to 418 syllables in Chinese.
[0105] The adjustment of the model parameters of the syllable detection model to be trained, based on the difference between the predicted syllable detection result and the labeled label corresponding to the sample audio signal, can be achieved by: using a preset loss function, determining the loss function value based on the difference between the predicted syllable detection result and the labeled label, and then using this loss function value to adjust the model parameters of the syllable detection model to be trained. For example, the preset loss function can be the cross-entropy loss function. The model parameters of the syllable detection model to be trained can be adjusted in the direction of reducing the loss function value. For instance, when training the syllable detection model, a backpropagation algorithm (such as stochastic gradient descent (SGD)) can be used to adjust the model parameters in the direction of reducing the loss function value. The training termination condition can be that the number of iterations reaches an iteration threshold (which can be set as needed), the loss function value is less than the loss function threshold (which can be set as needed), or the difference between two adjacent loss function values reaches a difference threshold (which can be set as needed).
[0106] In some possible implementations, the syllable detection model may include an acoustic feature extraction network, a convolutional neural network, a time-delay neural network, and a long short-term memory network cascaded in sequence; the step of inputting the audio signal into the syllable detection model for syllable detection and obtaining the syllable detection result includes: inputting the audio signal into the syllable detection model, and processing the audio signal sequentially through the acoustic feature extraction network, the convolutional neural network, the time-delay neural network, and the long short-term memory network of the syllable detection model to obtain the output syllable detection result.
[0107] Step (4): Based on the syllable detection results, determine the target syllable audio segment among the candidate audio segments.
[0108] This application embodiment uses syllable detection results to correct candidate audio segments, thereby determining the target syllable audio segment. The target syllable audio segment is the syllable audio segment among the candidate audio segments that meets the correction conditions. The correction conditions may be that the audio overlap between the syllable audio segment of the syllable detection result and the candidate audio segment meets the audio overlap threshold. Audio overlap characterizes the degree of overlap between two audio segments. For example, the audio overlap can be characterized by the degree of overlap between two audio segments in the time dimension. The audio overlap threshold can be set based on practical experience.
[0109] In some possible implementations, to further improve the accuracy of the target syllable audio segment and thus the accuracy of syllable boundaries, determining the target syllable audio segment among the candidate audio segments based on the syllable detection results may include: determining the audio overlap between the syllable audio segment indicated by the syllable detection results and the candidate audio segments; selecting candidate syllable audio segments from the candidate audio segments whose audio overlap satisfies a preset overlap condition; and selecting the target syllable audio segment from the candidate syllable audio segments that satisfies a preset confidence threshold based on the confidence level of the preset syllable corresponding to the candidate syllable audio segment.
[0110] The preset overlap condition indicates overlap, and this preset overlap condition is consistent with the representation method of audio overlap. The candidate syllable audio segment is an audio segment of a syllable selected from the candidate audio segments, and is used as a candidate. The audio overlap between the syllable audio segment and the candidate audio segment meeting the preset overlap condition indicates that the syllable audio segment and the candidate audio segment point to the same syllable.
[0111] For example, when determining the audio overlap between the syllable audio segment indicated by the syllable detection result and the candidate audio segment, the deviation between the start time of the syllable audio segment and the start time of the candidate audio segment can be determined to obtain a first time deviation; and the deviation between the end time of the syllable audio segment and the end time of the candidate audio segment can be determined to obtain a second time deviation; and then, based on the first time deviation and the second time deviation, the audio overlap between the syllable audio segment and the candidate audio segment can be determined. Specifically, if the first time deviation does not exceed a preset time deviation threshold and the second time deviation does not exceed the preset time deviation threshold, the audio overlap between the syllable audio segment and the candidate audio segment can be determined to be coincident, meaning that the syllable audio segment and the candidate audio segment point to the same syllable audio, meeting the preset overlap condition (i.e., coincident). In this case, the candidate audio segment is determined as the candidate syllable audio segment. Conversely, if the first time deviation exceeds the preset time deviation threshold and / or the second time deviation exceeds the preset time deviation threshold, the audio overlap between the syllable audio segment and the candidate audio segment can be determined to be non-coincident, thus failing the preset overlap condition, and the candidate audio segment is not considered as the target syllable audio segment. The preset time deviation threshold can be set based on practical experience, for example, as follows: 100 milliseconds. The above method determines the audio overlap between the syllable audio segment and the candidate audio segment by using the first and second time deviations. This can accurately locate the candidate syllable audio segment, thereby improving the accuracy of the target syllable audio segment, which in turn improves the accuracy of the syllable boundary.
[0112] Specifically, the preset syllables corresponding to the candidate syllable audio segments are the preset syllables in the syllable detection results corresponding to syllable audio segments whose audio overlap meets the preset overlap condition. This allows for the acquisition of corresponding confidence levels. Based on the confidence levels of the preset syllables corresponding to each candidate syllable audio segment, target syllable audio segments that meet a preset confidence threshold are selected from the candidate syllable audio segments. These target syllable audio segments can be candidate syllable audio segments whose corresponding confidence levels exceed the preset confidence threshold. The preset confidence threshold can be set based on practical experience, for example, it can be set to 0.9.
[0113] Step (5): Based on the preset syllables corresponding to the target syllable audio segment and the target EEG signal segment in the EEG signal corresponding to the target syllable audio segment, a training corpus is generated; the training corpus is used to train the EEG decoding model of the target object.
[0114] Specifically, the time boundary (start time and end time) of the target syllable audio segment can be obtained. Based on this time boundary, a segment of EEG signal that is time-aligned with it can be determined from the EEG signal as the target EEG signal segment corresponding to the target syllable audio segment. Then, training data in the EEG-syllable corpus of the target object can be generated based on the preset syllable corresponding to the target syllable and the target EEG signal segment.
[0115] Specifically, the training corpus can include the correspondence between target EEG signal segments and corresponding preset syllables, where the preset syllables serve as label information for the target EEG signal segment. In practice, the target EEG signal segment can be extracted from the EEG signal, and label information for that segment can be generated based on its corresponding preset syllables to obtain the first training corpus. For example, the generated label information could be (syl, 1), where syl represents the preset syllable, and "1" indicates that the target EEG signal segment is a speaking EEG signal. In practical applications, a second training corpus can also be generated based on the remaining EEG signal segments after extraction. This second training corpus includes the remaining EEG signal segments and label information indicating that the remaining EEG signal segments are non-speaking EEG signals, such as "0". Both the first and second training corpora are stored in the target object's EEG-syllable corpus, thereby improving the completeness of the training corpus in the target object's EEG-syllable corpus.
[0116] In other examples, data segmentation of EEG signals may not be performed when generating training corpora. In this case, after determining the target EEG signal segment aligned with the target syllable audio segment in the synchronized EEG signal based on the time boundary (start time and end time) of the target syllable audio segment, label information of the corresponding target EEG signal segment is generated based on the preset syllable corresponding to the target syllable audio segment. For example, the label information can be (syl,1), where syl represents the preset syllable and "1" indicates that the target EEG signal segment is an EEG signal in the speaking state. The label information is marked on the target EEG signal segment of the EEG signal. Label information indicating that the EEG signal segment is an EEG signal in the non-speaking state, such as "0", can also be marked on the EEG signal segments other than the target EEG signal segment. Thus, EEG signals labeled with the above label information are obtained, and EEG signals labeled with the above label information are used as training corpora in the EEG-syllable corpus. It should be noted that, in the embodiments of this application, there are no specific restrictions on the presentation of the correspondence between the target EEG signal segments and preset syllables in the training corpus. One or more presentation formats can be adopted based on the actual model training needs.
[0117] It should be noted that the EEG signals with labeled information in the EEG-syllable corpus of the target object in this application embodiment can also be referred to as sample EEG signals.
[0118] In some possible implementations, in order to further improve the quality of the training corpus, the aforementioned acquisition of the target object's audio signal and EEG signal may include: acquiring the original audio signal and synchronized original EEG signal output by the target object based on the training text; performing noise reduction processing on the original audio signal to obtain the audio signal; and performing EEG preprocessing on the original EEG signal to obtain the EEG signal.
[0119] Specifically, noise reduction processing of the original audio signal can remove low-frequency signals. For example, a high-pass filter can be used to filter out low-frequency signals below 100Hz, so that subsequent steps are all based on the noise-reduced audio signal.
[0120] EEG preprocessing of the raw EEG signal may include the following operations performed in sequence: (1) Removal of bad channels: Delete electrode data with excessive or insufficient impedance. For example, if the channel impedance is greater than 1M (no connection) or less than 1k (short circuit), the entire channel data is excluded; (2) Power frequency filtering: Remove power frequency interference. Use notch filtering to remove 50 Hz (or 60 Hz) and its harmonic components to suppress the pollution of the neural signal spectrum by power supply noise. In practical applications, if the system is battery powered and there is no obvious power frequency interference, this step is optional; (3) Common mode re-reference: For each time point, use the average (or median) of all channel signals as the reference signal and subtract it from each channel to suppress spatially shared noise components and enhance the distinguishability of local neural activities; (4) Bandpass filtering: Only retain the frequency bands related to speech (such as ECoG). (2) High-frequency γ band), to remove low-frequency drift and high-frequency noise, and focus on neural activities with speech meaning, for example, retaining the frequency band range related to speech as 70~170Hz; (5) Hilbert transform: apply Hilbert transform to the bandpass filtered signal to obtain the analytical signal, and calculate its instantaneous envelope to characterize the change of neural activity intensity over time, thereby obtaining the EEG signal. By performing the above EEG preprocessing on the original EEG signal, the quality of the EEG signal in the training corpus can be significantly improved, thereby improving the quality of the training corpus.
[0121] After constructing the EEG-syllable corpus of the target object, the articulation boundary model and EEG decoding model of the target object can be trained using the EEG-syllable corpus. Then, the articulation boundary model and EEG decoding model are used to process the EEG signal of the target object to obtain the EEG decoding result. Based on the EEG decoding result, the language decoding of the embodiment of this application is performed to obtain the language content to be output.
[0122] The following describes the process of training the articulation boundary model and EEG decoding model of the target object using the target object's EEG-syllable corpus.
[0123] The training corpus of the target object's EEG-syllable corpus includes sample EEG signals of the target object and label information corresponding to the sample EEG signals. The label information may include a first label, which indicates the target syllable corresponding to the corresponding sample EEG signal among multiple preset syllables. In some embodiments, training an EEG decoding model based on the target object's EEG-syllable corpus may include the following steps: (1) obtaining a training corpus set, wherein the training corpus set includes sample EEG signals of the target object and label information corresponding to the sample EEG signals. The label information includes a first label, which indicates the target syllable corresponding to the sample EEG signal among multiple preset syllables; (2) training a first neural network model for an EEG decoding task based on the sample EEG signals and the corresponding first label to obtain the target object's EEG decoding model; wherein the EEG decoding task is to predict the probability that the corresponding sample EEG signal belongs to each of the preset syllables.
[0124] Specifically, multiple training corpora can be obtained from the target object's EEG-syllable corpus to obtain a training corpus set. For each training corpus, the sample EEG signals are input into the first neural network model for EEG decoding processing to obtain the predicted EEG decoding result corresponding to the sample EEG signal. The predicted EEG decoding result includes the probability that the sample EEG signal belongs to each preset syllable. Based on the difference between the predicted EEG decoding result corresponding to each sample EEG signal and the first label, the model parameters of the first neural network model are adjusted, and the training continues iteratively based on the adjusted model parameters until the training termination condition is met, thus obtaining the trained EEG decoding model.
[0125] For example, adjusting the model parameters of the first neural network model based on the difference between the predicted EEG decoding result and the first label corresponding to each sample EEG signal can be achieved by: using a preset loss function and calculating the EEG decoding loss based on the difference between the predicted EEG decoding result and the first label corresponding to each sample EEG signal, and then using this EEG decoding loss to adjust the model parameters of the first neural network model. The preset loss function may include a cross-entropy loss function. When adjusting the model parameters of the first neural network model using this EEG decoding loss, the model parameters can be adjusted in the direction of reducing the EEG decoding loss for training. For example, when training the first neural network model for an EEG decoding task, a backpropagation algorithm (such as stochastic gradient descent (SGD)) can be used to adjust the model parameters in the direction of reducing the EEG decoding loss. The training termination condition can be that the number of iterations reaches an iteration threshold (which can be set as needed), the EEG decoding loss is less than a loss threshold (which can be set as needed), or the difference between two adjacent EEG decoding losses reaches a difference threshold (which can be set as needed). The first neural network model can be a recurrent neural network, such as a Long Short-Term Memory (LSTM) network.
[0126] The above implementation method uses training data from the target object's EEG-syllable corpus to train the EEG decoding model, thereby achieving personalized modeling for the target object. This can adapt to the differences in neurophysiology among different individuals and improve the accuracy and stability of the EEG decoding model.
[0127] In some implementations, the label information corresponding to the sample EEG signal further includes a second label, which indicates whether the sample EEG signal is an EEG signal in a speaking state. For example, the second label can use "0" to indicate an EEG signal in a non-speaking state and "1" to indicate an EEG signal in a speaking state. Then, training the articulation boundary model based on the EEG-syllable corpus of the target object can include: training a second neural network model on an effective EEG recognition task based on the sample EEG signal and the corresponding second label to obtain the articulation boundary model of the target object; the effective EEG recognition task is to predict whether the corresponding sample EEG information is an EEG signal in a speaking state.
[0128] Specifically, for each training corpus, the sample EEG signals are input into the second neural network model for effective EEG recognition processing to obtain the predicted effective EEG recognition result corresponding to the sample EEG signal. The predicted effective EEG recognition result includes the probability that the sample EEG signal is an EEG signal in a speaking state. Based on the difference between the predicted effective EEG recognition result and the second label corresponding to each sample EEG signal, the model parameters of the second neural network model are adjusted, and the training continues iteratively based on the adjusted model parameters until the training termination condition is met, thus obtaining the trained articulation boundary model.
[0129] For example, adjusting the model parameters of the second neural network model based on the difference between the predicted effective EEG recognition results and the second label corresponding to each sample EEG signal can be achieved by: using a preset loss function and calculating the recognition loss based on the difference between the predicted effective EEG recognition results and the second label corresponding to each sample EEG signal, and then using this recognition loss to adjust the model parameters of the second neural network model. The preset loss function may include a cross-entropy loss function. When adjusting the model parameters of the second neural network model using this recognition loss, the model parameters can be adjusted in the direction of reducing the recognition loss for training. For example, when training the second neural network model for effective EEG recognition, a backpropagation algorithm (such as stochastic gradient descent (SGD)) can be used to adjust the model parameters in the direction of reducing the recognition loss. The training termination condition can be that the number of iterations reaches an iteration threshold (which can be set as needed), the recognition loss is less than a loss threshold (which can be set as needed), or the difference between two adjacent recognition losses reaches a difference threshold (which can be set as needed). The second neural network model can be a deep neural network based on linear discriminant analysis.
[0130] The above implementation method uses training corpus from the target object's EEG-syllable corpus to train the articulation boundary model, thereby achieving personalized modeling for the target object. It can adapt to the differences in articulation habits of different individuals and improve the accuracy and stability of the articulation boundary model.
[0131] In some possible implementations, in order to improve the generalization ability of the articulation boundary model and the EEG decoding model under different noise and individual differences, after acquiring the training corpus, the method may further include: performing data augmentation processing on the sample EEG signals in the training corpus; the data augmentation processing includes one or more of the following combinations: adding random noise, stretching or compressing in the time dimension, perturbing in the frequency dimension, and randomly occluding some features.
[0132] Specifically, adding random noise to the sample EEG signals can simulate unavoidable physiological artifacts (such as slight electromyography and electrocardiography) and environmental noise during the acquisition process. For example, Gaussian white noise can be added.
[0133] The sample EEG signal can be stretched or compressed in the time dimension to perform nonlinear deformation in the time dimension, simulating the natural differences in speech rate among different individuals and the fluctuation of pronunciation duration of the same object in different states. For example, the time axis of the sample EEG signal can be stretched or compressed uniformly within a predetermined scaling range (such as [0.9, 1.1]). Of course, different degrees of stretching and compression can also be applied to different local time periods of the sample EEG signal to simulate the non-uniform changes in syllable duration in speech flow.
[0134] Perturbation of the sample EEG signal in the frequency dimension can simulate the changes in frequency response characteristics caused by individual physiological differences, slight changes in electrode impedance, or fluctuations in brain state. For example, the spectrum of the sample EEG signal can be slightly shifted along the frequency axis, or its amplitude can be slightly and randomly scaled.
[0135] Randomly occluding certain features of the sampled EEG signals can enhance the model's robustness to partial information loss or attentional distraction, simulating transient signal loss or attention shifts that may occur in real-world applications. Examples include temporal random occlusion and channel random occlusion. Temporal random occlusion involves randomly selecting one or more consecutive time segments on the time axis and setting the values of all channels or features within that time segment to zero or the average value of that channel. Channel random occlusion involves randomly selecting a portion of the acquisition channels (e.g., randomly blocking 10%-30% of the electrode channels) and setting their data to zero at all time steps to simulate scenarios of poor electrode contact or partial channel failure.
[0136] It should be noted that the label information corresponding to the enhanced sample EEG signal obtained after data augmentation is consistent with the label information of the sample EEG signal before data augmentation.
[0137] For example, to maximize the data augmentation effect, random noise can be added to a sample EEG signal sequentially, stretching or compressing it in the time dimension, perturbing it in the frequency dimension, and randomly occluding some features.
[0138] The above implementation methods can simulate the diversity and uncertainty of EEG signals in real-world scenarios, allowing the model to "see" richer and more complex EEG signals during training. This forces the model to learn deep features that remain stable even under noise, temporal distortion, frequency perturbations, and partial information loss. Consequently, the trained articulation boundary model can more reliably identify valid speech state EEG segments under various interferences, while the EEG decoding model can decode the enhanced, more diverse, and valid EEG signals. This significantly improves the generalization ability and overall decoding accuracy in complex environments during actual deployment.
[0139] The following describes the process of using the articulation boundary model and EEG decoding model of the target object obtained from the above training to process the EEG signal of the target object to obtain the EEG decoding result. The EEG signal processing process may include the following steps (1) to (3): Step (1): Obtain the EEG signal to be processed from the target object.
[0140] In some implementations, to improve the accuracy of language content decoding based on EEG signals, the process of acquiring the EEG signal to be processed of the target object may include: acquiring the original EEG signal of the target object; performing EEG preprocessing on the original EEG signal to obtain the EEG signal to be processed of the target object; the EEG preprocessing includes extracting the EEG time-domain envelope of a preset speech-related frequency band.
[0141] Specifically, raw EEG signals can be information on changes in brain neuron activity over time that can be directly acquired using a non-invasive brain-computer interface.
[0142] The preset speech-related frequency band refers to the range of frequencies in the original EEG signal that are related to speech. In this embodiment, the preset speech-related frequency band is 70~170Hz to remove low-frequency drift and high-frequency noise, focusing on neural activity with speech significance. The EEG time-domain envelope of the preset speech-related frequency band is used to characterize the change in the intensity of neural activity within this frequency band over time. During EEG preprocessing, Hilbert transform can be used to extract the EEG time-domain envelope of the preset speech-related frequency band.
[0143] In some examples, the preset speech association frequency band can be set according to the individual neural response characteristics of the target object, so that the EEG preprocessing can be adapted to the neurophysiological differences of different people, thereby improving the individual adaptability of the EEG signal to be processed, which is conducive to improving the accuracy of the final decoded language content.
[0144] In specific implementation, the EEG preprocessing of the raw EEG signal may include the following operations performed sequentially: (a) Removal of bad channels: Deleting electrode data with excessively high or low impedance. For example, if the channel impedance is greater than 1MΩ (no connection) or less than 1kΩ (short circuit), the entire channel data is excluded; (b) Power frequency filtering: Removing power frequency interference, using a notch filter to remove 50 Hz (or 60 Hz) interference. (c) Common-mode re-reference: For each time point, the average (or median) of all channel signals is used as the reference signal and subtracted from each channel to suppress spatially shared noise components and enhance the discriminability of local neural activity; (d) Bandpass filtering: Only the frequency band range related to speech, such as 70~170Hz, is retained to remove low-frequency drift and high-frequency noise and focus on neural activity with speech significance; (e) Hilbert transform: The Hilbert transform is applied to the bandpass filtered signal to obtain the analytic signal, and its instantaneous envelope is calculated to obtain the EEG signal to be processed to characterize the change of neural activity intensity over time.
[0145] By selecting the speech-related frequency band and extracting the EEG temporal envelope, the neural activity representations most relevant to the language temporal structure are initially extracted from the raw EEG signals, thereby effectively improving the efficiency of subsequent processing steps and the accuracy of the final language content decoding.
[0146] Step (2): Call the articulation boundary model of the target object to perform effective EEG recognition processing on the EEG signal to be processed, and obtain effective EEG recognition results.
[0147] The effective EEG recognition result indicates the effective EEG signal in the EEG signal to be processed, and the effective EEG signal is the EEG signal in the speaking state. For example, the effective EEG recognition result can be represented as ( , ),in, The start time of an effective EEG signal. The end time of an effective EEG signal.
[0148] Step (3): Call the EEG decoding model of the target object, perform EEG decoding processing based on the effective EEG recognition results, and obtain the EEG decoding results.
[0149] The EEG decoding result indicates the probability that the valid EEG signal belongs to each of the multiple preset syllables. Specifically, the EEG decoding result may include multiple preset syllables and the EEG decoding score of each preset syllable. The EEG decoding score indicates the probability of the corresponding preset syllable, which can also be referred to as the confidence level.
[0150] In some implementations, when calling the EEG decoding model of the target object and performing EEG decoding processing based on the effective EEG recognition results to obtain the EEG decoding results, the process may include: segmenting the EEG signal to be processed based on the effective EEG recognition results to obtain the effective EEG signal; normalizing the effective EEG signal to obtain effective EEG features; and inputting the effective EEG features into the EEG decoding model of the target object for EEG decoding processing to obtain the EEG decoding results.
[0151] Specifically, it can be based on the ( ) in the valid EEG recognition results , The EEG signal to be processed is segmented to extract the starting time. The end time is A segment of brainwave signal was used to obtain an effective brainwave signal.
[0152] The normalization process can be a standardization process performed on each channel or feature dimension, such as Z-score standardization, to make the mean of the data zero and the variance one, so as to obtain effective EEG features. Then, the effective EEG features are input into the EEG decoding model of the target object for EEG decoding processing to obtain the EEG decoding result.
[0153] By utilizing the articulation boundary model of the target object, background neural activity, resting-state signals, and most artifact interference unrelated to language decoding are effectively stripped away. This provides high signal-to-noise ratio and highly correlated effective EEG signals for subsequent EEG decoding steps. This allows the EEG decoding model to perform EEG decoding on effective EEG signals, enabling it to learn more focused neural representations and spatiotemporal patterns directly related to language production. This significantly reduces computational load and the possibility of confusion, thereby effectively improving the accuracy and reliability of syllable probability determination. Consequently, it greatly improves the accuracy and reliability of language content based on EEG signal decoding.
[0154] After obtaining the above EEG decoding results of the target subject, such as Figure 5 As shown, language decoding can be performed based on the language decoding method provided in this application embodiment, namely steps S201 to S209 of this application embodiment, to determine the language content to be output at the current time step, which will not be described in detail here.
[0155] This application also provides a language decoding device. Since the language decoding device provided in this application corresponds to the language decoding methods provided in the above-mentioned embodiments, the implementation methods of the aforementioned language decoding methods are also applicable to the language decoding device provided in this embodiment, and will not be described in detail in this embodiment.
[0156] Please see Figure 6The diagram shown is a structural schematic of a language decoding device provided in an embodiment of this application. This device has the function of implementing the language decoding method described in the above-described method embodiments. This function can be implemented in hardware or by hardware executing corresponding software. Figure 6 As shown, the language decoding device 600 may include: The EEG decoding result acquisition unit 610 is used to acquire the EEG decoding result at the current time step; the EEG decoding result is obtained based on the EEG signal to be processed of the target object at the current time step, and the EEG decoding result includes multiple preset syllables and the EEG decoding score corresponding to each preset syllable. The sentence expansion unit 620 is used to expand the historical candidate sentences of the previous time step for each preset syllable based on the candidate characters associated with the preset syllable, so as to obtain the expanded sentence set corresponding to each preset syllable; The integrated language decoding score determination unit 630 is used to determine the integrated language decoding score of each extended sentence corresponding to the preset syllable based on the historical integrated language decoding score of the previous time step, the first linguistic probability score corresponding to the extended sentence, and the EEG decoding score corresponding to the preset syllable; the first linguistic probability score represents the co-occurrence probability of the historical candidate sentences and candidate characters used for expansion. The candidate sentence determination unit 640 is used to determine the candidate sentence sequence at the current time step based on the comprehensive language decoding score of the extended sentences corresponding to each preset syllable. The language content to be output unit 650 is used to perform sentence end detection based on the candidate sentence sequence at the current time step, and determine the language content to be output at the current time step based on the detection result of the sentence end detection.
[0157] In some implementations, the integrated language decoding score determination unit 630 includes: The weight coefficient acquisition unit is used to acquire the first weight coefficient, the second weight coefficient, and the third weight coefficient; The weighted summation unit is used to perform a weighted summation of the historical comprehensive language decoding score of the previous time step, the first linguistic probability score corresponding to the extended sentence, and the EEG decoding score corresponding to the preset syllable based on the first weight coefficient, the second weight coefficient, and the third weight coefficient, to obtain the comprehensive language decoding score of the extended sentence.
[0158] In some implementations, the weight coefficient acquisition unit is specifically used to: determine the first weight coefficient, the second weight coefficient, and the third weight coefficient based on the current total number of characters in the expanded sentence; wherein the first weight coefficient and the second weight coefficient are both positively correlated with the current total number of characters, and the third weight coefficient is negatively correlated with the current total number of characters.
[0159] In some implementations, the output language content determination unit 650, when performing sentence end detection based on the candidate sentence sequence at the current time step, specifically performs the following: for each candidate sentence at the current time step, based on the historical comprehensive language decoding score and the second linguistic probability score of the previous time step, determines the sentence completeness score corresponding to the candidate sentence, where the second linguistic probability score represents the linguistic probability score of the historical candidate sentence corresponding to the candidate sentence as a complete sentence; based on the third linguistic probability score and the EEG decoding score of the corresponding preset syllable, determines the language decoding score when the preset syllable is the beginning of a sentence, where the third linguistic probability score represents the linguistic probability score of the candidate character of the preset syllable corresponding to the candidate sentence as the beginning of a sentence at the current time step; based on the sentence completeness score of the candidate sentence and the language decoding score when the corresponding preset syllable is the beginning of a sentence, determines the detection score corresponding to the candidate sentence; and based on the detection scores and comprehensive language decoding scores corresponding to each candidate sentence, determines the detection result of the sentence end detection.
[0160] In some implementations, the output language content determination unit 650, when determining the detection result of sentence end detection based on the detection score and comprehensive language decoding score corresponding to each candidate sentence, is specifically configured to: determine the highest score based on the detection score and comprehensive language decoding score corresponding to each candidate sentence; if the highest score is any of the detection scores, then determine that the detection result of sentence end detection is that the sentence is complete; if the highest score is any of the comprehensive language decoding scores, then determine that the detection result of sentence end detection is that the sentence is incomplete.
[0161] In some implementations, when the output language content determination unit 650 determines the output language content of the current time step based on the detection result of the sentence end detection, it is specifically used to: if the detection result of the sentence end detection is that the sentence is not completed, then take the candidate sentence corresponding to the highest score as the output language content of the current time step; if the detection result of the sentence end detection is that the sentence is completed, then determine the output language content of the current time step based on the language decoding score when each preset syllable is the beginning of the sentence, and update the candidate sentence sequence of the current time step.
[0162] In some implementations, the candidate sentence determination unit 640 is specifically used to: select a preset number of extended sentences with the highest integrated language decoding scores as the candidate sentence sequence for the current time step.
[0163] It should be noted that the apparatus provided in the above embodiments is only illustrated by the division of the above functional modules when implementing its functions. In actual applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above. In addition, the apparatus and method embodiments provided in the above embodiments belong to the same concept, and the specific implementation process can be found in the method embodiments, which will not be repeated here.
[0164] This application provides an electronic device including a processor and a memory. The memory stores at least one instruction or at least one program, which is loaded and executed by the processor to implement any of the language decoding methods provided in the above method embodiments.
[0165] Furthermore, Figure 7 A schematic diagram of the hardware structure of an electronic device for implementing a language decoding method provided in an embodiment of this application is shown. The electronic device may participate in or include the language decoding apparatus provided in the embodiment of this application. Figure 7 As shown, the electronic device 70 may include one or more (shown as 702a, 702b, ..., 702n) processors 702 (processors 702 may include, but are not limited to, microprocessors such as MCUs or programmable logic devices such as FPGAs), a memory 704 for storing data, and a transmission device 706 for communication functions. In addition, it may also include: a display, an input / output interface (I / O interface), a universal serial bus (USB) port (which may be included as one of the ports of the I / O interface), a network interface, a power supply, and / or a camera. Those skilled in the art will understand that... Figure 7 The structure shown is for illustrative purposes only and does not limit the structure of the electronic device described above. For example, the electronic device 70 may also include... Figure 7 The more or fewer components shown, or having the same Figure 7 The different configurations shown.
[0166] It should be noted that the aforementioned one or more processors 702 and / or other data processing circuits are generally referred to herein as "data processing circuits". These data processing circuits may be embodied, in whole or in part, in software, hardware, firmware, or any other combination thereof. Furthermore, the data processing circuits may be a single, independent processing module, or may be wholly or partially integrated into any other element within the electronic device 70 (or mobile device). As involved in the embodiments of this application, the data processing circuits serve as a processor control mechanism (e.g., selection of a variable resistor termination path connected to an interface).
[0167] The memory 704 can be used to store software programs and modules of application software, such as the program instructions / data storage device corresponding to the method described in the embodiments of this application. The processor 702 executes various functional applications and data processing by running the software programs and modules stored in the memory 704, thereby realizing the above-described training corpus generation method. The memory 704 may include high-speed random access memory, and may also include non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some instances, the memory 704 may further include memory remotely located relative to the processor 702, and these remote memories can be connected to the electronic device 70 via a network. Examples of the above-mentioned networks include, but are not limited to, the Internet, corporate intranets, local area networks, mobile communication networks, and combinations thereof.
[0168] The transmission device 706 is used to receive or send data via a network. Specific examples of the network described above may include a wireless network provided by the communication provider of the electronic device 70. In one example, the transmission device 706 includes a network interface controller (NIC), which can connect to other network devices via a base station to communicate with the Internet. In one embodiment, the transmission device 706 may be a radio frequency (RF) module for wireless communication with the Internet.
[0169] The display can be, for example, a touchscreen liquid crystal display (LCD) that allows a user to interact with the user interface of the electronic device 70 (or mobile device).
[0170] Embodiments of this application also provide a computer-readable storage medium, which can be disposed in an electronic device to store at least one instruction or at least one program related to implementing a language decoding method. The at least one instruction or the at least one program is loaded and executed by the processor to implement any of the language decoding methods provided in the above-described method embodiments.
[0171] This application also provides a computer program product, which includes a computer program stored in a computer-readable storage medium. An electronic device's processor reads the computer program from the computer-readable storage medium and executes the computer program, causing the electronic device to perform any of the language decoding methods provided in the above-described method embodiments.
[0172] Optionally, in this embodiment, the storage medium may include, but is not limited to, various media capable of storing program code, such as USB flash drives, read-only memory (ROM), random access memory (RAM), portable hard drives, magnetic disks, or optical disks.
[0173] It should be noted that the order of the embodiments described above is merely for descriptive purposes and does not represent the superiority or inferiority of the embodiments. Furthermore, specific embodiments have been described above. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps described in the claims can be performed in a different order than that shown in the embodiments and still achieve the desired result. Additionally, the processes depicted in the drawings do not necessarily require a specific or sequential order to achieve the desired result. In some embodiments, multitasking and parallel processing are also possible or may be advantageous.
[0174] The various embodiments in this specification are described in a progressive manner. Similar or identical parts between embodiments can be referred to mutually. Each embodiment focuses on describing the differences from other embodiments. In particular, the apparatus embodiments are basically similar to the method embodiments, so the description is relatively simple; relevant parts can be referred to the descriptions of the method embodiments.
[0175] Those skilled in the art will understand that all or part of the steps of the above embodiments can be implemented by hardware or by a program instructing related hardware. The program can be stored in a computer-readable storage medium, such as a read-only memory, a disk, or an optical disk.
[0176] The above description is only a preferred embodiment of this application and is not intended to limit this application. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the protection scope of this application.
Claims
1. A language decoding method, characterized in that, The method includes: Obtain the EEG decoding result at the current time step; the EEG decoding result is obtained based on the EEG signal to be processed of the target object at the current time step, and the EEG decoding result includes multiple preset syllables and the EEG decoding score corresponding to each preset syllable; For each preset syllable, based on the candidate characters associated with the preset syllable, the historical candidate sentences of the previous time step are expanded to obtain the set of expanded sentences corresponding to the preset syllable; For each extended sentence corresponding to the preset syllable, the comprehensive language decoding score of the extended sentence is determined based on the historical comprehensive language decoding score of the previous time step, the first linguistic probability score corresponding to the extended sentence, and the EEG decoding score corresponding to the preset syllable; the first linguistic probability score represents the co-occurrence probability of the historical candidate sentences and candidate characters used for extension. Based on the comprehensive language decoding score of the extended sentences corresponding to each preset syllable, the candidate sentence sequence for the current time step is determined; For each candidate sentence at the current time step, a sentence completeness score is determined based on the historical comprehensive language decoding score and the second linguistic probability score from the previous time step. The second linguistic probability score represents the linguistic probability score of the historical candidate sentence corresponding to the candidate sentence as a complete sentence. A language decoding score is determined based on the third linguistic probability score and the EEG decoding score of the preset syllable corresponding to the candidate sentence when the preset syllable is the beginning of a sentence. The third linguistic probability score represents the linguistic probability score of the candidate character corresponding to the preset syllable of the candidate sentence as the beginning of a sentence at the current time step. A detection score is determined based on the sentence completeness score of the candidate sentence and the language decoding score when the corresponding preset syllable is the beginning of a sentence. A detection result for sentence end detection is determined based on the detection score and comprehensive language decoding score of each candidate sentence, and the language content to be output at the current time step is determined based on the detection result of sentence end detection.
2. The method according to claim 1, characterized in that, The determination of the comprehensive language decoding score of the extended sentence based on the historical comprehensive language decoding score of the previous time step, the first linguistic probability score corresponding to the extended sentence, and the EEG decoding score corresponding to the preset syllable includes: Obtain the first weight coefficient, the second weight coefficient, and the third weight coefficient; Based on the first weighting coefficient, the second weighting coefficient, and the third weighting coefficient, the historical comprehensive language decoding score of the previous time step, the first linguistic probability score corresponding to the extended sentence, and the EEG decoding score corresponding to the preset syllable are weighted and summed to obtain the comprehensive language decoding score of the extended sentence.
3. The method according to claim 2, characterized in that, The process of obtaining the first weighting coefficient, the second weighting coefficient, and the third weighting coefficient includes: Based on the current total number of characters in the expanded sentence, the first weight coefficient, the second weight coefficient, and the third weight coefficient are determined; The first weighting coefficient and the second weighting coefficient are both positively correlated with the current total number of words, while the third weighting coefficient is negatively correlated with the current total number of words.
4. The method according to claim 1, characterized in that, Based on the detection scores and comprehensive language decoding scores corresponding to each candidate sentence, the detection results for sentence end detection include: The highest score is determined based on the detection score and comprehensive language decoding score corresponding to each candidate sentence; If the highest score is any of the detection scores, then the detection result of the sentence end detection is determined to be that the sentence is complete; If the highest score is any of the comprehensive language decoding scores, then the detection result of the sentence end detection is determined to be that the sentence is incomplete.
5. The method according to claim 4, characterized in that, The step of determining the language content to be output at the current time step based on the detection result of the sentence end detection includes: If the sentence end detection result is that the sentence is incomplete, then the candidate sentence corresponding to the highest score is taken as the language content to be output at the current time step. If the sentence end detection result indicates that the sentence is complete, then the language content to be output at the current time step is determined based on the language decoding score when each preset syllable is the beginning of the sentence, and the candidate sentence sequence at the current time step is updated.
6. The method according to claim 1, characterized in that, The determination of the candidate sentence sequence for the current time step based on the comprehensive language decoding score of the extended sentences corresponding to each preset syllable includes: A predetermined number of extended sentences with the highest comprehensive language decoding scores are selected as the candidate sentence sequence for the current time step.
7. A language decoding device, characterized in that, The device includes: The EEG decoding result acquisition unit is used to acquire the EEG decoding result at the current time step; the EEG decoding result is obtained based on the EEG signal to be processed of the target object at the current time step, and the EEG decoding result includes multiple preset syllables and the EEG decoding score corresponding to each preset syllable; The sentence expansion unit is used to expand the historical candidate sentences of the previous time step for each preset syllable based on the candidate characters associated with the preset syllable, so as to obtain the expanded sentence set corresponding to each preset syllable; The integrated language decoding score determination unit is used to determine the integrated language decoding score of each extended sentence corresponding to the preset syllable based on the historical integrated language decoding score of the previous time step, the first linguistic probability score corresponding to the extended sentence, and the EEG decoding score corresponding to the preset syllable; the first linguistic probability score represents the co-occurrence probability of the historical candidate sentences and candidate characters used for expansion; The candidate sentence determination unit is used to determine the candidate sentence sequence at the current time step based on the comprehensive language decoding score of the extended sentences corresponding to each preset syllable; The output language content determination unit is used to determine the sentence completeness score corresponding to each candidate sentence at the current time step based on the historical comprehensive language decoding score and the second linguistic probability score of the previous time step; the second linguistic probability score represents the linguistic probability score of the historical candidate sentence corresponding to the candidate sentence as a complete sentence; based on the third linguistic probability score and the EEG decoding score of the preset syllable corresponding to the candidate sentence, the unit determines the language decoding score when the preset syllable is the beginning of a sentence; the third linguistic probability score represents the linguistic probability score of the candidate character corresponding to the preset syllable of the candidate sentence as the beginning of a sentence at the current time step; based on the sentence completeness score of the candidate sentence and the language decoding score when the corresponding preset syllable is the beginning of a sentence, the unit determines the detection score corresponding to the candidate sentence; based on the detection scores and comprehensive language decoding scores corresponding to each candidate sentence, the unit determines the detection result of sentence end detection, and determines the output language content at the current time step based on the detection result of sentence end detection.
8. An electronic device, characterized in that, It includes a processor and a memory, wherein the memory stores at least one instruction or at least one program, the at least one instruction or the at least one program being loaded and executed by the processor to implement the language decoding method as described in any one of claims 1 to 6.
9. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores at least one instruction or at least one program, which is loaded and executed by a processor to implement the language decoding method as described in any one of claims 1 to 6.