Language function rehabilitation corpus screening method based on semantic embedding and neural coding
By constructing an encoding model using pre-trained semantic matching models and functional magnetic resonance imaging data, the problems of insufficient semantic coverage and high redundancy in existing rehabilitation corpus screening methods are solved, achieving personalized and efficient corpus screening and activation effects.
Patent Information
- Application Number
- CN202511371878.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-24
- Publication Date
- 2026-01-16
AI Technical Summary
Existing methods for screening rehabilitation corpora rely on human experience or traditional vocabulary lists, resulting in insufficient semantic coverage, high redundancy, difficulty in broadly activating the brain's language network, and a lack of individualization and scientific rigor.
By using a pre-trained semantic matching model for semantic embedding and clustering analysis, and combining functional magnetic resonance imaging data, a coding model is constructed from sentence semantic vectors to brain regions of interest related to language processing, thereby screening out rehabilitation corpora with high activation responses.
It achieves a corpus with a reasonable semantic structure, wide coverage, and controlled redundancy, and individually identifies brain regions related to language processing, reducing reliance on brain imaging and improving the efficiency and accuracy of corpus selection.
Smart Images

Figure CN121350233A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of semantic embedding and neural coding, and particularly relates to a language function rehabilitation corpus screening method based on semantic embedding and neural coding. BACKGROUND
[0002] With the development of neurolinguistics and brain science, brain imaging technologies such as functional magnetic resonance imaging (fMRI) are widely used to study human language processing mechanisms, and are gradually applied to the rehabilitation training of patients with language dysfunction. In actual rehabilitation process, the selection of stimulating corpus directly affects the activation effect of the brain language-related brain area of the patient, thereby affecting the efficiency and quality of the rehabilitation training.
[0003] The existing rehabilitation corpus screening methods mainly have two types: one type relies on artificial experience or traditional vocabulary table for selection, and the semantic coverage is insufficient, the content is single, and there is a high redundancy, so it is difficult to ensure that the selected corpus can activate the brain language network widely; the other type tries to use a specific type of corpus (such as action-related sentences) to activate part of the brain area, but the activation range is often limited to the motor cortex, and more brain areas related to language processing cannot be effectively covered. These deficiencies result in limited neural effectiveness of the rehabilitation corpus, and the rehabilitation process lacks individualization and scientificity.
[0004] Therefore, a new corpus screening method is needed, which can break through the limitations of artificial experience selection, systematically construct a corpus with diverse semantics and reasonable structure, and combine the brain imaging data of the subject, to establish the mapping relationship between semantic features and brain activation, so as to scientifically screen the core corpus that can produce high activation response in the language-related brain area. SUMMARY
[0005] The present disclosure provides a language function rehabilitation corpus screening method based on semantic embedding and neural coding.
[0006] In a first aspect, the present disclosure provides a language function rehabilitation corpus screening method based on semantic embedding and neural coding, comprising:
[0007] performing semantic embedding processing on each vocabulary in the vocabulary library based on a pre-trained semantic matching model to obtain a vocabulary semantic vector;
[0008] performing clustering analysis on the vocabulary semantic vectors to obtain a plurality of semantic categories;
[0009] screening natural language sentences corresponding to each of the semantic categories from a large-scale corpus library to form a candidate stimulating corpus;
[0010] acquiring functional magnetic resonance imaging data of a subject based on the candidate stimulating corpus, and constructing an encoding model from a sentence semantic vector to a language processing-related brain region of interest.
[0011] inputting the sentence semantic vectors in the candidate stimulus corpus into the encoding model to obtain predicted activation values of the interest region, and constructing a rehabilitation training corpus set according to the predicted activation values.
[0012] Optionally, the screening of natural language sentences corresponding to each semantic category from the large-scale corpus to form a candidate stimulus corpus further comprises:
[0013] obtaining a large-scale corpus, wherein the large-scale corpus comprises a plurality of sub-corpus sets, and each sub-corpus set comprises a plurality of sentences;
[0014] processing semantic embedding of each sentence in the large-scale corpus by using a pre-trained semantic matching model to obtain a semantic vector of each sentence;
[0015] obtaining similarity of each sentence to each semantic category by cosine similarity calculation;
[0016] for each sub-corpus set, screening out a sentence with the highest similarity to each semantic category to form a candidate stimulus corpus.
[0017] Optionally, the acquisition of functional magnetic resonance imaging data of a subject based on the candidate stimulus corpus to construct an encoding model from a sentence semantic vector to an interest region of a language processing related brain area further comprises:
[0018] determining the interest region of the language processing related brain area of the subject by comparing functional magnetic resonance imaging data of real sentences and pseudo sentences;
[0019] obtaining BOLD values of the interest region of the language processing related brain area of the subject under each sentence stimulus in the stimulus corpus, and constructing an encoding model from a sentence semantic vector to the interest region of the language processing related brain area by using the BOLD values.
[0020] Optionally, the determination of the interest region of the language processing related brain area of the subject by comparing functional magnetic resonance imaging data of real sentences and pseudo sentences further comprises:
[0021] using real sentences and pseudo sentences as visual stimulus materials to respectively scan the subject to obtain functional magnetic resonance imaging data;
[0022] the pseudo sentence is a sentence with no complete semantics formed by combining random words in a reasonable grammatical structure;
[0023] comparing brain region activation responses under the stimulus conditions of the real sentences and the pseudo sentences to construct a significance statistical map;
[0024] The saliency statistics are spatially overlapped with a standard language area template to obtain an individualized language processing related brain area region of interest of the subject.
[0025] Optionally, the constructing of the encoding model from the sentence semantic vector to the language processing related brain area region of interest by the BOLD value further comprises:
[0026] The sentence semantic vector in the candidate stimulus corpus is taken as an input variable;
[0027] The average BOLD value of the language processing related brain area region of interest under the corresponding sentence stimulus is taken as an output variable;
[0028] A regularized regression algorithm is adopted to establish a mapping relationship between the input variable and the output variable to train the encoding model from the sentence semantic vector to the language processing related brain area region of interest.
[0029] Optionally, the semantic modeling of each vocabulary in the vocabulary library based on the pre-trained semantic matching model to obtain a vocabulary semantic vector further comprises:
[0030] Each vocabulary in the vocabulary library is subjected to semantic embedding processing based on a pre-trained semantic matching model adopting a sentence-transformer architecture to obtain a semantic vector corresponding to each vocabulary.
[0031] Optionally, the clustering analysis of the vocabulary semantic vectors to obtain a plurality of semantic categories further comprises:
[0032] A similarity matrix between the vocabulary semantic vectors is constructed;
[0033] A spectral clustering algorithm is applied to the similarity matrix to divide the vocabularies into different semantic categories.
[0034] Optionally, the inputting of the sentence semantic vector in the candidate stimulus corpus into the encoding model to obtain a predicted activation value of the region of interest, and the constructing of a rehabilitation training corpus set according to the predicted activation value further comprises:
[0035] The predicted activation values of each sentence in the candidate stimulus corpus are sorted, and a target sentence ranking at the top is selected as a rehabilitation training corpus;
[0036] Alternatively, an activation intensity threshold is set to filter out a target sentence with a predicted activation value greater than the threshold to form the rehabilitation training corpus set.
[0037] In a second aspect, the present disclosure provides an electronic device, comprising a processor and a memory connected in communication with the processor.
[0038] The memory stores computer-executable instructions;
[0039] The processor executes the computer-executable instructions stored in the memory to implement the method in the present disclosure.
[0040] In a third aspect, the present disclosure provides a computer-readable storage medium, the computer-readable storage medium storing computer-executable instructions, the computer-executable instructions being executed by a processor to implement the method of the present disclosure.
[0041] The present disclosure has the following advantages compared with the prior art:
[0042] 1) The present application processes semantic embedding of large-scale vocabulary and corpus by pre-training a semantic matching model, and constructs semantic categories combined with clustering analysis, thereby forming a candidate corpus library with reasonable semantic structure, wide coverage and controlled redundancy.
[0043] 2) The present application uses functional magnetic resonance imaging (fMRI) data, combined with real sentence and pseudo sentence comparison experiments, to individually determine the language processing related brain region of interest area (ROI) of the subject, and based on the mapping relationship between semantic vectors and brain region activation intensity, to construct an encoding model. This method can accurately predict the activation effect of different corpora in the target brain region, thereby screening corpora with higher neural activation potential.
[0044] 3) After obtaining the initial fMRI data, the neural activation effect of the candidate corpus can be forward predicted by the semantic encoding model obtained by training, thereby avoiding the high-cost operation of repeatedly carrying out fMRI experiments on each corpus. This mechanism not only retains the scientific basis provided by brain imaging experiments, but also significantly reduces the dependence of the corpus screening process on brain imaging, enabling efficient screening of large-scale corpora while maintaining high prediction accuracy. BRIEF DESCRIPTION OF DRAWINGS
[0045] The accompanying drawings, which are incorporated in and constitute a part of the specification, illustrate embodiments consistent with the present disclosure and serve to explain the principles of the present disclosure together with the specification.
[0046] Figure 1 A language function rehabilitation corpus screening method based on semantic embedding and neural encoding provided by an embodiment of the present disclosure;
[0047] Figure 2 A candidate stimulus corpus library formation method flowchart provided by an embodiment of the present disclosure;
[0048] Figure 3 A coding model for constructing a sentence semantic vector to a language processing related brain region of interest area provided by an embodiment of the present disclosure;
[0049] Figure 4 This is a schematic diagram of a method for determining the brain regions of interest related to language processing in a subject, as provided in an embodiment of this disclosure.
[0050] The accompanying drawings have illustrated specific embodiments of this disclosure, which will be described in more detail below. These drawings and descriptions are not intended to limit the scope of the concept in any way, but rather to illustrate the concepts of this disclosure to those skilled in the art through reference to particular embodiments. Detailed Implementation
[0051] The present disclosure will be further described below with reference to the accompanying drawings. The following embodiments are only used to illustrate the technical solutions of the present disclosure more clearly, and should not be used to limit the scope of protection of the present disclosure.
[0052] Figure 1 This diagram illustrates a method for screening language function rehabilitation corpora based on semantic embedding and neural coding, provided as an embodiment of this disclosure. See also... Figure 1 The following is a detailed discussion of each step in conjunction with this embodiment.
[0053] S100. Based on the pre-trained semantic matching model, semantic embedding processing is performed on each word in the vocabulary to obtain the semantic vector of each word.
[0054] In step S100, a vocabulary for semantic embedding processing needs to be established first. This vocabulary can be derived from a general language lexicon, such as a list of commonly used words in modern Chinese, or it can be selected from specialized vocabulary related to daily life, social interaction, and medical rehabilitation, tailored to the rehabilitation training scenario. In this embodiment, the vocabulary size can be set between 20,000 and 30,000 words to ensure sufficient semantic coverage. For each word in the vocabulary, a pre-trained semantic matching model is used for semantic embedding processing. A preferred approach is to use a pre-trained semantic matching model with a sentence-transformer architecture to perform semantic embedding processing on each word in the vocabulary, obtaining a semantic vector corresponding to each word. For example, the Text2Vec model or a BERT derivative model can map words to a high-dimensional continuous semantic space. During implementation, the vocabulary is first pre-processed, including removing symbols, merging synonyms, and unifying word forms. Then, each word is input into the pre-trained model, outputting a semantic vector with dimensions from 256 to 768. The dimension of each vector corresponds to the feature weight distribution within the model, reflecting the positional relationship of the word in the semantic space. For example, in this embodiment, the cosine similarity between the semantic vector of the word "apple" and the semantic vector of "banana" is close to 1, while the similarity with "hospital" is low. This indicates that the semantic embedding method can effectively characterize the semantic relevance between words. Through the above steps, a set of semantic vectors for all words in the vocabulary is obtained. This set will serve as input data for subsequent semantic clustering analysis, laying the foundation for constructing a stimulus corpus with semantic diversity.
[0055] S200. Perform cluster analysis on the semantic vectors of each word to obtain multiple semantic categories.
[0056] In step S200, cluster analysis is performed on all the word semantic vectors obtained in step S100 to obtain multiple semantic categories. This includes: constructing a similarity matrix between the semantic vectors of each word; and applying a spectral clustering algorithm to the similarity matrix to divide each word into different semantic categories.
[0057] Specifically, before performing clustering analysis, the high-dimensional semantic vectors can be reduced in dimensionality to decrease computational complexity. For example, Principal Component Analysis (PCA) can be used to compress the original 768-dimensional semantic vectors to 50 dimensions, thereby reducing computational costs while preserving the main semantic features. It should be noted that, to preserve complete semantic information, this embodiment does not perform dimensionality reduction but directly uses the original-dimensional semantic vectors for subsequent analysis. In this embodiment, a similarity matrix is constructed based on the cosine similarity between word semantic vectors. The similarity value ranges from 0 to 1, with values closer to 1 indicating stronger semantic association. Subsequently, this similarity matrix is used as input to classify the words using a spectral clustering algorithm. Spectral clustering projects the data into a low-dimensional space by calculating the feature vectors of the similarity matrix and then uses the K-means method to complete the clustering. Specifically, in this embodiment, the number of clusters K is set to between 300 and 500 to ensure strong semantic consistency among words within each category while covering a sufficiently broad semantic range. After clustering, each semantic category can correspond to a group of words with highly related semantics. For example, "doctor, nurse, patient" can be classified into the same category, while "fruit, vegetable, beverage" can be classified into another category.
[0058] To further enhance the scientific rigor and relevance of clustering results to rehabilitation, manual screening and correction can be performed. Specifically, professionals check the rationality of words in each semantic category, remove abnormal words misassigned by the clustering algorithm, and supplement missing keywords based on rehabilitation training scenarios.
[0059] In semantic clustering, the size of different categories often varies greatly: some categories have only a dozen or so words, while others may have hundreds. Traditional clustering algorithms (such as spectral clustering combined with k-means) are easily influenced by the dominance of "large categories," causing the boundaries of "small categories" to be swallowed up, resulting in an imbalance in semantic distribution. Consequently, rehabilitation corpora may end up biased towards common semantics, neglecting key but rare semantic categories (such as rare but necessary contextual terms in medical rehabilitation). Therefore, in another embodiment, this disclosure proposes a category weight correction method based on harmonic mean radius to address the problem of small categories being easily overlooked in traditional clustering. Specifically, after completing spectral clustering, the average radius r of each semantic category is first calculated, which is the average distance from all semantic vectors within the category to the cluster center. Then, based on the harmonic mean formula, the radii of each category are normalized to obtain the category weights. Since the harmonic mean can assign higher weights to smaller radii, those semantic categories with smaller size but high cohesion can be preferentially retained in subsequent corpus selection. In this way, this embodiment ensures the diversity of the corpus while avoiding the excessive dominance of major categories, thereby improving the quality of the rehabilitation corpus in terms of coverage balance and semantic representativeness.
[0060] S300. Select natural language sentences corresponding to each of the semantic categories from a large-scale corpus to form a candidate stimulus corpus.
[0061] In step S300, natural language sentences that highly correspond to the aforementioned semantic categories need to be selected from a large-scale corpus to construct a candidate stimulus corpus. First, a suitable data source is selected as the corpus, such as publicly available news text libraries, daily dialogue datasets, and rehabilitation-related medical corpora. In this embodiment, CLUECorpus2020 and the Alpha Modern Chinese Corpus are preferred as corpus sources. The corpus can contain millions of natural language sentences to ensure sufficient semantic coverage. Second, all sentences in the corpus undergo text cleaning and standardization, removing punctuation, garbled characters, and grammatically incomplete samples, retaining grammatically correct and semantically complete candidate sentences. Then, using a pre-trained semantic matching model consistent with step S100 (e.g., the Text2Vec model based on the sentence-transformer architecture), semantic embedding is performed on each sentence to obtain a sentence semantic vector, and the cosine similarity between this semantic vector and the cluster center vectors of each semantic category is calculated. In this embodiment, for each semantic category, the sentence with the highest similarity ranking is selected from each sub-corpus as the corresponding candidate sentence, thus ensuring that each semantic category is covered. For cases with low or missing semantic category coverage, sentences are supplemented using stratified sampling in additional corpora to improve category balance. After forming the initial candidate set, further manual screening is performed to ensure grammatical correctness and natural expression, and to remove duplicate or irrelevant corpora that do not meet the needs of the rehabilitation scenario. Simultaneously, to avoid high redundancy among corpora, this embodiment calculates the cosine similarity of all candidate sentences pairwise, for example, setting a threshold of 0.75. When the similarity exceeds the threshold, only one sentence is retained, thereby controlling redundancy and ensuring corpus diversity.
[0062] In one alternative implementation, the process of selecting natural language sentences from a large-scale corpus that highly correspond to K semantic categories can be further defined. First, nᵢ candidate sentences are extracted for each semantic category, where the value of nᵢ can be adaptively adjusted based on the weight or importance of that semantic category in rehabilitation training. For example, more sentences can be extracted for categories involving daily communication, while the number of sentences can be appropriately reduced for categories involving peripheral domains, thus highlighting key semantic categories while ensuring coverage. The total number of candidate sentences across all categories is denoted as N, which is typically controlled between several hundred and several thousand to balance the richness and operability of the training corpus. Second, during the selection process, sentences with clear semantics, complete syntactic structure, and moderate length are prioritized, generally limited to 10 to 20 words, to ensure that patients can quickly understand them during rehabilitation training and avoid increasing cognitive burden due to overly long or complex sentences. Simultaneously, sentences containing ambiguity, low-frequency terms, or grammatical errors need to be removed to ensure the accuracy of language stimulation.
[0063] In another embodiment, this disclosure introduces readability constraints during the candidate stimulus corpus screening process to improve the rehabilitation suitability of the screened corpus. Specifically, after selecting preliminary candidate sentences through semantic similarity, the sentences are further subjected to readability testing, including two dimensions: vocabulary familiarity and sentence length. Vocabulary familiarity is determined based on a pre-established list of commonly used words, requiring that at least 80% of the words in the candidate sentences come from this list to ensure that the language expression is close to daily usage habits. Sentence length is preferably controlled between 10 and 20 words to avoid sentences that are too short to form complete semantics, or sentences that are too long to cause additional cognitive burden on patients during rehabilitation training. By introducing the above constraints, this embodiment can further improve the practicality and operability of the corpus in rehabilitation scenarios while ensuring semantic coverage and diversity.
[0064] S400. Based on the candidate stimulus corpus, collect functional magnetic resonance imaging data of the subjects and construct an encoding model from sentence semantic vectors to brain regions of interest related to language processing.
[0065] In step S400, as an optional implementation, a coding model can be constructed through joint analysis of the stimulus corpus and fMRI activation data. Specifically, the N sentences selected in step S300 are presented to the subject visually one by one. The presentation duration of each sentence can be set to t1 seconds (e.g., 4 seconds), and an interval of t2 seconds (e.g., 2 seconds) is set between sentence presentations to ensure sufficient recovery of blood oxygenation level-dependent signals. Throughout the process, functional magnetic resonance imaging (fMRI) is used to collect whole-brain scan data to obtain neural activity information of the subject during reading and understanding the corpus.
[0066] After the experiment, the acquired raw fMRI data underwent standardized preprocessing, including time correction, motion correction, spatial normalization, and signal smoothing. Brain regions closely related to language processing were extracted as Regions of Interest (ROIs), such as the left inferior frontal gyrus, superior temporal gyrus, and middle temporal gyrus. Within each ROI, the activation intensity sequences corresponding to N sentences were obtained by statistically summarizing the voxel signals. Simultaneously, these N sentences were transformed into semantic vectors using a semantic matching model consistent with step S100, thus establishing a one-to-one correspondence between semantic vectors and neural activation intensities.
[0067] Subsequently, a regression encoding model is trained using these paired data sets. The model input is the sentence semantic vector, and the output is the average activation intensity of the ROI. In this embodiment, regularized regression methods such as ridge regression or LASSO are preferably used to avoid overfitting and improve the model's generalization ability. After training, the model can represent the mapping relationship between the semantic space and neural activation.
[0068] Using the trained encoding model, forward prediction can be performed on all N sentences in the stimulus corpus, i.e., the semantic vector of the input sentence can be used to obtain its predicted activation intensity value in each ROI. In this way, researchers can quickly and accurately assess the neural activation potential of candidate corpora without having to repeat fMRI experiments on every corpus, providing a basis for the selection of subsequent rehabilitation training corpora.
[0069] In another embodiment, this application introduces a fatigue monitoring and segmented data acquisition strategy when collecting fMRI data based on a candidate stimulus corpus to improve the stability of experimental data. Specifically, during sentence presentation, in addition to the regular corpus, one or two short baseline sentences are periodically inserted. The activation patterns of these sentences are relatively stable within individuals. By comparing the changes in activation intensity of the baseline sentences during the experiment, the attention level and fatigue level of the subjects can be monitored in real time. When the monitoring results indicate that the subjects show a significant fatigue trend, the current experiment can be terminated early and supplementary data collection can be carried out in the subsequent stages, or multiple segmented experiments can be used to replace a long continuous experiment.
[0070] In another embodiment, this application proposes a state-adaptive calibration mechanism to address situations where individual subjects exhibit significant fluctuations in fMRI activation patterns across different scanning periods, attention levels, or physiological states. Existing technologies often assume that the neural activation patterns of the same subject remain stable at different times. However, in practical applications of rehabilitation training, a patient's attention, fatigue level, and even blood oxygen levels can affect the BOLD signal, leading to inconsistent activation results under the same sentence stimulus and reducing the robustness of the encoding model.
[0071] To address this, this application introduces a state calibration factor during the fMRI data preprocessing and modeling stages. Specifically, behavioral indicators (such as reaction time and task completion rate) or physiological signals (such as heart rate and respiratory rate) of the subjects are recorded simultaneously during the scan, and these indicators are mapped to state vectors. Subsequently, when constructing a regression model from semantic vectors to activation intensity, the state vectors are introduced as covariates, enabling the model to dynamically correct signal biases caused by state fluctuations. For example, when attention levels are low, the model automatically reduces the weight of noise dimensions and increases the weight of semantically relevant dimensions, thereby obtaining prediction results that more closely resemble the actual processing mechanism.
[0072] Through this mechanism, this application can significantly mitigate the impact of individual state fluctuations on model stability, ensuring that the coding model maintains high predictive consistency and reliability during long-term rehabilitation training. This not only improves the accuracy of corpus selection but also enhances the generalizability of the method in real-world clinical settings.
[0073] S500. Input the sentence semantic vectors in the candidate stimulus corpus into the encoding model to obtain the predicted activation values of the interest regions, and construct a rehabilitation training corpus set based on the predicted activation values.
[0074] In step S500, the semantic vectors of all sentences in the candidate stimulus corpus obtained in step S300 are sequentially input into the encoding model trained in step S400 to obtain the predicted activation value of each sentence in the brain regions of interest related to language processing. The predicted activation value can be understood as the intensity of the neural response that the sentence may elicit when the subject reads it, and is a core indicator for measuring the effectiveness of the corpus. To ensure the scientific rigor and practicality of the screening results, this embodiment sorts and filters all prediction results.
[0075] Specifically, all candidate sentences are first sorted from highest to lowest according to their predicted activation values, and the top p% (e.g., the top 10% or 20%) are selected as the target corpus. These sentences exhibit the highest activation potential in the model's predictions. Alternatively, an activation threshold τ can be pre-set, for example, τ=0.75. When the predicted activation value of a candidate sentence exceeds this threshold, that sentence is included in the rehabilitation training corpus. This approach avoids situations where the corpus is too large or too small, ensuring that the selection results are both representative and balanced.
[0076] In a further optimized embodiment, the prediction results of different regions of interest can be weighted and integrated. For example, core language brain regions such as the left inferior frontal gyrus and the posterior left temporal lobe can be assigned higher weights, while auxiliary language regions can be assigned lower weights, thereby constructing a comprehensive activation score. This score is then used to perform a secondary sorting or filtering of candidate sentences, ensuring that the final selected corpus can generate significant activation in multiple key language regions simultaneously.
[0077] Finally, the predicted activation values of each sentence in the candidate stimulus corpus are sorted, and the top-ranked target sentences are selected as rehabilitation training corpora; alternatively, an activation intensity threshold is set, and target sentences with predicted activation values greater than the threshold are selected to form the rehabilitation training corpus set. The final rehabilitation training corpus set can be controlled to a number of several hundred sentences, covering a wide range of semantic categories while ensuring that individuals can gradually accept and digest the information during rehabilitation training. This corpus set can be directly imported into a language rehabilitation system and used to guide aphasia patients or other people with language disorders in neurorehabilitation training through screen presentation or audio playback.
[0078] Figure 2 This is a schematic flowchart illustrating a method for forming a candidate stimulus corpus, provided as an embodiment of this disclosure. Now, in conjunction with... Figure 2 The specific embodiments of this application are further described below.
[0079] S310. Obtain a large-scale corpus, which includes multiple sub-corpus sets, each containing multiple sentences.
[0080] In step S310, a sufficiently large and comprehensive corpus is first required to ensure the diversity and comprehensiveness of subsequent corpus selection. This large-scale corpus can consist of multiple sub-corpus sets, each containing a large number of natural language sentences. In this embodiment, publicly available multi-domain corpus resources are preferably selected, such as CLUECorpus2020, the Alpha Modern Chinese Corpus, and corpora of dialogues, news reports, and medical rehabilitation scenarios. These sub-corpus sets can cover different contexts such as daily life, social interaction, medical health, and news reports, thereby ensuring the corpus has rich semantic diversity.
[0081] In the specific implementation process, to ensure the quality of the corpus, the acquired sub-corpus needs to be cleaned and preprocessed. Cleaning includes removing sentences with grammatical errors, incompleteness, or garbled characters; standardizing the encoding format; and eliminating excessively long (e.g., more than 50 words) or excessively short (less than 3 words) sentences to ensure the integrity of sentences in form and meaning. Furthermore, the corpus can be preliminarily screened according to the needs of rehabilitation applications, for example, prioritizing the retention of semantically clear sentences that are close to everyday communication, and reducing overly technical or obscure text.
[0082] In this embodiment, the resulting large-scale corpus can contain approximately several million sentences, with each sub-corpus containing at least one hundred thousand sentences. This hierarchical, multi-domain approach can provide sufficient candidate sentences for each semantic category in subsequent semantic matching and filtering stages, ensuring that the corpus can comprehensively cover the different semantic scenarios required for language rehabilitation.
[0083] S320. Calculate the semantic vector of each sentence in the large-scale corpus using a semantic matching model.
[0084] In step S320, the large-scale corpus obtained in step S310 is subjected to semantic embedding processing to obtain sentence semantic vectors for subsequent matching with semantic categories. Specifically, a pre-trained semantic matching model consistent with step S100 is used to map each sentence in the corpus to a corresponding semantic vector. In this embodiment, a semantic matching model based on the sentence-transformer architecture, such as Text2Vec or SimCSE, is preferably used. This type of model can effectively capture semantic features at the sentence level and transform natural language sentences into high-dimensional continuous vectors.
[0085] In practice, each sentence in the corpus first undergoes basic text preprocessing, including punctuation removal, synonym standardization, case unification, and word segmentation, to ensure input consistency and stability. Then, the preprocessed sentences are input into a semantic matching model to obtain semantic vectors of dimension d, where d is set to 768 dimensions to maintain good semantic expressiveness while accommodating the computational efficiency required for large-scale processing. In this way, millions of sentences can be transformed into a standardized set of semantic vectors. This set of semantic vectors will be used in subsequent steps to calculate cosine similarity with the cluster center vectors of semantic categories, thereby achieving automatic selection of candidate stimulus corpora.
[0086] S330. The similarity between each sentence and each of the semantic categories is obtained by calculating cosine similarity.
[0087] In step S330, the semantic vector of each sentence obtained in step S320 needs to be compared with the semantic category cluster center vector obtained in step S200 to determine the degree of matching between the sentence and different semantic categories. Specifically, cosine similarity is used as a metric to calculate the similarity between each sentence and the center of each semantic category. In this way, a similarity distribution can be obtained to reflect the degree of fit of the sentence under different semantic categories. For example, when a sentence has the highest similarity value with the "medical care scenario" category and exceeds a preset threshold, the sentence can be determined to belong to that semantic category. In large-scale corpus processing, this method can quickly and effectively find the most suitable semantic category for each sentence, thus providing basic data support for the subsequent screening of candidate stimulus corpora.
[0088] S340. For each of the sub-corpora, select the sentences with the highest similarity to each of the semantic categories to form a candidate stimulus corpus.
[0089] In step S340, based on the similarity results obtained in step S330, each sub-corpus is filtered separately. Specifically, for each semantic category, the similarity between the cluster center of that category and all sentences in the sub-corpus is calculated, and the similarity values are sorted. The sentence with the highest ranking is selected as the representative sentence of that category in the sub-corpus.
[0090] In a specific example, suppose that K=300 semantic categories are obtained through clustering in step S200, and each sub-corpus contains approximately 100,000 natural language sentences. After calculating similarity in step S330, the sentence that best matches each of the 300 semantic categories in each sub-corpus can be found. For example, in the "daily life corpus," the sentence with the highest similarity corresponding to the semantic category "medical care scenario" might be "A nurse is taking a patient's blood pressure," the sentence with the highest similarity corresponding to the semantic category "shopping scenario" might be "I want to buy a pound of fresh apples," and the sentence with the highest similarity corresponding to the semantic category "education scenario" might be "A teacher is explaining a text to students."
[0091] In this way, approximately 300 representative sentences covering different semantic categories can be obtained for each sub-corpus. If the corpus contains four sub-corpora, then after this step, the candidate stimulus corpus will contain approximately 1200 sentences. Subsequently, through redundancy control and manual review, sentences with excessively high semantic similarity or unnatural grammar are removed, ultimately forming a candidate stimulus corpus of approximately 1000 sentences with balanced coverage and strong semantic diversity. This candidate corpus will serve as an important input for subsequent fMRI experiments and coding model construction, thereby ensuring the scientific rigor and effectiveness of the rehabilitation training data.
[0092] Figure 3 This diagram illustrates an embodiment of the present disclosure of a coding model for constructing a language processing-related brain region of interest from a sentence semantic vector, further illustrating specific embodiments of the present application.
[0093] S410. By comparing functional magnetic resonance imaging data of real sentences and pseudo-sentences, the brain regions of interest related to language processing in the subject are determined.
[0094] In step S410, it is first necessary to individually locate the brain regions of interest related to language processing in the subjects to ensure that the subsequent construction of the encoding model can be carried out on the core regions of their brain's language network. To this end, this embodiment designed a comparative experiment of real sentences and pseudo-sentences. Real sentences are selected from corpora outside the candidate stimulus corpus formed in step S300 to ensure that subjects do not come into contact with the training sentences used for modeling in advance during the experiment; pseudo-sentences are random combinations of words that are grammatically acceptable but lack complete semantics, such as "the table looks at the sky" or "the apple runs in the book".
[0095] In this embodiment, real sentences and pseudo-sentences were used as stimulus materials, and functional magnetic resonance imaging (fMRI) scans were performed on the subjects. The pseudo-sentences were grammatically correct but lacked complete semantics, serving as a contrast to the real sentences. A saliency map was obtained by statistically comparing brain activation under real and pseudo-sentence conditions. Subsequently, this saliency map was spatially overlapped with a standard language area template, and the voxels with the highest saliency within each template region were selected as individualized regions of interest (ROIs). This approach combines the anatomical consistency of the group template with the subjects' own activation characteristics, ensuring the accuracy and individualization of subsequent semantic vector and neural activation intensity modeling.
[0096] Subsequently, the obtained statistical significance map was spatially overlapped with the standard language area template (including the left inferior frontal gyrus, the orbital part of the left inferior frontal gyrus, the left middle frontal gyrus, the posterior part of the left temporal lobe, and the middle and posterior part of the left temporal lobe).
[0097] Within each template region, voxels ranking in the top percentile (e.g., the top 10%) based on statistical significance are selected as the Regions of Interest (ROIs) for language processing in each individual subject. This approach not only ensures a high correlation between the ROIs and the language processing task but also eliminates interference from anatomical differences at the individual level, thus providing a stable and personalized neural basis for the subsequent mapping of semantic vectors to brain region activation intensity.
[0098] In another preferred embodiment, functional connectivity constraints can be introduced to further enhance the functional specificity and network representativeness of ROI localization. Specifically, within the ROI initially defined based on task activation intensity, the functional connectivity strength between each voxel and these nodes is calculated (e.g., quantified using Pearson correlation coefficient, partial correlation analysis, etc.) based on more clearly defined core nodes in the standard language region template (e.g., the left inferior frontal gyrus, the posterior left temporal lobe). By retaining voxels that exhibit high activation in language tasks and maintain high functional synchronization with the core nodes of the language network, spurious responses caused by non-specific activation or noise can be effectively suppressed, providing a more stable and interpretable neural signal basis for subsequent coding models.
[0099] S420. Obtain the BOLD value of the language processing-related brain region of interest for each sentence stimulus in the stimulus corpus, and construct an encoding model from the sentence semantic vector to the language processing-related brain region of interest using the BOLD value.
[0100] In step S420, based on the individualized language processing-related brain regions of interest (ROIs) determined in step S410, all sentences in the candidate stimulus corpus are presented to the subject one by one, and whole-brain fMRI data are continuously collected during the stimulus presentation process. The semantic vectors of the sentences in the candidate stimulus corpus are used as input variables, and the average BOLD value of the language processing-related brain regions of interest under the corresponding sentence stimulus is used as the output variable. By performing temporal correction, spatial registration, standardization, and denoising on the data, voxel signals within the ROI region under each sentence stimulus are extracted. To ensure data stability, this embodiment uses the voxel averaging method, that is, averaging the BOLD signals of all voxels within the ROI during the sentence presentation period, which is used as the activation intensity of the sentence under that ROI. Finally, a one-to-one data pair is formed between the semantic vector representation of each sentence and its neural activation intensity.
[0101] In the modeling phase, a regularized regression algorithm is used to establish a mapping relationship between input and output variables, thereby training the encoding model that maps sentence semantic vectors to language processing-related brain regions of interest. To establish a stable input-output mapping relationship, this embodiment preferably uses regularized regression methods such as ridge regression to train the encoding model.
[0102] In another preferred embodiment, this application proposes a transfer learning mechanism to address the problems of high data acquisition costs and limited data from individual patients in functional magnetic resonance imaging (fMRI). Specifically, this application can initially train a general semantic-neural coding model using large-scale fMRI data from multiple healthy subjects or rehabilitation patients. This model can characterize the general mapping relationship between natural language semantic vectors and the activation of language-related brain regions. After obtaining this general model, for newly admitted patients undergoing rehabilitation training, only a small amount of individualized fMRI sample data needs to be collected. Based on this general model, parameter fine-tuning can be performed to quickly adapt to the brain region activation characteristics of the new patients, thereby achieving individualized model customization under small sample conditions.
[0103] In its implementation, this application employs a combination of transfer learning and few-sample adaptive training. For example, a base model is first trained using a large sample of data from multiple subjects, with parameters including mapping weights from semantic vectors to multi-region BOLD responses. Subsequently, with limited data from new patients (e.g., ten to twenty sentence stimuli), the main parameters of the model are frozen, and only some high-level parameters or regularization terms are fine-tuned. This allows the model to learn individual differences from new patients while maintaining its general semantic-brain region mapping capabilities. This transfer learning mechanism can significantly shorten the modeling time for new patients and reduce reliance on repeated fMRI experiments.
[0104] By combining transfer learning with rapid customization using small samples, this application significantly reduces the number of fMRI experiments per patient while maintaining high predictive accuracy, solving the problem of "individualized models relying on large-scale neuroimaging data" in existing technologies. Compared to the traditional method of independent modeling per subject, this application's method improves the efficiency and scalability of rehabilitation corpus screening and enhances the practicality of the rehabilitation system in clinical settings. Figure 4 This is a schematic diagram illustrating a method for determining brain regions of interest related to language processing in test subjects, as provided in an embodiment of this disclosure. Now, in conjunction with... Figure 4 The specific embodiments of this application are further described below.
[0105] S411. Using real sentences and pseudo-sentences as visual stimulus materials, the subjects are scanned to obtain functional magnetic resonance imaging data.
[0106] In step S411, to determine the brain regions of interest related to language processing in the subjects, a contrastive experiment between real and pseudo-sentences is first designed. Real sentences are selected from an independent corpus different from the candidate stimulus corpus, ensuring semantic integrity and natural syntax, such as "The nurse is measuring the patient's blood pressure." Pseudo-sentences are randomly combined from common words; while conforming to grammatical rules in form, they lack complete semantics, such as "The table is singing a quiet blue." This set of stimulus materials can create a clear contrast between language processing and non-language processing conditions.
[0107] During the experiment, real and pseudo-sentences were presented to the subjects in a random order. Each sentence was presented for 4 seconds, with a 2-second interval between sentences to ensure sufficient response of the blood oxygenation signal. The stimulation process could employ either an event-related design or a block design; this embodiment preferred an event-related design to more accurately separate neural responses under single-sentence conditions. Subjects were required to maintain fixation and try to understand the presented content during the scanning process.
[0108] S412. Compare the brain region activation responses under the conditions of real sentences and pseudo-sentences, and construct a significance statistical map.
[0109] In step S412, based on the fMRI data collected in step S411 under both real and pseudo-sentence stimulus conditions, the brain activation response of the subjects is compared and analyzed. Specifically, the collected whole-brain BOLD signals are first standardized and preprocessed, including time slice correction, motion artifact correction, spatial normalization, and smoothing, to ensure the comparability of data between different scanning time periods and different subjects.
[0110] Subsequently, time-series regression models were established based on a generalized linear model (GLM) under real sentence and pseudo-sentence conditions. The real sentence condition was set as the experimental group, and the pseudo-sentence condition as the control group. Voxel signal differences between the two conditions across the entire brain were extracted. Through the above process, a significance map (t-map) was obtained, where areas with higher brightness or color represent voxel distributions with stronger activation under the real sentence condition compared to the pseudo-sentence condition. This significance map can intuitively reflect the activation patterns of subjects in brain regions related to language processing, providing a basis for spatial overlap analysis with standard language area templates in subsequent steps.
[0111] S413. Perform spatial overlap analysis between the saliency statistical graph and the standard language area template to obtain the individualized language processing-related brain regions of interest for the subject.
[0112] In step S413, the significance statistical map obtained in step S412 is spatially overlapped with the standard language area template to determine the individualized brain regions of interest (ROIs) related to language processing in the subjects. Specifically, the standard language area template can be a widely used ROI template of the functional language network in cognitive neuroscience, including the left inferior frontal gyrus, the orbital portion of the left inferior frontal gyrus, the left middle frontal gyrus, the posterior portion of the left temporal lobe, and the middle and posterior portions of the left temporal lobe. This template has a defined spatial range under the MNI standard space and can be used as a reference benchmark.
[0113] In practice, the saliency statistical map generated in step S412 is first spatially normalized to ensure its coordinate system is consistent with the standard language region template. Then, a spatial overlap operation is performed to intersect the significantly activated voxel regions in the statistical map with the corresponding regions in the template. This method determines which significantly activated voxels belong to language-related brain regions. To improve individualization accuracy, this embodiment further selects voxels with statistical values ranking in the top percentiles (e.g., the top 10% or top 20%) within each template region as the final ROI. This ensures that the selected regions are highly correlated with language processing while eliminating interference from marginal signals or noise.
[0114] The final result is a set of individualized brain regions of interest (ROIs) related to language processing for each subject, which preserves the anatomical consistency of the group template while incorporating the subject's own activation characteristics. These ROIs will be used in step S420 to extract BOLD signals under sentence stimuli and will further participate in the encoding and modeling of semantic vectors and neural activation intensity.
[0115] According to embodiments of this disclosure, an electronic device is also provided, which may include a processor, a communications interface, a memory, and a communication bus, wherein the processor, communications interface, and memory communicate with each other via the communication bus. The processor can invoke logical instructions in the memory to execute a configuration software-based software licensing implementation method.
[0116] Furthermore, the logical instructions in the aforementioned memory can be implemented as software functional units and sold or used as independent products, and can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this disclosure, in essence, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this disclosure. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0117] On the other hand, this disclosure also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, is implemented to perform the configuration software-based software licensing implementation methods provided by the above methods.
[0118] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Those skilled in the art can understand and implement this without any creative effort.
[0119] Through the above description of the embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus necessary general-purpose hardware platforms, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solutions, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods described in the various embodiments or some parts of the embodiments.
[0120] It should be understood that the above embodiments are only used to illustrate the technical solutions of this disclosure, and not to limit them; although this disclosure has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or make equivalent substitutions for some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this disclosure.
Claims
1. A corpus screening method for language function rehabilitation based on semantic embedding and neural encoding, characterized in that, The method comprises the following steps: performing semantic embedding processing on each word in the vocabulary based on a pre-trained semantic matching model to obtain a word semantic vector; performing cluster analysis on the word semantic vectors to obtain a plurality of semantic categories; filtering natural language sentences corresponding to each semantic category from a large-scale corpus to form a candidate stimulus corpus; collecting functional magnetic resonance imaging data of a test subject based on the candidate stimulus corpus to construct an encoding model from a sentence semantic vector to a language processing related brain region of interest; inputting the sentence semantic vector in the candidate stimulus corpus into the encoding model to obtain a predicted activation value of the region of interest, and constructing a rehabilitation training corpus set according to the predicted activation value.
2. The language function rehabilitation corpus screening method based on semantic embedding and neural coding according to claim 1, characterized in that, The method further comprises the following steps of filtering natural language sentences corresponding to each semantic category from a large-scale corpus to form a candidate stimulus corpus: obtaining a large-scale corpus, wherein the large-scale corpus comprises a plurality of sub-corpus sets, and each sub-corpus set comprises a plurality of sentences; performing semantic embedding processing on each sentence in the large-scale corpus based on a pre-trained semantic matching model to obtain a sentence semantic vector; obtaining the similarity of each sentence to each semantic category through cosine similarity calculation; for each sub-corpus set, filtering out the sentence with the highest similarity to each semantic category to form a candidate stimulus corpus.
3. The language function rehabilitation corpus screening method based on semantic embedding and neural coding according to claim 1, characterized in that, The method further comprises the following steps of collecting functional magnetic resonance imaging data of a test subject based on the candidate stimulus corpus to construct an encoding model from a sentence semantic vector to a language processing related brain region of interest: determining the language processing related brain region of interest of the test subject by comparing the functional magnetic resonance imaging data of real sentences and pseudo sentences; obtaining the BOLD value of the language processing related brain region of interest of the test subject under the stimulation of each sentence in the stimulus corpus, and constructing an encoding model from a sentence semantic vector to the language processing related brain region of interest through the BOLD value.
4. The language function rehabilitation corpus screening method based on semantic embedding and neural coding according to claim 3, characterized in that, The method further comprises the following steps of determining the language processing related brain region of interest of the test subject by comparing the functional magnetic resonance imaging data of real sentences and pseudo sentences: scanning the test subject using real sentences and pseudo sentences as visual stimulation materials to obtain functional magnetic resonance imaging data; the pseudo sentence is a sentence with reasonable grammar structure but without complete semantics formed by random words; comparing the brain region activation response under the stimulation of the real sentence and the pseudo sentence to construct a significance statistical map; performing spatial coincidence analysis on the significance statistical map and a standard language region template to obtain the language processing related brain region of interest of the test subject.
5. The language function rehabilitation corpus screening method based on semantic embedding and neural coding according to claim 3, characterized in that, The method further comprises the following steps of constructing an encoding model from a sentence semantic vector to the language processing related brain region of interest through the BOLD value: inputting the sentence semantic vector in the candidate stimulus corpus as an input variable; inputting the average BOLD value of the language processing related brain region of interest under the stimulation of the corresponding sentence as an output variable; The mapping relationship between the input variables and the output variables is established by using a regularization regression algorithm to train an encoding model from the sentence semantic vector to the brain region of interest related to the language processing.
6. The language function rehabilitation corpus screening method based on semantic embedding and neural coding according to claim 1, characterized in that, The method further includes: The pre-trained semantic matching model is based on a sentence-transformer architecture, and each word in the vocabulary is subjected to semantic embedding processing to obtain a semantic vector corresponding to each word.
7. The language function rehabilitation corpus screening method based on semantic embedding and neural coding according to claim 1, characterized in that, The method further includes: The method further includes: A similarity matrix is constructed between the semantic vectors of the words.
8. The language function rehabilitation corpus screening method based on semantic embedding and neural coding according to claim 1, characterized in that, A spectral clustering algorithm is applied to the similarity matrix to divide the words into different semantic categories. The method further includes: The predicted activation values of the sentences in the candidate stimulus corpus are sorted, and a target sentence with a high ranking is selected as a rehabilitation training corpus.
9. An electronic device comprising: Alternatively, an activation intensity threshold is set, and a target sentence with a predicted activation value greater than the threshold is selected to form the rehabilitation training corpus. A processor and a memory in communication with the processor; characterized in that: The memory stores computer execution instructions; 10. A computer-readable storage medium having stored thereon a computer program, characterized in that, The processor executes the computer execution instructions stored in the memory to perform the steps of the method of any one of claims 1-8. The program is executed by the processor to perform the steps of the method of any one of claims 1-8.