A method, system, device and medium for parsing human brain syntax functions based on machine learning

By constructing an extreme gradient enhancement algorithm regression model and Shapley's interpretation theory, combined with graph embedding algorithms, the syntactic structure in natural text corpora is analyzed. This solves the problem of insufficient understanding of human brain syntactic functions in existing technologies, and realizes a comprehensive analysis of human brain syntactic functions and regional differences.

CN116598020BActive Publication Date: 2026-08-25JIANGXI UNIVERSITY OF FINANCE AND ECONOMICS
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202310509864.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-05-08
Publication Date
2026-08-25
Estimated Expiration
2043-05-08

AI Technical Summary

Technical Problem

Existing technologies are insufficient to fully and deeply understand how the human brain processes syntactic functions, and the limitations of artificial corpora lead to fragmented experimental results, making it difficult to integrate multiple experimental results.

Method used

Using a machine learning-based approach, we construct an extreme gradient boosting algorithm regression model, combine it with Shapley's method of interpretation and graph embedding algorithms, analyze the syntactic structure in natural text corpora, build a syntactic network and perform region classification, and demonstrate syntactic functional connections.

Benefits of technology

It achieves a comprehensive and in-depth analysis of the syntactic functions of the human brain, and can intuitively display the functional distribution patterns and regional differences of syntactic processing, capture the differences in syntactic processing between voxels, and demonstrate the connection strength of syntactic functions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116598020B_ABST
    Figure CN116598020B_ABST
Patent Text Reader

Abstract

A kind of machine learning-based human brain syntax function analysis method, system, device and medium, method includes: first, construct extreme gradient boosting algorithm (XGBoost) regression model sample set, and train XGBoost regression model for voxel, then calculate each XGBoost regression model output signal by Shapley plus method interpretation theory, obtain the sample interaction matrix corresponding to XGBoost regression model and statistically obtain the syntax vector of voxel, to obtain the syntax function intensity of voxel and the functional brain area when human brain processes different syntax structure, finally, construct the syntax network corresponding to voxel by interaction matrix, and based on the structure of syntax network, the region classification of voxel is carried out, and the function correlation of different regions based on syntax is obtained;System, device and medium are used to realize a kind of machine learning-based human brain syntax function analysis method;The present application shows the contrast of the advantage processing brain area and functional connection of different syntax structure, can be in-depth, comprehensive understanding and explanation syntax function details in human brain and intuitive presentation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of interdisciplinary technology of machine learning and cognitive neuroscience, specifically to a method, system, device, and medium for parsing human brain syntactic functions based on machine learning. Background Technology

[0002] The development of human civilization often requires knowledge sharing, exchange, learning, and win-win cooperation among groups. Language, as the carrier of thought and a tool for exchanging information, plays an irreplaceable role in the development of civilization. Therefore, how the human brain encodes its own thoughts into language for output, and how it decodes others' statements to achieve intellectual exchange, has always been a highly focused research topic in neuroscience. Language cognition research is divided into semantic and syntactic directions. Semantic research focuses on the brain's processing mechanisms of the meaning represented by individual language units (such as Chinese characters or words), while syntactic research mainly focuses on how the brain combines a series of discrete language units into complete sentences for understanding through syntactic functions. Language without syntax is like a jumbled mess, making it difficult to accurately convey thoughts. Therefore, studying how to analyze the syntactic functions of the human brain is a worthwhile research question.

[0003] Syntax is key to the human brain's understanding of language. In existing research, some researchers have explored the brain's syntactic functions using natural text corpora, but these studies only treat syntax as a whole, identifying that syntactic and semantic functions share the same brain regions for language processing. Other researchers have designed artificial corpora to explore the brain's responses to specific syntactic phenomena, but these corpora have limitations and cannot integrate results from different experiments. Therefore, the details of syntactic functions in the human brain have not been fully and deeply understood and explained.

[0004] Previous studies exploring brain regions sensitive to syntactic function have found that syntactic and semantic functions share the same language regions within the brain (AJ Reddy and L. Wehbe, "Can fMRI reveal the representation of syntactic structure in the brain?", Advances in Neural Information Processing Systems, Vol. 34, pp. 9843–9856, 2021.) or that they have a high degree of regional overlap (S. Wang, J. Zhang, N. Lin, and C. Zong, "Probing Brain Activation Patterns by Dissociating Semantics and Syntax in Sentences", AAAI, Vol. 34, Issue 05, pp. 9201–9208, April 2020, doi:10.1609 / aaai.v34i05.6457.). However, these studies have struggled to demonstrate differences in syntactic function between different regions. Another group of studies uses artificially designed experimental corpora to explore how selected regions respond to a certain type of syntactic structure (e.g., adjective + noun) (M. Schell, E. Zaccarrella, and A.D. Friederici, *Differential cortical contribution of syntax and semantics: An fMRI study on two-word phrasal processing*, Cortex, Vol. 96, pp. 105–120, 2017). However, due to limitations in experimental conditions, such as the limitations of highly homogenized artificial corpora, the experimental results are often too fragmented, making it difficult to integrate the results of multiple experiments focusing only on a single linguistic phenomenon. Therefore, previous research still lacks a comprehensive and in-depth analysis of how the entire brain's language system processes syntax. Summary of the Invention

[0005] To overcome the shortcomings of the prior art, the present invention aims to provide a method, system, device, and medium for analyzing human brain syntactic functions based on machine learning. It uses words in natural text corpora as a sample set, constructs and calculates an extreme gradient boosting algorithm regression model for each voxel to obtain the interaction matrix between various features, then extracts the interaction relationships belonging to syntactic functions based on dependency grammar theory, and finally performs statistical analysis to analyze the operating mechanism of human brain syntactic functions, presenting the results intuitively. This invention has the characteristics of comprehensively and deeply understanding and explaining the details of syntactic functions in the human brain.

[0006] To achieve the above object, the technical solution adopted by the present invention is as follows:

[0007] A method for parsing the syntactic function of the human brain based on machine learning, comprising the following steps:

[0008] Step 1, construct a sample set of an Extreme Gradient Boosting (XGBoost) regression model, and train an Extreme Gradient Boosting regression model corresponding to each voxel;

[0009] Step 2, calculate the output signal of each Extreme Gradient Boosting regression model through the Shapley additive explanation theory, and obtain the sample interaction matrix corresponding to the Extreme Gradient Boosting regression model;

[0010] Step 3, based on the syntactic structure relationship of different words, perform numerical statistics on the sample interaction matrix obtained in Step 2 to obtain the syntactic vector of each voxel;

[0011] Step 4, through the syntactic vector obtained in Step 3, obtain the syntactic function strength of each voxel and the functional brain regions when the human brain processes different syntactic structures;

[0012] Step 5, construct a syntactic network corresponding to the voxel through the sample interaction matrix obtained in Step 2, and classify the voxels based on the structure of the syntactic network;

[0013] Step 6, obtain the syntactic vector corresponding to the region classified in Step 5 through the constructed Extreme Gradient Boosting regression model, and display the syntactic function connections of different regions.

[0014] The specific content of Step 1 is as follows:

[0015] Step 1.1, play a natural text story for the subject, and collect the activation signal of his brain through functional magnetic resonance imaging (fMRI) technology;

[0016] Step 1.2, when the total duration of playing the natural text story for the subject is N', according to the activation signal acquisition interval of functional magnetic resonance imaging (fMRI), take the activation signals of the voxel at N (N < N') moments as the output signal of the Extreme Gradient Boosting regression model

[0017]

[0018] In formula (1), yN For the Nth output signal

[0019] Step 1.3: Use M different words from the natural text story to form a word set. And design matrix A = (a i,j ) 1≤i≤M,1≤j≤N' When the subject listens to the story, if the i-th word appears at time j, then a i,j =1, otherwise a i,j =0;

[0020] Step 1.4: The i-th row vector a of matrix A... i,* Discrete convolution is performed with the Hemodynamic Response Function (HRF) to obtain the input signal value x of the i-th word to the human brain at each time step from the matrix A designed in step 1.3. i :

[0021]

[0022] In equation (2), b is the HRF function f = 0.452t 8.6 e -1.828t discrete values, x i,j Let be the input signal value of the i-th word to the human brain at time j;

[0023] Step 1.5, x i By downsampling based on the activation signal acquisition interval of fMRI, the input signal value x' of the i-th word to the brain at N time points is obtained. i :

[0024] x' i =(x i,1 ,x i,2 ,…,x i,N ) T Equation (3)

[0025] Step 1.6: Construct the input feature matrix X based on equation (3):

[0026] X = (x'1, x'2, ..., x') i ,…,x' M ) T ,X∈R M×N Equation (4)

[0027] Step 1.7: Represent the j-th column vector of matrix X as x. j And compared with the output signal y of the extreme gradient boosting algorithm regression model in equation (1) j Forming training data pairs (xj ,y j If the extreme gradient boosting algorithm regression model has an input sample set of , then the input sample set of the extreme gradient boosting algorithm regression model is . The output sample set of the extreme gradient boosting algorithm regression model is No less than 80% of the sample set is designated as the training set, and the remaining sample set is used as the test set.

[0028] Step 1.8: Use the activation signal of any voxel as the output signal of the corresponding extreme gradient boosting algorithm regression model. With input signal A training sample set is constructed, an XGBoost regression model is built, K tree models f are trained, and the predictions of all tree models for the same sample are summed to obtain the final prediction.

[0029]

[0030] In equation (5), x i It is the i-th sample, f k (x i ) is the k-th tree pair of sample x i The predicted value, It is a function space. This is the final predicted value, and the corresponding true value is y. i ;

[0031] Step 1.9: Using the final predicted value from Step 1.8 and the true value y i Construct the loss function l, and after adding a regularization term Ω to control the model structure risk, construct the objective function Obj that needs to be optimized:

[0032]

[0033] Step 1.10: Using the sample set divided in Step 1.7 and the objective function Obj constructed in Step 1.9, train the extreme gradient boosting algorithm regression model corresponding to each voxel.

[0034] Step 2 specifically includes:

[0035] Step 2.1: Use the TreeExplainer method in SHapley Additive exPlanations (SHAP) to decompose the output value of each sample of the extreme gradient boosting algorithm regression model, and generate a symmetric sample interaction matrix.

[0036] Step 2.2: Absolute-value the sample interaction matrices generated from all samples and sum them to obtain the final sample interaction matrix Mat. The elements of the sample interaction matrix Mat reflect the overall importance of all features and their interactions to the human brain activation signals.

[0037]

[0038] In equation (7), TreeE(x) i This refers to interpreting sample x using the TreeExplainer method. i This will generate a sample interaction matrix.

[0039] Step 3 specifically includes:

[0040] Step 3.1: Based on the syntactic relationships in the sentence, divide the word pairs into various syntactic structures;

[0041] Step 3.2: Use the natural language text processing library (spaCy) to find all word pairs in the natural text that have syntactic structural relationships;

[0042] Step 3.3: Each type of syntactic structure corresponds to a set of word pairs.

[0043] Step 3.4: Pair any word with a given syntactic structure The average value of the interaction values ​​of all word pairs (i,j) in the sample interaction matrix Mat is considered as the functional strength P of the voxel corresponding to the interaction matrix Mat in processing syntactic structure.

[0044]

[0045] In equation (8), L is the number of word pairs belonging to the syntactic structure, i.e.

[0046] Step 3.5: After calculating different syntactic structures and obtaining the corresponding functional strength P, combine all functional strengths P into a syntactic vector.

[0047] Step 4 specifically includes:

[0048] Step 4.1: Superimpose the values ​​of each dimension of the syntax vector corresponding to the voxel:

[0049] Total=∑Proc s ,s={nsubj,amod,…}Formula (9)

[0050] In equation (9), Total represents the syntactic functional strength of the voxels as a whole;

[0051] Step 4.2: Compare the values ​​of the same dimension of the syntactic vector corresponding to each voxel to obtain the functional strength when all voxels process the same syntactic structure;

[0052] Step 4.3: Display and compare the functional strengths obtained in Step 4.2 to obtain the functional brain regions when the human brain processes different syntactic structures.

[0053] Step 5 specifically includes:

[0054] Step 5.1: Combine all word pairs that have syntactic structural relations into a set. Construct the adjacency matrix W of the syntactic network based on the sample interaction matrix Mat:

[0055]

[0056] Then, using the adjacency matrix W of each voxel, a syntactic network corresponding to each voxel is constructed;

[0057] Step 5.2: Based on the graph embedding algorithm (graph2vec), transform all syntactic networks into vectors to form a vector space.

[0058] Step 5.3: Based on the classification rules of the K-means clustering algorithm, classify and divide all syntactic networks in the vector space formed in Step 5.2.

[0059] Step 6 specifically includes:

[0060] Step 6.1: Take the average value of all voxel signals in each region as the signal of that region, construct an extreme gradient boosting algorithm regression model for each region according to steps 1 to 3, and obtain the corresponding sample interaction matrix;

[0061] Step 6.2: Perform numerical statistics on the sample interaction matrix based on syntactic structure relations, and combine them into syntactic vectors corresponding to each region;

[0062] Step 6.3: Finally, calculate the Pearson correlation coefficient of the syntactic vectors corresponding to different regions to obtain the correlation between different regions.

[0063] A machine learning-based system for parsing human brain syntactic functions includes:

[0064] Sample set construction and modeling module: Constructs a sample set for the extreme gradient boosting algorithm regression model, and uses the sample set to construct the corresponding extreme gradient boosting algorithm regression model for each voxel;

[0065] Shapley's additive interpretation theory module: Calculates the output signal of each extreme gradient boosting algorithm regression model using Shapley's additive interpretation theory, and obtains the sample interaction matrix corresponding to the extreme gradient boosting algorithm regression model;

[0066] Numerical statistics module: By analyzing the syntactic structural relationships of different words, numerical statistics are performed in the interaction matrix to obtain the syntactic vector of each voxel;

[0067] Human Brain Syntactic Function Demonstration Module: Demonstrates the syntactic function mechanism of the human brain, including the overall syntactic functional brain regions and the functional brain regions when processing different syntactic structures;

[0068] Syntactic network construction module: Constructs syntactic networks corresponding to voxels through interaction matrices, and can classify voxels into regions based on the network structure;

[0069] Syntactic Function Connection Display Module: Displays the syntactic function connection relationships between different regions.

[0070] A machine learning-based device for parsing human brain syntactic functions includes:

[0071] Memory: for storing a computer program that implements the machine learning-based human brain syntactic function parsing method as described in any one of claims 1-7;

[0072] Processor: Used to implement the machine learning-based human brain syntactic function parsing method as described in any one of claims 1-7 when executing the computer program.

[0073] A computer-readable storage medium storing a computer program that, when executed by a processor, implements the steps of the machine learning-based human brain syntactic function parsing method.

[0074] Compared with the prior art, the beneficial effects of the present invention are as follows:

[0075] 1. By analyzing the processing intensity of each voxel, the differences in the advantages of different syntactic structures in processing brain regions can be demonstrated, thus providing an intuitive analysis of the functional distribution pattern of the human brain in syntactic processing.

[0076] 2. By constructing a syntactic network, the syntactic processing differences between voxels can be captured at the word pair level, thereby enabling the re-division of regions based on syntactic differences.

[0077] 3. Two types of regional syntactic functional connectivity: By comparing them, we can show the strength of syntactic functional connectivity in each region as well as the overall functional connectivity strength, and thus understand the collaborative mechanism of different brain regions in processing syntax from the perspective of functional connectivity. Attached Figure Description

[0078] Figure 1 (a) is a relational diagram showing the construction of the sample set, the output signal of the Shapley addition method interpretation theoretical calculation model, the acquisition of the sample interaction value matrix, and the construction of the syntactic network in this invention. Figure 1 (b) is a schematic diagram of the performance of the voxel-corresponding extreme gradient enhancement algorithm regression model of the present invention.

[0079] Figure 2 (a) is a schematic diagram of the overall syntactic functional strength of the present invention. Figure 2 (b) is a schematic diagram of the functional strength of a specific syntactic structure of the present invention.

[0080] Figure 3 (a) is a schematic diagram of the syntactic network of the present invention. Figure 3 (b) is a schematic diagram of the syntactic function connection of the present invention. Figure 3 (c) is a schematic diagram of the overall functional connection of the present invention. Detailed Implementation

[0081] The working principle of the present invention will now be described in detail with reference to the accompanying drawings.

[0082] See Figure 1 (a) A machine learning-based method for parsing human brain syntactic functions, comprising the following steps:

[0083] Step 1: Construct a sample set for the Extreme Gradient Boosting (XGBoost) regression model, and train the corresponding Extreme Gradient Boosting regression model for each voxel.

[0084] Step 2: Calculate the output signal of each extreme gradient boosting algorithm regression model using the Shapley addition method to obtain the sample interaction matrix corresponding to the extreme gradient boosting algorithm regression model;

[0085] Step 3: Based on the syntactic structure relationships of different words, perform numerical statistics on the sample interaction matrix obtained in Step 2 to obtain the syntactic vector of each voxel;

[0086] Step 4: Using the syntactic vectors obtained in Step 3, obtain the syntactic functional strength of each voxel, as well as the functional brain regions in the human brain when processing different syntactic structures.

[0087] Step 5: Construct the syntactic network corresponding to the voxels using the sample interaction matrix obtained in Step 2, and classify the voxels into regions based on the structure of the syntactic network;

[0088] Step 6: Obtain the syntactic vectors corresponding to the regions classified in Step 5 by constructing an extreme gradient boosting algorithm regression model, and display the syntactic function connections of different regions.

[0089] See Figure 1 (b), Step 1 is specifically as follows:

[0090] Step 1.1: Play a natural text story to the subject, and collect the activation signals of their brain through functional magnetic resonance imaging (fMRI) technology;

[0091] Step 1.2: When the total duration of playing the natural text story to the subject is N', according to the activation signal acquisition interval of functional magnetic resonance imaging (fMRI), take the activation signals of the voxel at N (N < N') moments as the output signals of the extreme gradient boosting algorithm regression model

[0092]

[0093] In formula (1), y N is the Nth output signal

[0094] Step 1.3: Use M different words in the natural text story to form a word set and design a matrix A = (a i,j ) 1≤i≤M,1≤j≤N' , when the ith word appears at the jth moment while the subject is listening to the story, then a i,j = 1, otherwise a i,j = 0;

[0095] Step 1.4: Discretely convolve the ith row vector a i,* of the matrix A with the hemodynamic response function (HRF), and obtain the input signal value x i of the ith word to the human brain at each moment from the matrix A designed in Step 1.3:

[0096]

[0097] In formula (2), b is the discrete value of the HRF function f = 0.452t 8.6 e -1.828t , and x i,j is the input signal value of the ith word to the human brain at the jth moment;

[0098] Step 1.5: Downsample x i according to the activation signal acquisition interval of fMRI to obtain the input signal value x' i of the ith word to the brain at N moments:

[0099] x' i =(x i,1 ,x i,2 ,…,x i,N ) T Equation (3)

[0100] Step 1.6: Construct the input feature matrix X based on equation (3):

[0101] X = (x'1, x'2, ..., x') i ,…,x' M ) T ,X∈R M×N Equation (4)

[0102] Step 1.7: Represent the j-th column vector of matrix X as x. j And compared with the output signal y of the extreme gradient boosting algorithm regression model in equation (1) j Forming training data pairs (x j ,y j If the extreme gradient boosting algorithm regression model has an input sample set of , then the input sample set of the extreme gradient boosting algorithm regression model is . The output sample set of the extreme gradient boosting algorithm regression model is No less than 80% of the sample set is designated as the training set, and the remaining sample set is used as the test set.

[0103] Step 1.8: Use the activation signal of any voxel as the output signal of the corresponding extreme gradient boosting algorithm regression model. With input signal A training sample set is constructed, an XGBoost regression model is built, K tree models f are trained, and the predictions of all tree models for the same sample are summed to obtain the final prediction.

[0104]

[0105] In equation (5), x i It is the i-th sample, f k (x i ) is the k-th tree pair of sample x i The predicted value, It is a function space. This is the final predicted value, and the corresponding true value is y. i ;

[0106] Step 1.9: Using the final predicted value from Step 1.8 and the true value y i Construct the loss function l, and after adding a regularization term Ω to control the model structure risk, construct the objective function Obj that needs to be optimized:

[0107]

[0108] Step 1.10: Using the sample set divided in Step 1.7 and the objective function Obj constructed in Step 1.9, train the extreme gradient boosting algorithm regression model corresponding to each voxel, that is, optimize the objective function Obj in Step 1.9 using the training data in Step 1.10.

[0109] Step 2 specifically includes:

[0110] Step 2.1: Use the TreeExplainer method in Shapley Additive exPlanations (SHAP) to decompose the output value of each sample in the extreme gradient boosting algorithm regression model, generating a symmetric sample interaction matrix. The length and width of the sample interaction matrix are consistent with the number of features. The diagonal values ​​of the sample interaction matrix reflect the contribution of a feature to the sample output individually, while the off-diagonal values ​​reflect the contribution of the interaction between two features to the sample output.

[0111] Step 2.2: Absolute-value the sample interaction matrices generated from all samples and sum them to obtain the final sample interaction matrix Mat. The elements of the sample interaction matrix Mat reflect the overall importance of all features and their interactions to the human brain activation signals.

[0112]

[0113] In equation (7), TreeE(x) i This refers to interpreting sample x using the TreeExplainer method. i This will generate a sample interaction matrix.

[0114] See Figure 2 (a) Figure 2 (b) Step 3 specifically includes:

[0115] Step 3.1: Based on the syntactic relationships in the sentence, divide the word pairs into various syntactic structures;

[0116] Step 3.2: Use the natural language text processing library (spaCy) to find all word pairs in the natural text that have syntactic structure relations, such as word pairs with the syntactic structure relation nsubj and word pairs with the syntactic structure relation amod.

[0117] Step 3.3: Each type of syntactic structure corresponds to a set of word pairs. For example, nsubj corresponds to the set of word pairs. All word pairs within it have an nsubj syntactic structure relation; amod corresponds to the set of word pairs. All word pairs within it exhibit amod syntactic structure relations;

[0118] Step 3.4: Pair any word with a given syntactic structure The average value of the interaction values ​​of all word pairs (i,j) in the sample interaction matrix Mat is considered as the functional strength P of the voxel corresponding to the interaction matrix Mat in processing syntactic structure.

[0119]

[0120] In equation (8), L is the number of word pairs belonging to the syntactic structure, i.e.

[0121] Step 3.5: After calculating different syntactic structures and obtaining the corresponding functional strength P, combine all functional strengths P into a syntactic vector. That is, each dimension of the vector represents the functional strength of a voxel when processing a certain syntactic structure.

[0122] Step 4 specifically includes:

[0123] Step 4.1: Superimpose the numerical values ​​of each dimension of the syntactic vector corresponding to the voxel; this gives the total syntactic functional strength of the voxel as a whole.

[0124] Total=∑Proc s ,s={nsubj,amod,…}Formula (9)

[0125] In equation (9), Total represents the syntactic functional strength of the voxels as a whole;

[0126] Step 4.2: Compare the values ​​of the same dimension of the syntactic vector corresponding to each voxel to obtain the functional strength when all voxels process the same syntactic structure;

[0127] Step 4.3: Display and compare the functional strengths obtained in Step 4.2 to obtain the functional brain regions when the human brain processes different syntactic structures, that is, to show which brain regions have greater voxel strength and which brain regions have smaller voxel strength.

[0128] See Figure 3 (a) Figure 3 (b) Step 5 specifically includes:

[0129] Step 5.1: Combine all word pairs that have syntactic structural relations into a set. Construct the adjacency matrix W of the syntactic network based on the sample interaction matrix Mat:

[0130]

[0131] Then, using the adjacency matrix W of each voxel, a syntactic network corresponding to each voxel is constructed. The nodes of the syntactic network are the words in all syntactic structures, and any two words have syntactic relationships.

[0132] Step 5.2: Based on the network transformation vector rules of the graph embedding algorithm (graph2vec), all syntactic networks are transformed into vectors to form a vector space. That is, different syntactic networks are mapped to different positions in the vector space. Two syntactic networks with similar positions represent that their network structures are similar, that is, the syntactic functions of the two voxels corresponding to the two syntactic networks are similar.

[0133] Step 5.3: Based on the classification rules of the K-means clustering algorithm, classify all syntactic networks in the vector space formed in Step 5.2, that is, obtain the classification of voxels corresponding to all syntactic networks. Specifically, the Euclidean distance metric is used to classify syntactic networks (voxels) that are close to each other in the vector space into the same category.

[0134] See Figure 3 (c) Step 6 specifically includes:

[0135] Step 6.1: Take the average value of all voxel signals in each region as the signal of that region, construct an extreme gradient boosting algorithm regression model for each region according to steps 1 to 3, and obtain the corresponding sample interaction matrix;

[0136] Step 6.2: Perform numerical statistics on the sample interaction matrix based on syntactic structure relations, and combine them into syntactic vectors corresponding to each region;

[0137] Step 6.3: Finally, calculate the Pearson correlation coefficient of the syntactic vectors corresponding to different regions to obtain the correlation between different regions, that is, their syntactic functional connection relationship. The larger the correlation coefficient, the greater their correlation.

[0138] A machine learning-based system for parsing human brain syntactic functions includes:

[0139] Sample set construction and modeling module: Constructs a sample set for the extreme gradient boosting algorithm regression model, and constructs the corresponding extreme gradient boosting algorithm regression model for each voxel using the sample set; this module is used for step 1 of the human brain syntactic function parsing method based on machine learning.

[0140] Shapley's additive interpretation theory module: Calculates the output signal of each extreme gradient boosting algorithm regression model using Shapley's additive interpretation theory to obtain the sample interaction matrix corresponding to the extreme gradient boosting algorithm regression model; this module is used in step 2 of the human brain syntactic function parsing method based on machine learning.

[0141] Numerical statistics module: By performing numerical statistics on the syntactic structural relationships of different words in the interaction matrix, the syntactic vector of each voxel is obtained; this module is used in step 3 of the human brain syntactic function parsing method based on machine learning.

[0142] Human brain syntactic function demonstration module: This module demonstrates the syntactic function mechanism of the human brain, including the overall syntactic function brain regions and the functional brain regions when processing different syntactic structures. This module is used in step 4 of the human brain syntactic function parsing method based on machine learning.

[0143] Syntactic network construction module: Constructs the syntactic network corresponding to voxels through the interaction matrix, and can classify voxels into regions based on the network structure; this module is used in step 5 of the human brain syntactic function parsing method based on machine learning.

[0144] Syntactic Function Connection Display Module: Displays the syntactic function connection relationships between different regions; this module is used in step 6 of the machine learning-based human brain syntactic function parsing method.

[0145] A machine learning-based device for parsing human brain syntactic functions includes:

[0146] Memory: for storing a computer program that implements the machine learning-based human brain syntactic function parsing method as described in any one of claims 1-7;

[0147] Processor: Used to implement the machine learning-based human brain syntactic function parsing method as described in any one of claims 1-7 when executing the computer program.

[0148] The processor can be a Central Processing Unit (CPU), or other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor can be a microprocessor, or any conventional processor. The processor is the control center of the device for the machine learning-based human brain syntactic function parsing method, connecting various parts of the device via various interfaces and lines.

[0149] When the processor executes the computer program, it implements the steps of the above-described machine learning-based human brain syntactic function parsing method.

[0150] Alternatively, when the processor executes the computer program, it implements the functions of each module in the above system, such as: sample set construction and modeling module; Shapley addition interpretation theory module; numerical statistics module; human brain syntactic function display module; syntactic function connection display module, and outputs the results of the human brain syntactic function parsing based on machine learning.

[0151] For example, the computer program can be divided into one or more modules / units, which are stored in the memory and executed by the processor to complete the present invention. The one or more modules / units can be a series of computer program instruction segments capable of performing preset functions, wherein the instruction segments describe the execution process of the computer program in the machine learning-based human brain syntactic function parsing device.

[0152] The device for machine learning-based human brain syntactic function parsing can be a desktop computer, laptop, handheld computer, or cloud server, among other computing devices. This device may include, but is not limited to, processors and memory. Those skilled in the art will understand that the above examples of machine learning-based human brain syntactic function parsing devices do not constitute a limitation on such devices. The device may include more components than described above, or combine certain components, or use different components. For example, the machine learning-based human brain syntactic function parsing device may also include input / output devices, network access devices, buses, etc.

[0153] The memory can be used to store the computer program and / or modules. The processor implements various functions of the machine learning-based human brain syntactic function parsing device by running or executing the computer program and / or modules stored in the memory and by calling the data stored in the memory.

[0154] The memory may primarily include a program storage area and a data storage area. The program storage area may store the operating system and at least one application program required for a given function (such as sound playback or image playback). The data storage area may store data created based on the use of the phone (such as audio data or a phonebook). Furthermore, the memory may include high-speed random access memory (RAM) and non-volatile memory, such as hard disks, RAM, plug-in hard disks, SmartMediaCards (SMC), Secure Digital (SD) cards, flash cards, at least one disk storage device, flash memory device, or other volatile solid-state storage devices.

[0155] The present invention also provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the steps of the machine learning-based human brain syntactic function parsing method.

[0156] If the system integration module / unit of the machine learning-based human brain syntactic function parsing method is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium.

[0157] This invention implements all or part of the processes in the above-described machine learning-based human brain syntactic function parsing method. It can also be accomplished by a computer program instructing related hardware. The computer program can be stored in a computer-readable storage medium. When executed by a processor, the computer program implements the steps of the above-described machine learning-based human brain syntactic function parsing method. The computer program includes computer program code, which can be in the form of source code, object code, executable file, or a preset intermediate form, etc.

[0158] The computer-readable storage medium may include: any entity or device capable of carrying the computer program code, recording media, USB flash drive, portable hard drive, magnetic disk, optical disk, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signal, telecommunication signal, and software distribution medium, etc.

[0159] It should be noted that the content contained in the computer-readable storage medium may be appropriately increased or decreased according to the requirements of legislation and patent practice in the jurisdiction. For example, in some jurisdictions, according to legislation and patent practice, the computer-readable storage medium does not include electrical carrier signals and telecommunication signals.

[0160] It should be noted that embodiments of the present invention can be implemented using hardware, software, or a combination of both. The hardware portion can be implemented using dedicated logic; the software portion can be stored in memory and executed by an appropriate instruction execution system, such as a microprocessor or dedicated hardware.

[0161] Those skilled in the art will understand that the above-described devices and methods can be implemented using computer-executable instructions and / or included in processor control code, for example, such code provided on a carrier medium such as a disk, CD, or DVD-ROM, a programmable memory such as read-only memory (firmware), or a data carrier such as an optical or electronic signal carrier. The devices and modules of the present invention can be implemented by hardware circuitry of semiconductors such as very large-scale integrated circuits or gate arrays, logic chips, transistors, etc., or programmable hardware devices such as field-programmable gate arrays, programmable logic devices, etc., or by software executed by various types of processors, or by a combination of the above-described hardware circuitry and software, such as firmware.

[0162] The above description is merely a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any modifications, equivalent substitutions, and improvements made by those skilled in the art within the scope of the technology disclosed in the present invention, and within the spirit and principles of the present invention, should be covered within the scope of protection of the present invention.

Claims

1. A method for parsing human brain syntactic functions based on machine learning, characterized in that, Includes the following steps: Step 1: Construct a sample set for the extreme gradient boosting algorithm regression model, and for each voxel, train the extreme gradient boosting algorithm regression model corresponding to that voxel; Step 2: Calculate the output signal of each extreme gradient boosting algorithm regression model using the Shapley addition method to obtain the sample interaction matrix corresponding to the extreme gradient boosting algorithm regression model. The specific steps are as follows: Step 2.1: Use the tree interpretation method in the Shapley method to decompose the output value of each sample of the extreme gradient boosting algorithm regression model, and generate a symmetric sample interaction matrix; Step 2.2: Convert the absolute values ​​of the sample interaction matrices generated by all samples and sum them to obtain the final sample interaction matrix. Sample interaction matrix The elements reflect the overall importance of all features and their interactions to the activation signals in the human brain: In equation (7), Refers to interpreting samples using the tree interpretation method. This will generate a sample interaction matrix; Step 3: Based on the syntactic structure relationships of different words, perform numerical statistics on the sample interaction matrix obtained in Step 2 to obtain the syntactic vector of each voxel; Step 4: Using the syntactic vectors obtained in Step 3, obtain the syntactic functional strength of each voxel, as well as the functional brain regions in the human brain when processing different syntactic structures. Step 5: Construct the syntactic network corresponding to the voxels using the sample interaction matrix obtained in Step 2, and perform region classification on the voxels based on the structure of the syntactic network. The specific steps are as follows: Step 5.1: Combine all word pairs that have syntactic structural relations into a set. Based on the sample interaction matrix Constructing the adjacency matrix of the syntactic network : Equation (10) Then through the adjacency matrix of each voxel Construct the syntactic network corresponding to each voxel; Step 5.2: Based on the graph embedding algorithm, the network transformation vector rules are used to transform all syntactic networks into vectors to form a vector space; Step 5.3: Based on the classification rules of the K-means clustering algorithm, classify and divide all syntactic networks in the vector space formed in Step 5.2; Step 6: Obtain the syntactic vectors corresponding to the regions classified in Step 5 by constructing an extreme gradient boosting algorithm regression model, and display the syntactic function connections of different regions.

2. The method for parsing human brain syntactic functions based on machine learning according to claim 1, characterized in that, Step 1 specifically includes: Step 1.1: Play natural text stories to the subjects and collect activation signals in their brains using functional magnetic resonance imaging (fMRI). Step 1.2: When the total duration of playing the natural text story to the subjects is... At that time, according to the activation signal acquisition interval of magnetic resonance imaging, voxels are placed in... ( The activation signal at time step () is used as the output signal of the regression model of the extreme gradient boosting algorithm. : In equation (1), For the first Output signal ; Step 1.3: Using natural text stories A word set composed of different words And design matrix When the first person listens to the story, The word is in When the moment occurs, then ,otherwise ; Step 1.4: Calculate the matrix The row vectors Discrete convolution with the hemodynamic response function, from the matrix designed in step 1.3 In the middle, it won the first The input signal value of each word to the human brain at each moment : In equation (2), HRF function discrete values, For the first The word is in The constant input signal values ​​to the human brain; Step 1.5, Downsampling was performed based on the activation signal acquisition interval of fMRI to obtain the first... The word is in The input signal value to the brain at any given moment : Step 1.6: Construct the input feature matrix based on equation (3) : Step 1.7: Calculate the matrix The Each column vector is represented as And compared with the output signal of the extreme gradient boosting algorithm regression model in equation (1) Composition of training data pairs Then the input sample set of the extreme gradient boosting algorithm regression model is The output sample set of the extreme gradient boosting algorithm regression model is No less than 80% of the sample set shall be designated as the training set, and the remaining sample set shall be designated as the test set. Step 1.8: Use the activation signal of any voxel as the output signal of the corresponding extreme gradient boosting algorithm regression model. , and input signal Construct a training sample set, build an XGBoost regression model, and train it. Tree model The predictions from all tree models for the same sample are then summed to obtain the final prediction. : Equation (5) In equation (5), It is the first One sample, It is the first Tree samples The predicted value, It is a function space. This is the final predicted value, and the corresponding actual value is... ; Step 1.9: Using the final predicted value from Step 1.8 and the true value Constructing the loss function and in the appended regular expression After controlling for model structural risks, construct the objective function to be optimized. : Step 1.10, the sample set partitioned in Step 1.7, and the objective function constructed in Step 1.

9. For each voxel, train the extreme gradient boosting algorithm regression model corresponding to that voxel.

3. The method for parsing human brain syntactic functions based on machine learning according to claim 1, characterized in that, Step 3 specifically includes: Step 3.1: Based on the syntactic relationships in the sentence, divide the word pairs into various syntactic structures; Step 3.2: Find all word pairs with syntactic structural relationships in the natural language text using a natural language text processing library; Step 3.3: Each type of syntactic structure corresponds to a set of word pairs. ); Step 3.4: Pair any word with a given syntactic structure All word pairs (i,j) in the set, in the sample interaction matrix The average value of the corresponding interaction values ​​is considered as the interaction matrix. The functional strength of the corresponding voxels in processing syntactic structures : In equation (8), L is the number of word pairs belonging to the syntactic structure, i.e., L = | |; Step 3.5: Calculate different syntactic structures to obtain the corresponding functional strength. Then, all functional strengths P are combined into a syntactic vector.

4. The method for parsing human brain syntactic functions based on machine learning according to claim 1, characterized in that: Step 4 specifically includes: Step 4.1: Superimpose the values ​​of each dimension of the syntax vector corresponding to the voxel: In equation (9), To reflect the overall syntactic functional strength of voxels; Step 4.2: Compare the values ​​of the same dimension of the syntactic vector corresponding to each voxel to obtain the functional strength when all voxels process the same syntactic structure; Step 4.3: Display and compare the functional strengths obtained in Step 4.2 to obtain the functional brain regions when the human brain processes different syntactic structures.

5. The method for parsing human brain syntactic functions based on machine learning according to claim 1, characterized in that: Step 6 specifically includes: Step 6.1: Take the average value of all voxel signals in each region as the signal of that region, construct an extreme gradient boosting algorithm regression model for each region according to steps 1 to 3, and obtain the corresponding sample interaction matrix; Step 6.2: Perform numerical statistics on the sample interaction matrix based on syntactic structure relations, and combine them into syntactic vectors corresponding to each region; Step 6.3: Finally, calculate the Pearson correlation coefficient of the syntactic vectors corresponding to different regions to obtain the correlation between different regions.

6. A human brain syntactic function parsing system based on machine learning, implemented using the method described in any one of claims 1-5, characterized in that, include: Sample set construction and modeling module: Constructs a sample set for the extreme gradient boosting algorithm regression model, and uses the sample set to construct the corresponding extreme gradient boosting algorithm regression model for each voxel; Shapley's additive interpretation theory module: Calculates the output signal of each extreme gradient boosting algorithm regression model using Shapley's additive interpretation theory, and obtains the sample interaction matrix corresponding to the extreme gradient boosting algorithm regression model; Numerical statistics module: By analyzing the syntactic structural relationships of different words, numerical statistics are performed in the interaction matrix to obtain the syntactic vector of each voxel; Human Brain Syntactic Function Demonstration Module: Demonstrates the syntactic function mechanism of the human brain, including the overall syntactic functional brain regions and the functional brain regions when processing different syntactic structures; Syntactic network construction module: Constructs syntactic networks corresponding to voxels through interaction matrices, and can classify voxels into regions based on the network structure; Syntactic Function Connection Display Module: Displays the syntactic function connection relationships between different regions.

7. A machine learning-based device for parsing human brain syntactic functions, characterized in that, include: Memory: for storing computer programs that implement the machine learning-based human brain syntactic function parsing method as described in any one of claims 1-5; Processor: Used to implement the machine learning-based human brain syntactic function parsing method as described in any one of claims 1-5 when executing the computer program.

8. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that, when executed by a processor, implements the steps of the machine learning-based human brain syntactic function parsing method as described in any one of claims 1-5.