Neuroscience literature processing method and system for constructing brain connection mapping knowledge domain
By combining the pointer annotation network and the position attention mechanism, the problem of complex brain region naming in neuroscience literature is solved, high-precision entity recognition and relationship extraction are achieved, and a directional brain connection knowledge graph is constructed to support efficient knowledge acquisition and analysis.
Patent Information
- Application Number
- CN202510673559.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-23
- Publication Date
- 2025-09-05
AI Technical Summary
The naming of brain regions in neuroscience literature is complex and varied. Existing methods have difficulty in automatically linking across databases, entity recognition is unstable, and the directionality of brain region connections cannot be accurately identified. Traditional knowledge graph construction methods do not provide directional information on brain region connections, resulting in low accuracy in knowledge integration.
An entity recognition method based on a pointer annotation network and a relationship extraction method that integrates fine-grained entity features and position attention are adopted. The model is trained to automatically identify entities and connection relationships in neuroscience literature, and the fuzzy matching algorithm is combined to achieve entity standardization and knowledge graph construction.
It significantly improves the accuracy of entity recognition and relationship extraction, realizes efficient knowledge integration and reasoning, provides more accurate knowledge acquisition and analysis tools for neuroscience research, and supports literature retrieval, brain region circuit analysis, and research hotspot identification.
Smart Images

Figure CN120596672A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to data processing in the field of neuroscience, and in particular to a neuroscience literature processing method and system for constructing a brain connectivity knowledge graph. Background Art
[0002] The brain is a complex network of interconnected neurons. A comprehensive understanding of brain connectivity is crucial to understanding its structure and function.
[0003] With the rapid development of neuroscience, researchers are systematically mapping brain connections across species, from nematodes to fruit flies to mice. This rich knowledge of brain connectivity provides an important foundation for a deeper understanding of the brain. However, this knowledge is often scattered across different literature and databases, and there are significant differences under different nomenclature systems. Systematically integrating this knowledge requires not only relying on expert experience, but also facing challenges such as a huge workload and frequent knowledge updates. When traditional methods extract knowledge from massive amounts of literature, it is often difficult to accurately and efficiently obtain brain connectivity knowledge using a small amount of expert-annotated data. Although knowledge graph technology provides a solution for integrating massive amounts of literature knowledge, how to quickly interact with literature through the constructed brain connectivity knowledge graph to achieve efficient literature retrieval, knowledge verification, and discovery remains a difficult problem in current research.
[0004] With the continuous advancement of neuroscience research, data-driven methods are becoming an important means of exploring brain connectivity. Researchers are delving into connectivity patterns between brain regions through various methods, including literature mining, neuroimaging analysis, and neuroanatomical experiments. In this context, knowledge graphs, as a structured knowledge representation method, have been widely used in fields such as biomedicine, artificial intelligence, and big data analytics.
[0005] Research methods based on knowledge graphs mainly include the following aspects:
[0006] First, entity recognition and relationship extraction: With the help of transfer learning technology, the language features of large-scale texts are learned by training the model, and fine-tuned in a variety of downstream tasks to achieve high-precision entity recognition and relationship extraction. For example, the invention patent with authorization announcement number CN112256828B discloses a medical entity relationship extraction method, device, computer equipment and readable storage medium, which uses a first model to perform medical named entity recognition on the data in the medical text to obtain the entity recognition results corresponding to each data to be processed; and based on the entity recognition results, entity relationship extraction is performed to obtain entity pairs with entity relationships, calculate the confidence of the entity pairs based on the entity relationships, and generate target data based on each entity pair, entity relationship and corresponding confidence, which solves the problem of time-consuming, labor-intensive and low efficiency of manual extraction of medical entity relationships in the existing technology. However, traditional entity recognition methods mostly use sequence labeling mechanisms (such as BIO / BIOES labeling schemes) combined with structures such as BERT and BiLSTM-CRF. These methods face significant bottlenecks when dealing with long vocabularies, nested entities, or entities with ambiguous boundaries. This is particularly true in neuroscience literature, where the naming of brain entities is complex and their boundaries are highly uncertain, making misidentification prone. Furthermore, mainstream relationship extraction methods are often based on classification frameworks, masking the positional information representing the entity to be identified and classifying a predefined set of relationships using special identifiers. These methods are generally unable to effectively model the name, position, or semantic dependencies between entities, and perform particularly poorly when dealing with cross-sentence, long-distance relationships, or when the corpus is unevenly distributed.
[0007] Second, knowledge graph construction: For example, the invention patent with publication number CN112541086A discloses a method for constructing a knowledge graph for stroke, which includes acquiring medical data; extracting medical entities, their attributes, and their relationships; designing an ontology model to complete knowledge graph ontology construction; and mapping entities based on the extracted medical entities, their attributes, and their relationships in combination with the ontology model to construct a knowledge graph for stroke. By constructing a knowledge graph for stroke, medical knowledge on stroke is integrated, resolving the issue of incomplete information, providing effective assistance in stroke diagnosis, and addressing the lack of medical knowledge among some doctors.
[0008] Although existing research has made significant progress in the automated extraction and analysis of knowledge, the systematic modeling and analysis of brain connectivity in current neuroscience research still faces the following key technical bottlenecks:
[0009] 1. The terminology in neuroscience literature is complex and varied, with significant differences in nomenclature across species, literature, and databases. The aforementioned methods do not integrate multiple types of terms as search keywords for medical texts, and thus cannot support automatic linking between brain region nomenclatures across databases. Furthermore, traditional rule-based or shallow machine learning methods struggle to accurately identify entity boundaries, resulting in unstable entity recognition results.
[0010] 2. In existing technologies, single-label classification is often used for relationship prediction. This is only applicable to associations between any two entity categories generated based on simple dependency relationships. It does not fully utilize entity names and location information, and cannot accurately identify the relative positions and directional relationships of entities. As a result, when identifying brain region connectivity, the model has difficulty accurately distinguishing the association weights of different entities.
[0011] 3. Existing knowledge graph construction methods do not provide directional information about brain region connections. The connections between brain regions are usually directional, but most current methods focus on whether connections exist between brain regions, while ignoring the input and output information of neural circuits.
[0012] 4. Since different databases and documents use different naming systems, the existing entity linking method is difficult to effectively solve the problem of cross-system matching, resulting in low accuracy of knowledge integration. When constructing knowledge graphs, cross-naming system mapping is difficult. Summary of the Invention
[0013] Therefore, in order to solve the above problems, the present invention provides a neuroscience literature processing method and system for constructing a brain connectivity knowledge graph. This method constructs a brain connectivity knowledge graph based on an entity recognition method of a pointer annotation network and a relationship extraction method that integrates fine-grained entity features and position attention. It can significantly improve the accuracy of entity recognition and relationship extraction, while achieving efficient knowledge integration and reasoning, providing more accurate and efficient knowledge acquisition and analysis tools for neuroscience research.
[0014] The present invention is achieved through the following technical solutions:
[0015] The neuroscience literature processing method for constructing a brain connectivity knowledge graph includes the following steps:
[0016] Obtaining the first model based on a pointer annotation network, trained on a large-scale neuroscience corpus, for automatically identifying entities in neuroscience literature;
[0017] Combining the entity results identified by the first model and introducing a positional attention mechanism to obtain a second model, which is used to identify entity connections in neuroscience literature and automatically determine the directionality of the connections;
[0018] Screen neuroscience literature from multi-source literature databases and standardize the literature formats to obtain standardized literature;
[0019] Identifying entities in the standardization document based on the first model, and extracting connection relationships and directionality of entities in the standardization document based on the identified entities using the second model;
[0020] The extracted entities are mapped to the standard ontology library in this field using fuzzy matching algorithm, and a unified entity ontology library is constructed by integrating multiple ontology libraries;
[0021] The connection relationships between entities are stored in the form of triples in a graph database, and a directed brain connection knowledge graph is constructed. At the same time, the selected neuroscience literature and its structured information are stored in a document database and a relational database respectively.
[0022] Preferably, the first model adopts an entity recognition model based on a pointer annotation network, and the entity recognition model includes:
[0023] Input layer: used to accept neuroscience corpus containing entity annotation information;
[0024] Data preprocessing layer: used to standardize the corpus, using BIO annotation method and using dictionary tree to correct boundary inconsistencies;
[0025] Pointer annotation network layer: used to reallocate BIO labels and start and end position pointers to the sequence after word segmentation, and train based on the hybrid cross entropy loss function;
[0026] Output layer: Output entity recognition results, including text sequences and their corresponding label sequences.
[0027] Preferably, the second model adopts a relationship extraction model that integrates the entity features identified by the entity recognition model and the position attention mechanism, and the relationship extraction model includes:
[0028] Input layer: used to input corpus containing entity before and after position identifier information;
[0029] Position Attention Network Layer: This layer fine-tunes the pre-trained parameters of the entity recognition model, selects the entity's preceding and following position identifiers as position attention features, and combines convolutional and pooling layers to learn entity and position features to generate high-quality feature vectors.
[0030] Output layer: used to output relationship classification results, supporting binary or multi-classification tasks.
[0031] Preferably, the neuroscience corpus received by the input layer of the entity recognition model includes a gold standard corpus of neurons and neurotransmitters constructed by manual back-to-back annotation and a gold standard corpus of brain regions constructed from the existing project WhiteText; the corpus containing entity location information received by the input layer of the relationship extraction model is a brain region connection relationship corpus with directional information constructed through rules and manual back-to-back annotation.
[0032] Preferably, the multi-source literature database includes at least a PubMed database and a PMC database, and the "collecting neuroscience literature from the multi-source literature database" means: screening out PubMed literature abstracts related to brain regions from the PubMed database based on keywords, and screening out PMC literature full texts related to brain regions from the PMC database, and obtaining relevant literature data.
[0033] Preferably, the terms defined in the Allen ontology database, the BAMS ontology database, the NeuroNames ontology database, and the NeuroLex ontology database are used as keywords to screen out PubMed literature abstracts related to brain regions from the PubMed database, and screen out PMC literature full texts related to brain regions from the PMC database. The screened literature data at least includes title, abstract, author, journal name, PMID, and publication time.
[0034] Preferably, the screened neuroscience literature is stored in a document-type database in CSV format, and the structured information of the neuroscience literature is stored in a relational database in JSON format.
[0035] Preferably, the triples used to represent entity connection relationships are stored in a graph database in the form of directed edges, and the graph database is a Neo4j graph database.
[0036] Preferably, application analysis is also carried out based on the brain connection knowledge graph, and the application analysis at least includes: intelligent literature retrieval, input and output loop analysis of specific brain regions, author cooperation network analysis and research hotspot identification.
[0037] Brain connection knowledge graph system, including:
[0038] The data input layer is used to obtain information about brain connectivity from multi-source literature data. Terms defined in the Allen ontology database, BAMS ontology database, NeuroNames ontology database, and NeuroLex ontology database are used as keywords. PubMed literature abstracts related to brain regions are screened from the PubMed database, and PMC literature full texts related to brain regions are screened from the PMC database. Literature data including title, abstract, author, journal name, PMID, and publication date are obtained.
[0039] The knowledge acquisition layer includes an entity recognition module and a relationship extraction module. The entity recognition module uses an entity recognition model based on a pointer annotation network to automatically identify entities in neuroscience literature. The relationship extraction module combines the entity output results of the entity recognition module with a positional attention mechanism to extract entity connection relationships in neuroscience literature, automatically determine the connection direction, and output structured relationship data.
[0040] The knowledge graph evaluation layer includes an entity standardization module and a knowledge storage module. The entity standardization module uses a fuzzy matching algorithm to map the extracted entities to a standard ontology library in the field, and builds a unified entity ontology library by integrating multiple ontology libraries. The knowledge storage module stores the connection relationships of each entity in the form of triples in a graph database and constructs a directed brain connection knowledge graph. At the same time, the selected neuroscience literature and its structured information are stored in a document database and a relational database respectively.
[0041] The knowledge graph retrieval application layer includes a literature retrieval module, a brain region circuit analysis module, and an author community analysis module. The literature retrieval module adopts the Vue front-end framework and the Flask back-end service architecture, and realizes the visualization of the knowledge graph based on the Relation-Graph; the brain region circuit analysis module is used to generate a whole-brain input and output connection path diagram for a specific brain region; the author community analysis module constructs a community map based on the relationship between literature authors and brain regions, and uses the Leuven community detection algorithm to perform community division and identify cooperative relationships and research hotspots between research teams.
[0042] The beneficial effects of the technical solution of the present invention are mainly reflected in:
[0043] 1. The present invention draws on transfer learning technology to learn the language features of large-scale texts through a biomedical language model trained based on a pointer annotation network, and obtains lexical semantic information and grammatical dependencies. On this basis, by introducing a pointer network, boundary word classification is combined with sequence classification, and the start and end positions of sequence labels and entities are predicted at the same time, thereby improving the accuracy of entity boundary recognition and the recognition ability of long entities and nested entities. It has a significant improvement in the accuracy of entity boundary recognition and relationship classification, significantly alleviating problems such as inaccurate boundary prediction and relationship recognition errors, and showing stronger robustness and accuracy in the recognition of complex named entities in neuroscience.
[0044] 2. The present invention designs a relationship extraction model that integrates fine-grained entity features and position attention mechanism. The model explicitly introduces: word-level representation and type encoding of entities; entity name, relative position and relationship weight features; and integrates attention mechanism to strengthen the interaction modeling between entity semantics and context. Among them, the pointer annotation network provides high-precision, noise-resistant entity boundary and type information for the relationship extraction model, while the position attention network uses this information to strengthen the contextual interaction and spatial dependency modeling between entities. Compared with the existing relationship extraction model, it significantly improves the recognition ability of long-distance dependency relationships, and can effectively cope with the challenges of uneven distribution of relationship labels and unbalanced sample size in the corpus, achieve better performance in the brain area direction relationship extraction task, and has better generalization ability.
[0045] 3. Before training the entity recognition model and the relationship extraction model, to accurately identify various types of entities, the entity recognition model was trained based on a gold standard corpus of neurons and neurotransmitters constructed using manual back-to-back annotation, as well as a gold standard corpus of brain regions constructed from the existing project WhiteText. This solved the problem of inaccurate automatic recognition of such entities in the literature. At the same time, to accurately identify directional relationships between brain regions, a corpus of brain region connection relationships with directional information was constructed through rules and manual back-to-back annotation. The relationship extraction model was trained based on this corpus of brain region connection relationships with directional information, solving the problem that traditional methods cannot accurately identify directional relationships between brain regions.
[0046] 4. When screening neuroscience literature, we use terms defined in various brain region ontology libraries as keywords to automatically extract entities such as brain regions, neurons, and neurotransmitters from massive neuroscience literature. We also standardize the format of the literature content and construct a text dataset with a unified structure to ensure data consistency and scalability in subsequent processing steps.
[0047] 5. The present invention adopts fuzzy matching and word vector similarity calculation methods. The fuzzy matching algorithm adds a weighted word matching method on the basis of accurate matching, so that more qualified entities can be mapped to the ontology, realizing the standardization and linking of entities. The extracted neuroscience entities are mapped to the standard ontology library through the fuzzy matching algorithm to achieve the unification of naming specifications, solve the naming ambiguity and inconsistency problems in different documents or databases, and thus improve the consistency and standardization of entities in the knowledge graph.
[0048] 6. The application analysis of knowledge graphs supports literature retrieval, supports comprehensive exploration of the connection relationships between complex brain regions, provides efficient information integration and intuitive visualization, and helps researchers discover new connection patterns and research trends. At the same time, the input and output loop analysis of specific brain regions can be used to analyze literature data of any brain region in the entire brain, quickly locate the input and output loop information of the brain region, research hotspots, and neurons, neurotransmitters, genes and diseases related to the brain region, etc. It can not only systematically reflect the research hotspots of different author groups in the literature, but also identify important research directions, reveal potential research gaps, and provide valuable references for the future development of neuroscience research.
[0049] 7. The brain connectivity knowledge graph system achieves efficient organization and utilization of knowledge, as well as standardized integration, visualization and intelligent analysis of brain connectivity information through the collaborative cooperation of four modules: data acquisition, knowledge extraction, knowledge graph evaluation, and knowledge graph retrieval and application. This helps promote the development of related applications such as research on brain area connectivity mechanisms, functional analysis of brain circuits, and literature-driven brain map construction. BRIEF DESCRIPTION OF THE DRAWINGS
[0050] Figure 1 This is a flowchart of the neuroscience literature processing method used to build a brain connectivity knowledge graph;
[0051] Figure 2 It is a flowchart for constructing the brain connection knowledge graph;
[0052] Figure 3 This is a schematic diagram of the architecture of the first model;
[0053] Figure 4 This is a schematic diagram of the architecture of the second model;
[0054] Figure 5 It is a flowchart of a neuroscience literature processing method for constructing a brain connectivity knowledge graph, which includes obtaining standardized literature and identifying entities, entity connection relationships, and directional information in the standardized literature;
[0055] Figure 6 It is a statistical diagram of the number of different naming systems in the process of integration; Figure 6(A) shows the quantitative changes during the integration of different nomenclature systems in the BAMS ontology library, and is distinguished by mouse and rat. Figure 6 (B) is the quantitative change during the integration of different naming systems in the NeuroNames ontology database;
[0056] Figure 7 It is a search interface for the brain-connected knowledge graph; Figure 7 (A) is the search result of VTA brain region in the brain connection knowledge graph retrieval system; Figure 7 (B) Shows a summary of literature showing connectivity with the VTA brain region; Figure 7 (C) Displays the annotation information of various entities that have a connection relationship with the VTA brain area;
[0057] Figure 8 It is the PVH brain area output loop diagram generated by the brain area connection knowledge graph;
[0058] Figure 9 It is a hotspot map of brain regions that different author groups are focusing on; Figure 9 (A) The brain region research hotspot network of different author groups; Figure 9 (B) Figure 9 (A) The author groups in the box and the brain regions they mainly studied; Figure 9 (C) The top three brain regions studied by different author groups and the corresponding number of studies;
[0059] Figure 10 1 is a flowchart of a neuroscience literature processing method for constructing a brain connectivity knowledge graph (including step S7). DETAILED DESCRIPTION
[0060] To more clearly and in detail illustrate the objectives, advantages, and features of the present invention, the following non-limiting description of preferred embodiments is provided for illustration and explanation. This embodiment is merely a typical example of the application of the technical solution of the present invention. Any technical solution formed by equivalent substitution or equivalent transformation falls within the scope of protection claimed by the present invention.
[0061] It is also stated that the terms "first" and "second" in this solution are used for descriptive purposes only and should not be understood to indicate or imply a ranking of importance or implicitly specify the number of technical features shown. Therefore, features defined as "first" or "second" may explicitly or implicitly include one or more of such features. In the present invention, "plurality" means two or more, unless otherwise specifically defined.
[0062] The present invention discloses a neuroscience literature processing method for constructing a brain connectivity knowledge graph, comprising the following steps:
[0063] Step S1: Obtain a first model based on a pointer annotation network, which has been trained using a large-scale neuroscience corpus for automatically identifying entities in neuroscience literature.
[0064] Among them, the entities include brain regions, neurons, neurotransmitters, etc. in neuroscience literature; the first model is a training model based on a pointer annotation network. The pointer annotation network model is suitable for processing long texts and complex entities with nested structures. It can significantly improve the boundary recognition ability of entities such as brain regions, neurons, and neurotransmitters. By training on a large-scale neuroscience corpus, it can effectively capture different types of neuroscience entities, improve the boundary division accuracy and overall extraction performance of complex entities, automatically identify entities such as brain regions, neurons, and neurotransmitters in the literature, ensure high accuracy of entity extraction, and lay the foundation for subsequent entity standardization and relationship extraction.
[0065] like Figure 3 As shown, in some embodiments, the first model adopts an entity recognition model based on a pointer annotation network, and the entity recognition model includes:
[0066] Input layer: used to accept neuroscience corpus containing entity annotation information;
[0067] Data preprocessing layer: used to standardize the corpus, using BIO annotation method and using dictionary tree to correct boundary inconsistencies;
[0068] Specifically, BIO tags (such as B-brain region, I-brain region, O) are used to annotate entities to ensure that each token belongs to only one entity or non-entity. At the same time, a dictionary tree is used to store all brain region entities in the ontology set. Then, each entity in the corpus is checked sequentially through the dictionary tree, and its annotation times and total number of occurrences are recorded. Subsequently, in the secondary check process, the dictionary tree adds or deletes entities with high or low matching frequencies according to the set threshold to eliminate the annotation noise in the corpus while retaining a certain degree of fault tolerance.
[0069] Pointer annotation network layer: used to reallocate BIO labels and start and end position pointers to the sequence after word segmentation, and train based on the hybrid cross entropy loss function;
[0070] Specifically, for each token, its probability distribution on the BIO label and the probability of its starting and ending positions as an entity are predicted respectively, and then training is performed based on a hybrid cross-entropy loss function, wherein the hybrid cross-entropy loss function includes BIO classification loss and pointer classification loss. The BIO classification loss adopts standard cross entropy to measure the difference between the BIO label distribution predicted by the model and the true label; the pointer classification loss adopts binary cross entropy loss for two binary classification tasks for the starting and ending positions; the total loss can then be calculated based on the BIO classification loss and the pointer classification loss; the specific training steps of the hybrid cross entropy loss function can refer to the existing technology and will not be repeated here.
[0071] Output layer: Output entity recognition results, including text sequences and their corresponding label sequences.
[0072] In some embodiments, the neuroscience corpus received by the input layer of the entity recognition model is a gold standard corpus of neurons and neurotransmitters constructed using a manual back-to-back annotation method; in addition, the input layer of the entity recognition model also includes a gold standard corpus of brain regions obtained from the existing project WhiteText.
[0073] Step S2: Combining the entity results identified by the first model and introducing the position attention mechanism to obtain a second model, the second model is used to identify the entity connection relationship in neuroscience literature and automatically determine the connection direction.
[0074] like Figure 4 As shown, in some embodiments, the second model adopts a relationship extraction model that integrates the entity features recognized by the entity recognition model and the position attention mechanism, and the relationship extraction model includes:
[0075] Input layer: used to input corpus containing entity before and after position identifier information;
[0076] In one embodiment, the corpus containing entity location information received by the input layer of the relationship extraction model is a brain region connection relationship corpus with directional information constructed through rules and manual back-to-back annotation; wherein, the input layer splits the original content of the neuroscience literature, and annotates the starting and ending positions of the two entities of the relationship to be predicted for each instance sentence of the neuroscience literature, so as to facilitate the addition of special identifiers before and after each entity in the subsequent process, and then use these four position identifiers to predict the relationship category.
[0077] Position Attention Network Layer: This layer fine-tunes the pre-trained parameters of the entity recognition model, selects the entity's preceding and following position identifiers as position attention features, and combines convolutional and pooling layers to learn entity and position features to generate high-quality feature vectors.
[0078] Among them, the position attention network focuses on key modifiers by calculating the interaction weight between entity pairs and each context word; first, the model combines convolutional layers and pooling layers to work together, and through multi-stage feature extraction and fusion, generates a high-quality feature vector that integrates local semantics and position awareness; among them, the convolutional layer first slides within the context window around the entity, using the convolution kernel to capture local dependency patterns, such as connectives or directional prepositions between entities, and outputs a multi-channel feature map, where each channel corresponds to a local semantic pattern; then, the pooling layer (usually maximum pooling) reduces the dimensionality of the convolution feature map, retains the most significant feature response in each channel, filters out redundant noise, and forms a compact representation of the entity's local context; in this process, the convolution-pooling module explicitly encodes the relative position information (such as distance, direction) of the entity pair. Finally, the pooled local features are concatenated with the entity type embedding and the global attention context vector, and then fused through the fully connected layer. The generated feature vector simultaneously contains fine-grained local semantics, entity type, and the relative position and relationship direction of the entity.
[0079] In a preferred embodiment of the present invention, in the position attention network layer, the four position vectors h obtained by the model through self-attention learning 11 、h 12 、h 21 、h 22 , these position vectors record the position features and the preceding and following semantic features of the entities in the sentence. These features are the key to relationship classification. When making relationship classification predictions, the model extracts the features of these four position vectors to enhance the classification effect. In order to fully extract the information in each position vector, these vectors are first convolved and pooled to fuse the features. The specific steps are as follows:
[0080] Convolutional layer: For each position vector h ij , use the convolution kernel to gradually extract local features from the starting position within a specific window size, and the features extracted by the convolution kernel are represented as H i , the feature vector obtained after the convolution operation is expressed as This operation can capture local continuous semantic information in the position vector and improve the model's ability to perceive the position;
[0081] Pooling layer: Next, the convolutional features are compressed through the maximum pooling layer to extract the most important features. The vector obtained by the pooling operation is represented as V i , represents the streamlined features obtained after pooling of each position vector;
[0082] Feature fusion: The vector V after pooling iThey are concatenated to form a final feature vector V, which contains the fusion features of all position vectors and is the result of learning through the position attention mechanism.
[0083] Output layer: used to output relationship classification results, supporting binary or multi-classification tasks;
[0084] like Figure 4 As shown in the figure, the pre-trained parameters of the entity recognition model are transferred to the relationship extraction model. The pre-trained parameters include the names of several entities appearing in neuroscience literature and their relative positions. These parameters are optimized through fine-tuning. At the same time, the position attention mechanism is introduced, and the directional relationship of each entity is extracted using the relationship extraction model. The entity and position features are learned in combination with the convolutional layer and the pooling layer. The sensitivity to directional words is enhanced in the convolutional layer, thereby outputting the relationship classification result. By fusing entity features with the position attention mechanism, the recognition ability of brain area connection relationships and directionality can be improved, and the model's perception of the relative positions of entities and relationship directions in the text can be enhanced, thereby adapting to the contextual changes of complex neuroscience literature.
[0085] Step S3: Screen neuroscience literature from the multi-source literature database and standardize the literature format to obtain standardized literature.
[0086] The multi-source literature database in this method is based on neuroscience literature data. In some embodiments, the multi-source literature database includes at least PubMed database and PMC database, which are important sources of neuroscience knowledge. The "collecting neuroscience literature from the multi-source literature database" means: screening out PubMed literature abstracts related to brain regions from the PubMed database based on keywords, and screening out PMC literature full texts related to brain regions from the PMC database, such as Figure 5 As shown, in one embodiment, the E-utilities API is used to screen out 1.35 million abstracts related to brain regions from 37 million abstracts in PubMed, and 190,000 full texts related to brain regions are screened out from 9.8 million full texts in PMC; wherein, the "standardization of the document format" means: the screened documents are uniformly stored in CSV format, and the structured information in the documents are uniformly stored in JSON format.
[0087] In a preferred embodiment, during the standardized mapping process of entities, multiple ontology libraries related to the entities need to be added. Commonly used ontology databases include Allen, NeuroNames, BAMS, and NeuroLex. These databases contain key terms and synonyms in the field of neuroscience. The Allen ontology contains 927 words and synonyms related to mouse brain regions. The BAMS and NeuroNames ontology libraries cover the naming conventions of brain regions of multiple species, such as Swanson-1992, Swanson-1998, Hof-2000, Paxinos-2001, and Dong-2007. NeuroLex, as a semantic wiki system, includes more than 25 000 unique entities in the field of neuroscience, covering terms in multiple aspects such as brain regions, neurons, neurotransmitters, behavioral paradigms and anatomy. These ontology databases ensure the standardization and consistency of entity names; among them, the terms defined in the Allen ontology database, BAMS ontology database, NeuroNames ontology database and NeuroLex ontology database are used as keywords, and PubMed literature abstracts related to brain regions are screened from the PubMed database, and PMC literature full texts related to brain regions are screened from the PMC database. The screened literature data at least contains key information such as title, abstract, author, journal name, PMID and publication time. This information ensures the comprehensiveness and richness of the literature data.
[0088] Step S4: identifying entities in the standardized document based on the first model, and extracting the connection relationship and directionality of each entity in the standardized document through the second model based on the identified entities.
[0089] like Figure 5 As shown in the figure, it is the specific acquisition process of entities and their connection relationships and directions. Entity recognition uses the entity recognition model and PTC system to extract neuroscience entities and biomedical entities in the literature; relationship extraction uses the relationship extraction model to obtain brain region connection relationships, direction relationships, and co-occurrence relationships between brain regions and other entities from the literature.
[0090] Step S5: Use a fuzzy matching algorithm to map the extracted entities to the standard ontology library in this field, and build a unified entity ontology library by integrating multiple ontology libraries.
[0091] Since there is no unified ontology in neuroscience, for example, entity information may come from multiple ontology libraries, and neurons may be named according to brain regions, transmitters, morphology, function or conduction mode. Therefore, the naming ambiguity and inconsistency problems in different literature or databases require the integration of multiple naming systems. Among them, the fuzzy matching algorithm supports the mapping of brain region naming systems across databases, improving the automation level and matching efficiency of entity standardization. Specifically, the fuzzy matching algorithm adds a weighted word matching method on the basis of accurate matching, so that more qualified entities can be mapped to the ontology; for each word W to be matched, it is segmented into <w1,w2,...,w n >, then gradually apply the exact matching and fuzzy matching algorithms, and obtain the matching score through four steps: whole word matching, matching without direction words, matching without stop words, and matching using SciSpacy. In fuzzy matching, the algorithm will further decompose the words into subwords and remove special symbols. By calculating the matching degree between the subwords and the ontology vocabulary and the weight of synonym matching, the ontology identifier with the highest score is finally selected. Figure 6 As shown, in one embodiment, a unified brain region ontology database is constructed by integrating entities whose abbreviations and full names are consistent to improve the consistency and standardization of entities in the knowledge graph, which contains a total of 2354 mouse and rat brain region records, including unique IDs, full names, abbreviations, synonyms and description information; Figure 6 (A) shows the quantitative changes during the integration of different nomenclature systems in the BAMS ontology library, where different nomenclature systems in the BAMS ontology library are distinguished by mouse and rat; Figure 6 (B) is the quantitative change during the integration of different naming systems in the NeuroNames ontology library.
[0092] Step S6: Store the entity connection relationships in the form of triples in the graph database, and construct a directed brain connection knowledge graph. At the same time, store the selected neuroscience literature and its structured information in the document database and relational database respectively.
[0093] In one embodiment, the extracted entity connection relationship is stored in the form of a triple of (brain area A, connection relationship, brain area B), wherein the triple used to represent the entity connection relationship is stored in the form of a directed edge in a graph database, and the graph database is a Neo4j graph database, which constructs a directed relationship graph to facilitate subsequent query and analysis, ensuring efficient storage and access of data; in addition, the screened neuroscience literature is stored in a document database in CSV format, and the structured information of the neuroscience literature is stored in a relational database in JSON format, wherein the CSV format uses plain text delimiters to facilitate the rapid processing of small data sets, and the JSON format facilitates the storage of complex nested structured data, and is compatible with document databases such as MongoDB for flexible query, thereby achieving efficient storage and rapid access to complex data sets.
[0094] like Figure 10 As shown, the neuroscience literature processing method for constructing a brain connectivity knowledge graph also includes step S7: conducting application analysis based on the brain connectivity knowledge graph; the application analysis includes but is not limited to: intelligent literature retrieval, input and output loop analysis of specific brain regions, author collaboration network analysis and research hotspot identification.
[0095] The literature search is designed to support the comprehensive exploration of complex brain region connectivity, provide efficient information integration and intuitive visualization, and help researchers discover new connectivity patterns and research trends. The front-end interface is built using the Vue framework, the back-end service is implemented using the Flask framework, and the dynamic presentation of the knowledge graph is achieved through the Relation-Graph force-directed graph. It supports multi-type brain region connectivity analysis, full-text literature retrieval, and highlighting of multi-type neuroscience entities. The retrieval interface of the knowledge graph is as follows: Figure 7 As shown in , after entering the name of a brain region, the link relationship between the current brain region and other brain regions, the literature abstracts that have a connection relationship with the current brain region, and the entity annotation information will be displayed respectively; Figure 7 (A)- Figure 7 (C) Figure 7 (A) Shows the search results for the VTA brain region in the brain connectivity knowledge graph retrieval system; Figure 7 (B) Shows a summary of literature showing connectivity with the VTA brain region; Figure 7 (C) Displays the annotation information of various entities that have connection relationships with the VTA brain area.
[0096] The specific brain region input-output loop analysis can be used to analyze literature data of any brain region in the entire brain, quickly locate the brain region's input-output loop information, research hotspots, and neurons, neurotransmitters, genes, and diseases related to the brain region; this information provides important literature prior knowledge for brain region research, helping researchers to quickly and accurately acquire input-output loop knowledge.
[0097] At the same time, when verifying the results of biological experiments, prior knowledge based on literature can demonstrate existing research results and provide a reliable reference for experimental comparison and verification; Figure 8 As shown, the output loop diagram of the PVH brain region generated by the brain connection knowledge graph is displayed. Brain regions of the same category are marked with the same color, such as dark pink in the midbrain, red in the hypothalamus, and green in the olfactory bulb. In the study of the whole-brain output loop of the PVH, it was found that the PVH brain region mainly projects to the VTA, EW, and PAGvl in the midbrain; and mainly projects to the NTS, DMX, and MDRN in the medulla. In existing studies, the study by Geerling et al. also showed the output loop information of the PVH brain region, and most of the loops have been verified in the brain connection knowledge graph.
[0098] However, due to differences in the naming systems of some brain regions, the corresponding connection knowledge cannot be found; for example, the brain regions RTN and CPA do not appear in the brain region dictionary, so the corresponding connection knowledge cannot be found; for another example, Caudal C1 Catecholamine Neuron is classified as a neuron rather than a brain region in entity recognition, so the connection knowledge cannot be obtained.
[0099] The author community group analysis reveals the research focus of different author groups through knowledge graphs and analyzes the changes in research teams' attention to brain regions during a specific period. This analysis method can not only systematically reflect the research hotspots of different author groups in the literature, but also identify important research directions, reveal potential research gaps, and provide valuable references for the future development of neuroscience research.
[0100] In the statistics of the relationship between authors and brain regions, in order to improve the accuracy of the analysis, a minimum threshold for the frequency of occurrence of each node can be set to filter out authors and brain region information with fewer occurrences. The minimum threshold for the frequency of occurrence of each node can be adjusted according to actual needs and will not be elaborated here. In one embodiment, 177 author nodes and 126 brain region nodes were finally obtained, reflecting the core authors and important brain regions active in mouse neuroscience research.
[0101] like Figure 9As shown in (A), the Leuven community discovery algorithm was used to divide the author nodes, resulting in 17 different communities with a modularity parameter of 0.883. This indicates that each community has good cohesion and can better reflect the cooperative relationship or common research interests among author groups. For example, communities 7 and 15 are led by Schachner M and Watanabe M, respectively. They have published a large number of research papers on brain regions in the field of neuroscience, with 345 and 293 studies respectively. Further analysis of community 7 revealed that Schachner M, Wurst W, Yanagawa Y and others belong to the same community, indicating that they have close cooperation or overlapping research interests.
[0102] By counting the research hotspots of different author groups, we can reveal the main research directions of each community, such as Figure 9 As shown in (B), the brain regions that Community 5 mainly focuses on include the hippocampus, cerebral cortex, and striatum, which indicates that the group has a strong interest in studying these brain regions.
[0103] Figure 9 (C) Shows the research hotspots of each community. Specifically, community 15 focuses on CBX, CB, and HIP, with 892, 600, and 599 studies conducted, respectively; while community 13 pays more attention to PH, HY, and ARH, with 161, 147, and 101 studies conducted, respectively; among all communities, the most studied brain regions are HPF, CBX, and CTXpl, which play an important role in mouse neuroscience research; by analyzing the author groups, researchers can more intuitively understand the distribution of interests of different research groups and identify brain regions that may become the focus of research in the future.
[0104] The present invention also discloses a brain connection knowledge graph system, comprising:
[0105] The data input layer is used to obtain information about brain connectivity from multi-source literature data. Terms defined in the Allen ontology database, BAMS ontology database, NeuroNames ontology database, and NeuroLex ontology database are used as keywords. PubMed literature abstracts related to brain regions are screened from the PubMed database, and PMC literature full texts related to brain regions are screened from the PMC database. Literature data including title, abstract, author, journal name, PMID, and publication date are obtained.
[0106] The knowledge acquisition layer includes an entity recognition module and a relationship extraction module. The entity recognition module uses an entity recognition model based on a pointer annotation network to automatically identify entities in neuroscience literature. The relationship extraction module combines the entity output results of the entity recognition module with a positional attention mechanism to extract entity connection relationships in neuroscience literature, automatically determine the connection direction, and output structured relationship data.
[0107] The knowledge graph evaluation layer includes an entity standardization module and a knowledge storage module. The entity standardization module uses a fuzzy matching algorithm to map the extracted entities to a standard ontology library in the field, and builds a unified entity ontology library by integrating multiple ontology libraries. The knowledge storage module stores the connection relationships of each entity in the form of triples in a graph database and constructs a directed brain connection knowledge graph. At the same time, the selected neuroscience literature and its structured information are stored in a document database and a relational database respectively.
[0108] The knowledge graph retrieval application layer includes the literature retrieval module, the brain circuit analysis module, and the author community analysis module. Figure 7 As shown in the figure, the literature retrieval module adopts the Vue front-end framework and the Flask back-end service architecture, and realizes the visualization of the knowledge graph based on the Relation-Graph; Figure 8 As shown, the brain region circuit analysis module is used to generate a whole-brain input-output connection path diagram for a specific brain region; Figure 9 As shown, the author community analysis module constructs a community map based on the relationship between literature authors and brain regions, and uses the Leuven community detection algorithm to divide communities and identify cooperative relationships and research hotspots between research teams.
[0109] There are many implementation methods of the present invention, and all technical solutions formed by equivalent transformation or equivalent transformation fall within the protection scope of the present invention.
Claims
1. A neuroscience literature processing method for constructing a brain connectivity knowledge graph, characterized by: The following steps are involved: Obtaining the first model based on a pointer annotation network, which has been trained on a large-scale neuroscience corpus, for automatically identifying entities in neuroscience literature; Combining the entity results identified by the first model and introducing a positional attention mechanism to obtain a second model, which is used to identify entity connections in neuroscience literature and automatically determine the directionality of the connections; Screen neuroscience literature from multi-source literature databases and standardize the literature formats to obtain standardized literature; Identifying entities in the standardization document based on the first model, and extracting connection relationships and directionality of entities in the standardization document based on the identified entities using the second model; The extracted entities are mapped to the standard ontology library in this field using fuzzy matching algorithm, and a unified entity ontology library is constructed by integrating multiple ontology libraries; The connection relationships between entities are stored in the form of triples in a graph database, and a directed brain connection knowledge graph is constructed. At the same time, the selected neuroscience literature and its structured information are stored in a document database and a relational database respectively.
2. The neuroscience literature processing method for constructing a brain connectivity knowledge graph according to claim 1, characterized in that: The first model adopts an entity recognition model based on a pointer annotation network, and the entity recognition model includes: Input layer: used to accept neuroscience corpus containing entity annotation information; Data preprocessing layer: used to standardize the corpus, using BIO annotation method and using dictionary tree to correct boundary inconsistencies; Pointer annotation network layer: used to reallocate BIO labels and start and end position pointers to the sequence after word segmentation, and train based on the hybrid cross entropy loss function; Output layer: Output entity recognition results, including text sequences and their corresponding label sequences.
3. The neuroscience literature processing method for constructing a brain connectivity knowledge graph according to claim 2, characterized in that: The second model adopts a relation extraction model that integrates the entity features identified by the entity recognition model and the position attention mechanism. The relation extraction model includes: Input layer: used to input corpus containing entity before and after position identifier information; Position Attention Network Layer: This layer fine-tunes the pre-trained parameters of the entity recognition model, selects the entity's preceding and following position identifiers as position attention features, and combines convolutional and pooling layers to learn entity and position features to generate high-quality feature vectors. Output layer: used to output relationship classification results, supporting binary or multi-classification tasks.
4. The neuroscience literature processing method for constructing a brain connectivity knowledge graph according to claim 2 or 3, characterized in that: The neuroscience corpus received by the input layer of the entity recognition model includes a gold standard corpus of neurons and neurotransmitters constructed using manual back-to-back annotation and a gold standard corpus of brain regions constructed from the existing project WhiteText; the corpus containing entity location information received by the input layer of the relationship extraction model is a brain region connection relationship corpus with directional information constructed through rules and manual back-to-back annotation.
5. The neuroscience literature processing method for constructing a brain connectivity knowledge graph according to claim 1, characterized in that: The multi-source literature database includes at least a PubMed database and a PMC database. The "collecting neuroscience literature from the multi-source literature database" means: screening out PubMed literature abstracts related to brain regions from the PubMed database based on keywords, and screening out PMC literature full texts related to brain regions from the PMC database, and obtaining relevant literature data.
6. The neuroscience literature processing method for constructing a brain connectivity knowledge graph according to claim 5, characterized in that: Terms defined in the Allen ontology database, BAMS ontology database, NeuroNames ontology database, and NeuroLex ontology database were used as keywords. PubMed literature abstracts related to brain regions were screened from the PubMed database, and PMC literature full texts related to brain regions were screened from the PMC database. The screened literature data at least included title, abstract, author, journal name, PMID, and publication time.
7. The neuroscience literature processing method for constructing a brain connectivity knowledge graph according to claim 6, characterized in that: The screened neuroscience literature is stored in a document-based database in CSV format, and the structured information of the neuroscience literature is stored in a relational database in JSON format.
8. The neuroscience literature processing method for constructing a brain connectivity knowledge graph according to claim 1, characterized in that: The triples used to represent the entity connection relationship are stored in the form of directed edges in a graph database, which is a Neo4j graph database.
9. The neuroscience literature processing method for constructing a brain connectivity knowledge graph according to claim 1, characterized in that: It also includes application analysis based on the brain connection knowledge graph, and the application analysis at least includes: intelligent literature retrieval, input and output loop analysis of specific brain regions, author cooperation relationship network analysis and research hotspot identification.
10. Brain connection knowledge graph system, characterized by: include: The data input layer is used to obtain information about brain connectivity from multi-source literature data. Terms defined in the Allen ontology database, BAMS ontology database, NeuroNames ontology database, and NeuroLex ontology database are used as keywords. PubMed literature abstracts related to brain regions are screened from the PubMed database, and PMC literature full texts related to brain regions are screened from the PMC database. Literature data including title, abstract, author, journal name, PMID, and publication date are obtained. The knowledge acquisition layer includes an entity recognition module and a relationship extraction module. The entity recognition module uses an entity recognition model based on a pointer annotation network to automatically identify entities in neuroscience literature. The relationship extraction module combines the entity output results of the entity recognition module with a positional attention mechanism to extract entity connection relationships in neuroscience literature, automatically determine the connection direction, and output structured relationship data. The knowledge graph evaluation layer includes an entity standardization module and a knowledge storage module. The entity standardization module uses a fuzzy matching algorithm to map the extracted entities to a standard ontology library in the field, and builds a unified entity ontology library by integrating multiple ontology libraries. The knowledge storage module stores the connection relationships of each entity in the form of triples in a graph database and constructs a directed brain connection knowledge graph. At the same time, the selected neuroscience literature and its structured information are stored in a document database and a relational database respectively. The knowledge graph retrieval application layer includes a literature retrieval module, a brain region circuit analysis module, and an author community analysis module. The literature retrieval module adopts the Vue front-end framework and the Flask back-end service architecture, and realizes the visualization of the knowledge graph based on the Relation-Graph; the brain region circuit analysis module is used to generate a whole-brain input and output connection path diagram for a specific brain region; the author community analysis module constructs a community map based on the relationship between literature authors and brain regions, and uses the Leuven community detection algorithm to perform community division and identify cooperative relationships and research hotspots between research teams.
Citation Information
Patent Citations
Methods, apparatus, computer equipment and readable storage media for extracting medical entity relationships
CN112256828B
Knowledge graph construction method for cerebral apoplexy
CN112541086A
Cited By
Multi-level intelligent mind map generation device based on large model
CN121616683A