Deep learning-based mRNA field literature mining system

By designing a document mining system in the field of mRNA based on deep learning, the problem of difficulty in extracting core entities and their semantic relationships in the field of mRNA in the field of mRNA is solved in the prior art, efficient entity recognition and relationship extraction are achieved, and the efficiency and accuracy of literature analysis are significantly improved.

CN119990143AInactive Publication Date: 2025-05-13RICE RES INST GUANGDONG ACADEMY OF AGRI SCI

Patent Information

Application Number
CN202510464941.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-15
Publication Date
2025-05-13
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

The prior art is difficult to quickly and accurately extract core entities and their complex semantic relationships in the mRNA field from massive biomedical literature, especially when facing information overload, complex semantic understanding and multitasking fusion needs.

Method used

A document mining system in the field of mRNA based on deep learning is designed, using a customized optimization deep learning model, combined with a relation classification module, to achieve efficient collaborative optimization of named entity recognition (NER) and relationship extraction (RE). The system includes a data preprocessing module, a model training and inference module, and a result verification and storage module.

Benefits of technology

It significantly improves the analytical ability of unique languages ​​in mRNA literature, improves the performance of entity recognition and relationship extraction tasks, and can efficiently mine core entities and their complex semantic relationships in mRNA research, providing strong support for subsequent data analysis and knowledge graph construction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119990143A_ABST
    Figure CN119990143A_ABST
Patent Text Reader

Abstract

The invention provides an mRNA field literature mining system based on deep learning, relates to the technical field of bioinformatics and natural language processing, and can effectively identify mRNA-related entities including genes, proteins, diseases and the like and semantic relationships thereof. According to the system, the latest BioBERT model and the Transform-CRF model are combined, efficient collaborative optimization of named entity recognition and relation extraction is achieved through a multi-task learning architecture, and the precision and efficiency of literature mining in the mRNA field are remarkably improved; specific terms and complex syntactic structures in the field of biomedicine can be effectively processed, redundancy is reduced through multi-task optimization, and the overall performance is improved; by fusing the pre-training knowledge of BioBERT and the sequence labeling capability of Transform-CRF, an innovative technical support is provided for deep mining and knowledge discovery of mRNA related literatures, and the method has a wide application prospect.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of bioinformatics and natural language processing, and in particular to a document mining system in the mRNA field based on deep learning. Background Art

[0002] With the rapid development of high-throughput sequencing and molecular biology, the study of mRNA has become an important means to reveal gene regulatory networks and disease mechanisms. However, the massive amount of biomedical literature makes it difficult for researchers to obtain and analyze the latest research results in a timely manner.

[0003] The important role of mRNA in gene expression regulation, protein synthesis and disease treatment is becoming increasingly prominent. For example, the successful application of messenger RNA vaccines in virus prevention and control has demonstrated its great potential in the medical field. However, the number of research papers related to this has increased dramatically. In the PubMed database alone, the number of mRNA-related papers has exceeded hundreds of thousands. This phenomenon has led researchers to face the following challenges: 1. Information overload: It is difficult for researchers to obtain high-quality information related to their own research in a timely manner, and traditional manual reading and screening methods are inefficient.

[0004] 2. Difficulty in understanding complex semantics: mRNA-related literature often involves complex terminology, context-dependent relational expressions, and multidisciplinary background knowledge, which are difficult to handle with traditional rule-based or dictionary-based mining methods.

[0005] 3. Technical requirements for multi-task fusion: Researchers usually need to not only extract named entities from literature, but also understand the semantic relationships between entities (such as the regulatory effects of genes and proteins), which puts higher demands on traditional single-task models.

[0006] The existing technology also has the following limitations: Traditional information extraction methods rely on manually designed rules, such as regular expressions or template-based matching algorithms. Although such methods show a certain degree of accuracy in specific tasks, they lack flexibility and generalization ability. When the syntactic structure is complex or the terms are diverse, rule-based methods tend to fail.

[0007] Models such as Hidden Markov Model (HMM) and Support Vector Machine (SVM), although they overcome certain rule limitations, are inefficient when faced with large-scale corpora and complex relationships due to their reliance on manual feature engineering.

[0008] Deep learning-based language models (such as Word2Vec, ELMo, and BERT) have made significant progress in natural language processing tasks. However, these models are usually trained on general corpora (such as Wikipedia) and fail to be optimized for corpora in the biomedical field, resulting in unsatisfactory performance in domain tasks.

[0009] Therefore, building an intelligent system that can quickly and accurately mine entities and their relationships in mRNA-related literature is of great significance for promoting basic and applied research. Summary of the invention

[0010] The present invention combines the latest deep learning technology to design a literature mining system, method and medium adapted to the needs of mRNA research for the first time. The system adopts a custom optimized deep learning model, combined with a relationship classification module, to achieve efficient collaborative optimization of named entity recognition (NER) and relationship extraction (RE). Compared with traditional literature mining methods, the present invention can not only extract core entities in mRNA research (such as genes, proteins, diseases, etc.), but also efficiently mine the complex semantic relationships between them (such as regulation, activation, inhibition, etc.), providing support for subsequent data analysis and knowledge graph construction.

[0011] To achieve the above object, the technical solution adopted by the present invention is: on one hand, a deep learning-based mRNA field literature mining system is provided, comprising: The data preprocessing module is used to clean the biomedical literature, perform word and sentence segmentation, generate candidate entities, and extract semantic features to provide structured input for model training and reasoning. Model training and reasoning module, which is used to complete named entity recognition and relationship extraction tasks by combining deep semantic modeling methods with sequence optimization technology; The result verification and storage module is used to call the verification algorithm to check the consistency of the mining results and store the output in a structured format.

[0012] Preferably, the specific step of data cleaning includes using regular expressions to remove noise information in the document, including irrelevant characters, HTML tags and specific format symbols; the specific step of text segmentation and sentence segmentation includes using a sentence segmentation method based on punctuation and context rules to split the document into independent sentences; using a word segmentation algorithm to decompose complex terms into semantic sub-units and retain context features; the specific step of candidate entity generation includes extracting candidate entities based on rule-based pattern matching and domain-specific dictionaries; the specific step of semantic feature extraction includes using a deep semantic coding method to perform feature modeling on the context of candidate entities.

[0013] Preferably, the model training and reasoning module includes: The semantic modeling module is used to deeply model the semantic dependencies of texts through a bidirectional context modeling method; The sequence optimization module is used to improve the boundary accuracy and label consistency of entity recognition by using a global optimization method based on conditional random fields (CRF); The relation classification module is used to classify the semantic relations between entity pairs using a relation extraction method based on a multi-head attention mechanism, using a custom annotated data table and performing secondary relation extraction through a one-to-one mapping method; The task collaborative optimization module is used to adopt a multi-task learning framework to collaboratively optimize the named entity recognition (NER) and relation extraction (RE) tasks, and use shared features to improve the overall performance; The result fusion module is used to combine the prediction results based on deep learning and improve the accuracy of the mining results through cross-validation or weighted methods.

[0014] More preferably, the semantic modeling module adopts an improved Transformer architecture, captures contextual semantic dependencies through a multi-layer self-attention mechanism, and is optimized for research needs in the mRNA field.

[0015] More preferably, the sequence optimization module combines the advantages of the Transformer and CRF models, captures the global semantic features in the text through the self-attention mechanism, and uses the CRF layer to model the transfer relationship of the label sequence.

[0016] Preferably, in the data preprocessing module, a rule-based and statistical method is used to generate candidate entities, thereby ensuring a high recall rate of the candidate entity set.

[0017] Preferably, in the model training and reasoning module, an improved multi-task learning framework is used in combination with shared features to achieve collaborative optimization of named entity recognition and relationship extraction.

[0018] Compared with the prior art, the beneficial effects of the present invention include the following: 1. Deep learning model adapted for mRNA research: This paper designs a pre-trained model optimized specifically for the mRNA field, taking into account the unique semantics, terminology, and text structure of mRNA. The model is innovative in structure, adopts a hierarchical encoding layer, and is particularly optimized for biomedical entity relationships such as genes and proteins. These innovative designs can effectively capture complex terms and long-distance dependencies, improve the ability to parse the unique language in mRNA literature, and thus improve the performance of entity recognition and relationship extraction tasks.

[0019] 2. Multi-task learning framework: The multi-task learning framework is used to achieve task sharing and collaborative optimization of entity recognition (NER) and relation extraction (RE), solving the problem of task separation in traditional methods.

[0020] 3. Optimize the accuracy and reliability of the model: This paper introduces a sequence annotation method based on a custom optimization mechanism, combines global context information and local semantic features, and significantly improves the performance of the model in processing mRNA literature through joint optimization. At the same time, the system also verifies the results by calling the API of the large language model to ensure the reliability and accuracy of the mining results.

[0021] By introducing these innovative technologies, the present invention not only solves the lack of adaptability of existing methods in the mRNA field, but also improves the application effect of the model in complex biomedical tasks such as gene regulatory networks. BRIEF DESCRIPTION OF THE DRAWINGS

[0022] Figure 1 This is a framework diagram of a deep learning-based mRNA field literature mining system of the present invention.

[0023] Figure 2 This is an overall flow chart of a method for literature mining in the mRNA field based on deep learning of the present invention.

[0024] Figure 3 It is the task sharing layer model architecture diagram of the present invention.

[0025] Figure 4 It is a diagram of the task-specific layer model architecture of the present invention. DETAILED DESCRIPTION

[0026] See also Figure 1 As shown, the present invention provides, in a first aspect, a system for document mining in the field of mRNA based on deep learning, comprising: The data preprocessing module is used to clean the biomedical literature, perform word and sentence segmentation, generate candidate entities, and extract semantic features to provide structured input for model training and reasoning. In this invention, data preprocessing is the first step to achieve efficient literature mining, which mainly includes the following technical processes: 1. Data cleaning: Since biomedical literature comes from a wide range of sources, it may contain irrelevant characters, HTML tags, or specific format symbols. Therefore, regular expressions (Regex) are first used to clean the original text to remove interfering information and retain the core content.

[0027] 2. Text Sentence and Word Segmentation: Using an adaptive sentence segmentation algorithm based on punctuation marks such as the period ("."), question mark ("?"), exclamation mark ("!") and a vocabulary of biological terms, the sentence segmentation method can be optimized according to the context, domain terminology and structure. Specifically, when encountering complex molecular biology terms, the algorithm checks the vocabulary and analyzes the sentence structure to adjust the sentence boundaries, thereby improving the accuracy of sentence segmentation processing. This sentence segmentation strategy improves analysis efficiency while avoiding cross-sentence analysis redundancy.

[0028] o Word segmentation: Combine the vocabulary and language model to perform subword level word segmentation on the text. In the present invention, the word segmentation stage relies on other tools or algorithms, such as the WordPiece word segmentation algorithm, which helps to process the splitting of specific terms.

[0029] 3. Entity candidate generation: Use pattern matching methods (such as regular expressions or dictionary matching) to identify possible entities (such as genes, proteins). The purpose of this stage is to screen candidate entities and provide candidate inputs for subsequent models.

[0030] Model training and reasoning module, which is used to complete named entity recognition and relationship extraction tasks by combining deep semantic modeling methods with sequence optimization technology; The core model structure of this system has been optimized and innovated many times, such as Figure 3 and Figure 4 As shown in the figure, it mainly includes the following modules: encoder module, context modeling module, attention fusion module and sequence labeling module.

[0031] The encoder module uses a multi-layer custom Transformer encoder, which fully considers the modeling requirements of long texts in its design and optimizes and improves the traditional BERT in structure.

[0032] The encoder module consists of an embedding layer + a multi-layer custom Transformer encoder + a fully connected layer.

[0033] The input format of the encoder module: The input text is processed by word segmentation and lemmatization and converted into a sequence of word vectors. The word vector sequence contains three types of embeddings: word embedding, position embedding, and segment embedding.

[0034] The output format of the encoder module is as follows: The output is a sequence of feature vectors encoded by a multi-layer Transformer. These feature vectors provide a deep semantic representation of each word in the text.

[0035] Specific operation process of the encoder module: Input processing (embedding layer): After the input text is segmented and lemmatized, it is first converted into a sequence of word vectors. In this process, the model uses three embeddings (word embedding, position embedding, and paragraph embedding) to encode the input to ensure that each word can be fully represented in terms of semantics, position, and paragraph structure. The input processing process combines the results of text sentence segmentation and word segmentation to generate a sequence of word vectors for subsequent Transformer encoder processing. The input processing part involves the preprocessing steps of the text, including text sentence segmentation and word segmentation. This part refers to the text processed by the sentence segmentation algorithm and the word segmentation algorithm, which, after further word vectorization steps, serves as the input of the multi-layer custom Transformer encoder.

[0036] Multi-layer Transformer encoder: The input word vector sequence is processed by 12 layers of deep Transformer encoder in turn. In each layer, a self-attention mechanism and a feedforward neural network are included to capture the complex dependencies in the text. In particular, the present invention fuses the output of each layer, avoiding the defect that the traditional model only relies on the output of the last layer.

[0037] Inter-layer feature fusion: The outputs of different layers are fused through concatenation operations to ensure that the semantic information of different layers is effectively combined, thereby generating a comprehensive word vector sequence containing global information and details. The features of each layer are fused through a carefully designed weighting strategy to enhance the ability to capture long-distance dependencies.

[0038] Fully connected dimensionality reduction: In order to meet the needs of downstream tasks, the concatenated word vector sequence is input into the first fully connected layer for dimensionality reduction mapping, converting the high-dimensional vector into a low-dimensional representation suitable for downstream tasks.

[0039] The context modeling module includes a BiLSTM network and a context modeling layer; The input format of the context modeling module is: the input is the feature vector sequence from the encoder module (the first word vector sequence) and the word vector sequence of its input (the second word vector sequence).

[0040] The output format of the context modeling module: The output is a word vector sequence containing context information (the third word vector sequence).

[0041] The specific process of the context modeling module is as follows: Bidirectional Long Short-Term Memory (BiLSTM): Based on the output of the encoder module, the BiLSTM layer is used to further capture the contextual information of the sequence. This module captures long-range dependencies in the text through forward and reverse LSTM units, and ensures that the model can handle more complex contextual information, improving the model's ability to handle subtle differences in biomedical text.

[0042] The attention fusion module includes improved attention mechanism (Improve-Attention), multi-head attention mechanism, context vector fusion and fully connected layer.

[0043] The input format of the attention fusion module is: the input is the third word vector sequence from the context modeling module.

[0044] The output format of the attention fusion module is: the output is a semantically enhanced entity label sequence (the fourth word vector sequence).

[0045] The specific operation process of the attention fusion module: Improved attention mechanism: The improved attention mechanism of the present invention dynamically calculates the semantic association between each pair of word vectors by introducing a feedforward neural network, thereby overcoming the limitation that the traditional attention mechanism only relies on fixed weights. Specifically, the traditional attention mechanism usually determines the association by calculating the dot product between word units, while the improved method of the present invention first uses a feedforward neural network to process each word vector (the third word vector sequence) to generate a more accurate semantic relevance score. This score reflects the semantic relationship between words in the context, and these scores are normalized by the Softmax function to obtain the weight of each word. Then, these weights are used to perform weighted fusion on the word vectors to generate a context vector. Compared with the traditional attention mechanism, the improved mechanism can capture long-distance dependencies more flexibly and can dynamically adjust weights according to changes in context, thereby improving the model's ability to understand complex text contexts. This enhanced attention mechanism can effectively improve the accuracy of entity and relationship recognition in long texts, especially when processing complex sentences containing multiple entities.

[0046] The scoring formula is as follows: , in: and They are query word vector (Query) and key word vector (Key), which come from the input word vector sequence. is the weight matrix of the feedforward neural network, is the bias term. is the activation function , which aims to improve the nonlinear modeling capabilities of the model. It means to sum the word vectors of query and key, which increases the semantic combination and thus more fully represents the association between them.

[0047] Weight calculation: First, use the scoring function to calculate the semantic relevance between the current word and other words, and normalize these relevances through the Softmax function to obtain the weight.

[0048] ▪Context vector generation: The weighted word vector is concatenated with the original word vector to generate a context vector. This vector is further processed by a nonlinear activation function (such as tanh) and finally reduced in dimension by a second fully connected layer to generate a semantically enhanced word vector sequence.

[0049] Multi-head attention mechanism: Through the multi-head self-attention mechanism, the model can focus on multiple semantic patterns at the same time and capture diverse contextual information. The multi-head attention mechanism calculates different attention weights based on the context vector, and then in the context vector fusion module, the outputs of multiple heads are fused together. The fully connected layer then performs dimensionality reduction and nonlinear processing on these fused vectors. The formula is as follows: , in, , , are query matrix, key matrix and value matrix respectively, is the dimension of the key vector, which indicates the processing of the context vector. The multi-head mechanism can capture different semantic relationships in parallel and improve the generalization ability of the model.

[0050] The sequence labeling module includes a conditional random field (CRF) layer and a label transfer optimization layer; The input format of the sequence labeling module is the semantically enhanced entity label sequence from the attention fusion module.

[0051] The output format of the sequence labeling module: the output is the final entity label sequence.

[0052] Conditional Random Field (CRF) layer: In sequence labeling tasks, the CRF layer is used to optimize the entity label sequence. The CRF layer can improve the accuracy of entity boundaries and label consistency by modeling the transfer relationship between labels. From a global perspective, the model can optimize the entity label sequence and output the final label. The loss function formula of the CRF layer is as follows: , in, is a given input and tag sequence The score function when , the score function can be composed of the emission score and the transfer score: , in, is the emission fraction, indicating the label (such as "gene A" or "protein X") and input features (e.g., “the role of gene A in regulating the cancer process”); is the transfer score, indicating the label (such as "Gene A") and The legitimacy between the two (such as "protein B") reflects whether there is a biological regulatory or interactive relationship between genes and proteins. In the mRNA literature mining task, the label represents biological entities in the literature, such as genes, proteins, or diseases, and the input features is the corresponding text fragment (such as a word or phrase in a sentence). Therefore, the emission score It is a measure of the degree of match between the occurrence of a biological entity in the text and the entity label. For example, if the label yiy_iyi is "gene A" and the text segment If it contains "Gene A plays a role in regulating the cancer process", then the emission score will be high, indicating that the label "Gene A" has a good match with the semantics in the text. In the mRNA literature, labels usually refer to different types of biological entities, and the relationships between these entities are diverse. For example, there may be a regulatory relationship between "Gene A" and "Protein B", and there may be an inductive relationship between "Gene A" and "Disease C". The legal transfer relationship between these entities is affected by biological laws. For example, the regulatory relationship between a gene and a protein is biologically conventional, but the inhibitory relationship between a gene and a protein may be illegal.

[0053] The specific operation process of the sequence labeling module is as follows: Label transfer optimization: In the conditional random field (CRF), label transfer optimization improves the accuracy of label sequences by modeling the transition probability between labels. Specifically, CRF learns the transition relationship between labels (i.e., the legitimacy and possibility of changing from one label to another) and adjusts the position of each label in combination with the input feature information to ensure the coherence and consistency of the label sequence. This optimization process not only considers the emission probability of the label (the degree of match between the word and the label), but also comprehensively considers the legal transfer rules between labels, so as to accurately determine the boundaries and labels of each entity, and finally generate an accurate and biologically consistent entity label sequence.

[0054] Global optimization: In the CRF layer, the entire label prediction process is optimized by modeling the global dependencies between labels. Specifically, CRF not only considers the matching degree of each label with the current input features, but also considers the legal transfer relationship between labels, thereby optimizing the label sequence globally. This global optimization ensures the overall consistency and coherence of the label sequence, avoids the possible mislabeling and unreasonable label transfer caused by relying only on local features, and ultimately generates a more accurate entity label sequence.

[0055] In the model optimization and task hierarchical design, the present invention adopts a multi-task learning framework to jointly optimize entity recognition and relationship extraction tasks in the task sharing layer and the task specific layer. The task specific layer is used to optimize the NER (named entity recognition) task and the RE (relation extraction) task respectively through the text features extracted by the sharing layer.

[0056] Task sharing layer: The text features extracted by the encoder will be passed to the NER and RE modules in the downstream tasks for shared use by the two tasks.

[0057] Task specific layers: NER task: In the NER task, the system first uses the encoder to generate the word vector representation of the text, and then optimizes the label sequence of the entity through the conditional random field (CRF) layer. Specifically, the CRF layer not only considers the label probability of each word unit, but also introduces the transfer relationship between labels, that is, considers the dependency between adjacent word unit labels. By modeling the conditional probability of label transfer, the CRF layer can effectively avoid erroneous jumps in label prediction (for example, illegal order of entity labels), thereby improving the accuracy of entity boundaries and label consistency. Finally, the CRF layer outputs the final optimized entity label sequence by maximizing the conditional probability of the entire label sequence. This global optimization mechanism provides significant performance improvements in sequence labeling tasks, especially for entity recognition of complex and long texts.

[0058] RE task: In the RE task, the system predicts the relationship between the identified entity pairs through the classification layer. Specifically, the goal of the RE task is to predict the semantic relationship between the entity pairs based on the contextual information between them. The classification layer receives the entity pairs (i.e., the target entities in the text) and their context vectors output by the NER task, and uses this information to classify the potential relationship between the entity pairs, such as "regulation", "activation", or "inhibition". During the classification process, the system first encodes the contextual information of the entity pairs through a feedforward neural network, and calculates the probability distribution of each relationship through the Softmax function, and finally outputs the most likely relationship label. In this way, the RE task can accurately capture and predict the complex semantic relationships between entity pairs, further enhancing the model's ability to understand and extract interactions between entities in biomedical texts.

[0059] 2.3 Loss Function In the process of mRNA literature mining, entity recognition (such as identifying "gene" or "protein") and relationship extraction (such as identifying "regulatory relationship between gene and protein") are complementary tasks. For example, the model needs to identify the two entities "gene A" and "protein B", and then identify whether there is a "regulatory" or "activation" relationship between them through the relationship extraction task. The loss function will combine these tasks to improve the accuracy of the overall model by optimizing the prediction of entities and relationships.

[0060] Loss function: The joint objective function combining NER loss and RE loss weighs the training effects of the two tasks. The formula is as follows: , is the loss of the entity recognition task, is the loss of the relation extraction task, is the weight adjustment coefficient between the two. The NER task uses the conditional random field (CRF) layer to optimize the sequence labeling results. Its loss function is in the form of negative log-likelihood (NLL), and the formula is: , in, Represents a given input sequence When predicting the label sequence The conditional probability is calculated as: , here, is the score function of the label sequence, which includes the weighted sum of the transfer score and the emission score. The specific formula is as follows: , Emission fraction ( ): indicates the Location tags With input features The degree of match is determined by the output probability of the model. It measures the likelihood of each token being assigned to a certain label.

[0061] Transfer score ( ): indicates adjacent labels and For example, some combinations of labels (such as "B-gene" followed by "E-gene") may conform to semantic rules better than other combinations.

[0062] In the mRNA literature, the relation extraction task aims to identify the mutual relationships between different entities. The RE task uses the cross entropy loss function to optimize the classification accuracy of the relationship between entity pairs. The formula is: , Among them, the real label at this time represents the true relationship between entity pairs (e.g., the “activation” relationship between gene A and protein B), The true label (0 or 1) of the sample, the predicted value is the relationship label predicted by the model (e.g., “regulate” or “inhibit”), The predicted probability of a sample. is the total number of samples. By calculating the cross entropy loss function, the model optimizer adjusts the weights to minimize the error between the predicted relationship and the true label, thereby achieving accurate relationship extraction.

[0063] Through the joint optimization of the above loss functions, the system further improves the classification performance of the relationship extraction task while maintaining the accuracy of entity recognition.

[0064] Pre-training and fine-tuning: The model is initialized using pre-training data in the biomedical field and fine-tuned on annotated mRNA literature data to improve the model's adaptability to specific fields. The goal of fine-tuning is to minimize the loss function of the target task, which can be expressed as: , in, is a task-specific loss function (such as cross entropy loss), It's a label. is the label predicted by the model and N is the number of samples.

[0065] Reasoning phase: oInput processing: After the documents to be analyzed are processed by sentence and word segmentation, they are sent to the model in batches for processing.

[0066] oEntity Recognition: The NER module uses the features output by the encoder to perform entity recognition and optimizes sequence labeling through the CRF layer.

[0067] o Relation Extraction: The RE module extracts the relations between entity pairs through the classification layer.

[0068] Result verification and storage 1. Result verification: Call the big model API to automatically verify the extracted entities and relationships to ensure the accuracy and consistency of the output results.

[0069] 2. Result storage: oUse JSON format to save entity and relationship data for subsequent research.

[0070] oProvide database-based storage options to achieve efficient management of large-scale literature mining results.

[0071] like Figure 2 As shown, in the second aspect of the present invention, the present invention also provides a method for document mining in the field of mRNA based on deep learning, comprising the following steps: Data preprocessing: Clean biomedical literature and remove noise information; generate independent semantic units through sentence and word segmentation methods, and extract candidate entities and their context features; Model training and reasoning: Use deep learning methods based on bidirectional context modeling for semantic encoding; Use conditional random fields (CRFs) to optimize boundary predictions for named entity recognition tasks; combine multi-head attention mechanisms to classify and predict the semantic relationships of entity pairs; The result verification and storage module is used to call the verification algorithm to check the consistency of the mining results and store the output in a structured format.

[0072] Result verification and storage: Call the verification algorithm to automatically verify the mining results to ensure semantic consistency and logical accuracy; store the mining results in JSON format or database to provide support for knowledge graph construction; Output fusion: Perform weighted fusion or cross-validation on the prediction results of multiple models to improve the reliability of mining results.

[0073] The following is a preferred embodiment of the present invention to explain the method steps of the present invention: 1. Document input: Input the mRNA document to be processed, and the system automatically extracts the text and performs sentence processing. Each sentence will be broken down into a series of tokens and processed at the sub-word level through a custom word segmentation algorithm to ensure accurate identification of biological terms and key points in long texts.

[0074] 2. Entity Recognition: The system first processes each token (i.e., word segment or subword) in the text through a deep learning encoder. The encoder is based on a multi-layer custom Transformer architecture that can effectively capture long-distance dependencies and contextual information. Each token is processed layer by layer in the encoder to generate its semantic representation. Subsequently, the task-specific decoder is responsible for label prediction of these semantic representations and identifying the precise entities in the text, such as genes, proteins, etc.

[0075] 3. Relationship extraction: After completing entity recognition, the system performs contextual analysis on each pair of recognized entities to predict the semantic relationship between them. The relationship extraction module combines context vectors and global information to accurately identify interactive relationships between entities, such as "activation" and "inhibition". This process relies on a deep learning model structure that is enhanced by a specific optimization mechanism to effectively identify complex biomedical relationships.

[0076] 4. Result fusion: The system fuses the results of the entity recognition and relationship extraction stages. By combining the outputs of multi-layer encoders and task-specific decoders, more accurate entity recognition and relationship extraction results can be obtained. The fusion process uses weighted average or intersection methods to improve the accuracy and reliability of predictions. In addition, the system calls the large language model to verify the prediction results to further improve accuracy. The verification step of the large language model can feedback the verification results and iteratively optimize by adjusting model parameters or input data to improve the accuracy of subsequent processing.

[0077] 5. Result output: The verified entity and relationship results are output and saved as structured data, which is convenient for users to conduct subsequent research and data analysis. The output data format can flexibly support multiple storage methods, such as database storage and JSON format storage, to meet the needs of large-scale literature mining.

[0078] According to a third aspect of the present invention, a computer-readable storage medium is provided, wherein the computer program implements the above-mentioned system or method when executed by a processor.

[0079] The innovation of this system is that it is the first to design a deep learning model architecture adapted to mRNA research. This architecture not only significantly improves the accuracy of entity recognition, but also shows higher computational efficiency in relationship extraction tasks, especially when dealing with complex long texts and complex biomedical relationships, surpassing the performance of traditional methods. These innovations provide more efficient and accurate technical support for text analysis and data mining in the biomedical field, and have broad application potential.

[0080] The above implementation modes are merely descriptions of the preferred implementation modes of the present invention, and are not intended to limit the scope of the present invention. Without departing from the design spirit of the present invention, various modifications and improvements made to the technical solutions of the present invention by ordinary engineering and technical personnel in the field shall fall within the protection scope determined by the claims of the present invention.

Claims

1. A deep learning-based mRNA field literature mining system, characterized in that: Includes the following modules: The data preprocessing module is used to clean the biomedical literature, perform word and sentence segmentation, generate candidate entities, and extract semantic features to provide structured input for model training and reasoning. Model training and reasoning module, which is used to complete named entity recognition and relationship extraction tasks by combining deep semantic modeling methods with sequence optimization technology; The result verification and storage module is used to call the verification algorithm to check the consistency of the mining results and store the output in a structured format.

2. According to claim 1, a deep learning-based mRNA field literature mining system is characterized in that: The specific steps of data cleaning include using regular expressions to remove noise information in the document, including irrelevant characters, HTML tags and specific format symbols; the specific steps of text segmentation and sentence segmentation include using a sentence segmentation method based on punctuation and context rules to split the document into independent sentences; using a word segmentation algorithm to decompose complex terms into semantic sub-units and retain context features; the specific steps of candidate entity generation include extracting candidate entities based on rule-based pattern matching and domain-specific dictionaries; the specific steps of semantic feature extraction include using a deep semantic coding method to perform feature modeling on the context of candidate entities.

3. The mRNA field literature mining system based on deep learning according to claim 1, characterized in that: The model training and reasoning module includes: The semantic modeling module is used to deeply model the semantic dependencies of texts through a bidirectional context modeling method; The sequence optimization module is used to improve the boundary accuracy and label consistency of entity recognition by using a global optimization method based on conditional random fields (CRF); The relation classification module is used to classify the semantic relations between entity pairs using a relation extraction method based on a multi-head attention mechanism, using a custom annotated data table and performing secondary relation extraction through a one-to-one mapping method; The task collaborative optimization module is used to adopt a multi-task learning framework to collaboratively optimize the named entity recognition (NER) and relation extraction (RE) tasks, and use shared features to improve the overall performance; The result fusion module is used to combine the prediction results based on deep learning and improve the accuracy of the mining results through cross-validation or weighted methods.

4. The mRNA field literature mining system based on deep learning according to claim 3, characterized in that: The semantic modeling module adopts an improved Transformer architecture, captures contextual semantic dependencies through a multi-layer self-attention mechanism, and is optimized for research needs in the mRNA field.

5. The mRNA field literature mining system based on deep learning according to claim 3, characterized in that: The sequence optimization module combines the advantages of the Transformer and CRF models, captures the global semantic features in the text through the self-attention mechanism, and uses the CRF layer to model the transfer relationship of the label sequence.

6. The mRNA field literature mining system based on deep learning according to claim 1, characterized in that: In the data preprocessing module, a rule-based and statistical method is used to generate candidate entities, ensuring a high recall rate of the candidate entity set.

7. The mRNA field literature mining system based on deep learning according to claim 1, characterized in that: In the model training and reasoning module, an improved multi-task learning framework is used to combine shared features to achieve collaborative optimization of named entity recognition and relationship extraction.

Citation Information

Patent Citations

  • Entity relationship mining method based on biomedical literature

    CN111428036A

  • RoBERTa-BiLSTM-CRF voice dialogue text named entity recognition system fused with attention mechanism

    CN117010387A

  • Data mining analysis method based on entity recognition and relation extraction

    CN118550987A

Cited By

  • Professional technology maturity evaluation method based on artificial intelligence

    CN120234581A

  • An artificial intelligence-based professional technology maturity assessment method

    CN120234581B