Genome variation detection method and system based on few-sample learning
By introducing few-sample learning, adaptive memory module, supervised contrast learning module and orthogonal transfer module into the genomic variant detection model, combining external knowledge and optimizing deep learning technology, the problem of insufficient accuracy and robustness of existing models in few-sample learning is solved, and more efficient and accurate genomic variant detection is achieved.
Patent Information
- Application Number
- CN202510076029.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-17
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2045-01-17
AI Technical Summary
Existing genomic variant detection models are difficult to effectively combine external knowledge in small sample learning scenarios, resulting in insufficient accuracy and robustness of the detection.
Using a method that combines few-sample learning, adaptive memory module, supervised contrast learning module and orthogonal transfer module, we effectively combine external knowledge and improve model optimization capabilities through supervised contrast learning and orthogonal transfer optimization deep learning technology.
It improves the accuracy and robustness of genomic variant detection, and significantly improves the accuracy and generalization ability of detection under the condition of few samples.
Smart Images

Figure CN119993261A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of natural language processing and biomedical information extraction and analysis, and in particular to a genome variation detection method and system based on few-sample learning. Background Art
[0002] The detection and analysis of genomic variation is a core issue in modern biomedical research and clinical diagnosis. With the rapid development of genomics, precision medicine and big data technology, especially the rapid growth of genomic data, clinical data and academic literature, how to efficiently identify and process key information in these complex text data has become a major challenge. The timely discovery and accurate detection of genomic variation is of great significance for the early diagnosis of diseases, the formulation of personalized treatment plans and the evaluation of patient prognosis. It can help improve the treatment effect of diseases and reduce the risks in treatment.
[0003] Early genome variation detection relied on expert experience and manual analysis. Although this method can guarantee a certain degree of accuracy, it is time-consuming and labor-intensive, and is easily affected by expert subjective factors and cannot handle massive amounts of genomic data.
[0004] In recent years, the rise of deep learning technology has brought new solutions to genomic variation detection. Models such as deep neural networks (DNN), convolutional neural networks (CNN), and recurrent neural networks (RNN) can automatically learn and extract complex features from data, thereby capturing the potential semantic associations in the text and improving the accuracy and efficiency of detection. In particular, pre-trained language models based on self-attention mechanisms (such as BERT, BioBERT, etc.) have demonstrated excellent capabilities in processing biomedical texts. These models can capture the complex relationship between genetic variation and related clinical information through large-scale pre-training learning.
[0005] Although deep learning has made significant progress in genomic variation detection, it still faces some challenges. Most existing models focus on the analysis of internal features of text, ignoring the introduction and integration of external knowledge. The detection of genomic variation requires not only the identification of entities in the text, but also the understanding of the relationship between these entities and their biomedical background. Therefore, it is difficult to comprehensively improve the accuracy and robustness of detection by relying solely on text data for detection, especially in the scenario of few-sample learning.
[0006] Based on this, how to effectively combine external knowledge in genome variation detection and use deep learning technology to comprehensively improve the model's understanding and reasoning capabilities of complex semantic information has become a technical problem that needs to be urgently solved in this field. Summary of the invention
[0007] In view of this, the present invention provides a genome variation detection method and a related system that combine few-sample learning, an adaptive memory module, a supervised contrastive learning module and an orthogonal migration module, aiming to improve the accuracy and robustness of genome variation detection and achieve efficient and precise gene variation detection through the optimization of supervised contrastive learning and orthogonal migration and deep learning technology, effectively combining external knowledge and efficient model optimization.
[0008] In order to achieve the above object, the present invention adopts the following technical solution:
[0009] A genome variation detection method based on few-sample learning includes the following steps:
[0010] Constructing a genome variation detection model for few-sample learning, wherein the genome variation detection model includes a data preprocessing module, an adaptive memory module, a supervised contrastive learning module, an orthogonal migration module, and a decoding module;
[0011] The data preprocessing module is used to extract the location information and span information of the genomic variation entity in the genomic variation related text to be analyzed, and map each word in the text into a vector representation of a fixed dimension through a pre-trained language model to generate a word embedding matrix;
[0012] The adaptive memory module is used to calculate and update the memory state by weighted summing the input features, and output the adaptive memory update features;
[0013] The supervised contrastive learning module is used to generate contrastive learning optimization features under few-sample conditions by using contrastive learning strategies and based on positive sample pairs and negative sample pairs in adaptive memory update features;
[0014] The orthogonal migration module is used to receive contrastive learning optimization features, optimize feature space through orthogonal constraints, and output orthogonal constraint optimization features;
[0015] A decoding module is used to fuse the contrastive learning optimization features output by the supervised contrastive learning module and the orthogonal constraint optimization features output by the orthogonal migration module to generate the final genome variation detection results;
[0016] The constructed genome variation detection model is trained, and the trained genome variation detection model is used to complete genome variation detection.
[0017] Preferably, the specific contents of data preprocessing in the data preprocessing module include:
[0018] S11: Extract genomic variant entities in the text, mark their starting and ending positions in the text, and generate entity position span information [start i ,endi ], where start i and end i Respectively represent the starting and ending indexes of the i-th entity;
[0019] S12: Based on the marked entity positions, calculate and encode the position span information of each entity; at the same time, assign relevant category information C to each entity i , where C i represents the category of the i-th entity;
[0020] S13: converting the extracted entity position span information and category information into matrix representations respectively;
[0021] S14: Based on the entity position span information, the entities in the text are combined with their corresponding context information to generate an entity feature embedding matrix E.
[0022] Through the position span information, the preprocessing module in the genome variation detection model can extract the corresponding entity vocabulary and its adjacent context information from the original text, thereby generating the entity feature embedding matrix E.
[0023] Preferably, the adaptive memory module outputs the adaptive memory update feature, specifically comprising the following steps:
[0024] S21: Input the entity embedding matrix E into the adaptive memory module and use the attention mechanism to generate the context vector c t ;
[0025] S22: Use gating mechanism to control the dynamic update of new information and memory matrix;
[0026] S23: through the gate vector g t Update the memory matrix to get a new memory matrix M t+1 :M t+1 =g t ·c t +(1-g t )·M t ;
[0027] where · represents element-by-element multiplication, g t is the gate control vector, c t is the context vector generated by the attention mechanism, M t ∈R m×d is the memory matrix at time step t;
[0028] S24: Using the updated memory matrix M t+1 The input embedding matrix E is refined to obtain the refined embedding feature E'=Attention(E,M t+1 ,θ'attn ), and as the adaptive memory update feature output by the adaptive memory module, where θ' attn Indicates the updated memory M t+1 Parameters of the attention mechanism integrated with the input embedding E, where E represents the embedding matrix.
[0029] The adaptive memory module strengthens the model's ability to handle long-distance dependencies by dynamically storing and updating contextual information, helping the model to retain the memory of key information in a few-sample learning environment.
[0030] Preferably, in the supervised contrastive learning module, a contrastive learning strategy is used to generate contrastive learning optimization features under few-sample conditions based on positive sample pairs and negative sample pairs in the genomic variation data to be analyzed, specifically including:
[0031] S31: The refined embedding feature E i 'As input; each feature vector e' i Corresponding to a labeled sample, the label set Y = {y1,y2,…,y n} means, where y i is the category label of the i-th sample;
[0032] S32: For the labeled sample e' i , define the same category label y i =y j The sample pair (e' i ,e' j ) is a positive sample pair, and sample pairs of different categories are negative sample pairs;
[0033] Assume that the positive sample set is p(i) and the negative sample set is A(i), that is: P(i) = {j|y i =y j ,j≠i},A(i)={j|y i ≠y j}, where e' i Represents the feature vector of the i-th sample, which comes from the output of the adaptive memory module, y i is the category label of the i-th sample, p(i) is the positive sample set, including all positive samples with e' i The sample index of the same category, A(i) is the negative sample set, including all samples with the same category as e' i Sample indexes of the same category;
[0034] S33: Construct a supervised contrast loss function to maximize the similarity of positive sample pairs and minimize the similarity of negative sample pairs. The supervised contrast loss function is:
[0035]
[0036] Where I represents the sample set; sim() represents the cosine similarity function; τ is the temperature parameter, which is used to adjust the smoothness of the similarity distribution and is usually a positive value less than 1; exp is an exponential function, which is used to amplify the influence of similarity; |P(i)| usually represents the size of the positive sample set p(i), which is used to normalize the positive sample pairs;
[0037] S34: By minimizing the supervised contrast loss L SCL , optimize the feature representation and generate the feature embedding representation E' optimized by contrastive learning SCL , and as the contrastive learning optimization feature output by the supervised contrastive learning module, the feature embedding representation E' SCL The specific expression is:
[0038] E' SCL,i =Refine(e' i ,L SCL θ SCL );
[0039] Among them, θ SCL is the parameter of the supervised contrastive learning module, E' SCL It is the embedding representation optimized by contrastive learning, and Refine represents the function for optimizing features.
[0040] Preferably, the orthogonal migration module in the method of the present invention can realize efficient migration of cross-task knowledge by using the orthogonal constraint mechanism, specifically including:
[0041] S41: Embed the features optimized by contrastive learning into representation E' SCL As input, where E' SCL ={e' SCL,1 ,e' SCL,2 ,…,e' SCL,n};
[0042] S42: Embed the features optimized by contrastive learning into representation E' SCL Projected into the shared latent space, the orthogonal transformation matrix W is used for projection, and the transformed feature vector Z i The calculation formula is: i =We' SCL,i , where W is an orthogonal matrix with dimension d×d, d is the dimension of the eigenvector, satisfying WW T =I,Z i is the feature representation after projection, e' SCL,i represents the i-th feature vector after supervised contrastive learning optimization;
[0043] In the above steps, the orthogonality constraint is maintained by the following regularization term:
[0044]
[0045] Among them, ||·|| F represents the Frobenius norm; L orth is the orthogonality constraint loss, which is used to constrain W to maintain orthogonality; I is the unit matrix with dimension d×d;
[0046] In the above steps, the total loss function L is defined, and the task-specific loss L task With the orthogonality constraint loss L orth Combination: L = L task +λL orth , where λ is a hyperparameter that balances the two losses; L is the total loss function, L task is the task-specific loss function; L orth It is an orthogonal constraint loss function, which effectively transfers knowledge while maintaining feature independence by minimizing L;
[0047] S43: The feature Z after orthogonal projection i The orthogonal constrained optimization features are output from the final orthogonal migration module.
[0048] The orthogonal transfer module ensures knowledge sharing between different tasks while avoiding interference between tasks by imposing orthogonal constraints.
[0049] Preferably, the decoding module is used to fuse the contrastive learning optimization features output by the supervised contrastive learning module and the orthogonal constraint optimization features output by the orthogonal migration module, specifically including:
[0050] S51: Input the orthogonal constraint optimization feature Z obtained after processing by the orthogonal migration module into the joint learning framework; Z = {z1, z2, …, z n}, where z i is the feature vector of the i-th entity;
[0051] S52: Use the joint learning framework to optimize the contrastive learning feature E' output by the supervised contrastive learning module SCL , fused into Z to comprehensively utilize the information learned from different modules. The fused feature representation F is calculated as follows:
[0052] F=Z+E' SCL ;
[0053] Where F represents the final entity feature vector, Z is the feature representation after being processed by the orthogonal migration module, and E' SCL is the optimized feature obtained from the adaptive memory module and the supervised contrastive learning module
[0054] S53: Input the fused feature representation F into the linear classifier W cls In the example, the Softmax function is used to predict the entity category label. Among them, W cls is the weight matrix of the classifier, b cls is the bias vector, is the predicted entity category, and F is the fused feature representation;
[0055] S54: Define the cross entropy loss function L for the classification task ent , used to optimize the accuracy of entity recognition: Among them, y i is the true label, is the predicted probability, N is the number of samples;
[0056] S55: By minimizing the loss function L ent , optimize the parameters of the entire model. Finally, the predicted genomic variation entity category is output.
[0057] The decoding module in the present invention integrates the output features of the adaptive memory module, the supervised contrastive learning module and the orthogonal migration module, and adopts a joint learning strategy to simultaneously optimize the small sample entity recognition task. Through collaborative optimization, the genome variation detection model of few-sample learning can eventually generate genome variation detection results.
[0058] Preferably, the text related to the genomic variation to be analyzed is obtained by the following steps:
[0059] Collect text data related to the genomic variation to be detected from genomic databases (such as tmVar, BRONCO, EMU, etc.) and public scientific research literature.
[0060] Preferably, the above method also includes a comprehensive test and evaluation of the model after completing the training of the genomic variation detection model. The test process uses an independent test set, and the evaluation indicators include accuracy, recall, F1 value and other evaluation indicators for comprehensive analysis. By adjusting the key hyperparameters in the model (such as the number of memory slots of the adaptive memory module, the temperature coefficient of the supervised contrast learning module, the regularization coefficient of the orthogonal migration module, etc.), the model performance is optimized. Further combining the output features of the adaptive memory module, the supervised contrast learning module and the orthogonal migration module, the detection accuracy and generalization ability of the model under the condition of few samples are improved.
[0061] The optimized model can accurately identify variant entities in complex genomic variation texts, providing reliable technical support for clinical diagnosis and genomics research.
[0062] Preferably, the specific rules for improving the model using the AdamW optimization algorithm are:
[0063]
[0064] Among them, m t 、v t are the estimated values of the first-order moment and the second-order moment respectively; β1 and β2 are the attenuation factors of the first-order moment and the second-order moment, usually β1 = 0.9 and β2 = 0.999; ▽ θ L(θ t ) represents the gradient of the loss function L with respect to the model parameter θ; η is the learning rate, which is usually set to a small value, such as η = 1e-4; and ε is a small constant to prevent division by zero, usually set to ε = 1e-8; θ t+1 The optimized updated model parameters. This optimization algorithm adaptively adjusts the learning rate and combines the first-order and second-order moment estimation to efficiently update the parameters, ensuring stable training in few-sample learning scenarios and improving the model's convergence speed and prediction accuracy.
[0065] On the other hand, the present invention also discloses a genome variation detection system based on few-sample learning, comprising: a data preprocessing module, an adaptive memory module, a supervised contrast learning module, an orthogonal migration module and a decoding module;
[0066] The data preprocessing module is used to extract the location information and span information of the genomic variation entity in the genomic variation related text to be analyzed, and map each word in the text into a vector representation of a fixed dimension through a pre-trained language model to generate a word embedding matrix;
[0067] The adaptive memory module is used to calculate and update the memory state by weighted summing the input features, and output the adaptive memory update features;
[0068] The supervised contrastive learning module is used to generate contrastive learning optimization features under few-sample conditions by using contrastive learning strategies and based on positive sample pairs and negative sample pairs in adaptive memory update features;
[0069] The orthogonal migration module is used to receive contrastive learning optimization features, optimize feature space through orthogonal constraints, and output orthogonal constraint optimization features;
[0070] The decoding module is used to fuse the output features of the supervised contrastive learning module and the orthogonal migration module to generate the final genome variation detection results.
[0071] It can be seen from the above technical solutions that, compared with the prior art, the present invention provides a genome variation detection method and system based on few-sample learning. By introducing an adaptive memory module, a supervised contrastive learning module, and an orthogonal migration module, a deep understanding and efficient detection of genome variation texts are achieved. Compared with traditional technologies that rely only on rule methods or simple neural network architectures, the present invention exhibits stronger generalization ability and higher detection accuracy under few-sample conditions. The present invention adopts a multi-module collaboration strategy that fully combines the semantic embedding ability of pre-trained language models (such as BioBERT), the long-distance dependency modeling ability of adaptive memory modules, the feature differentiation ability of supervised contrastive learning modules, and the cross-task knowledge sharing advantages of orthogonal migration modules. By comprehensively utilizing the global semantic information of texts and the local features of genome variation entities, the present invention can effectively capture complex language patterns and biomedical background knowledge, and significantly improve accuracy and robustness in the identification tasks of multiple types of genome variations. In addition, the dynamic memory update mechanism and orthogonal constraint strategy can further optimize the adaptability and expansion ability of the model, so that the present invention has strong application potential in genome variation detection tasks in real scenarios. BRIEF DESCRIPTION OF THE DRAWINGS
[0072] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings required for use in the embodiments or the prior art description are briefly introduced below. Obviously, the drawings described below are only embodiments of the present invention, and for ordinary technicians in this field, other drawings can be obtained based on the provided drawings without creative work.
[0073] Figure 1 The accompanying drawing is a schematic diagram of the overall framework topology structure of a genome variation detection method based on few-sample learning provided by the present invention;
[0074] Figure 2 The accompanying drawing is a schematic diagram of the structure of the pre-trained language model BioBERT provided by the present invention;
[0075] Figure 3 The accompanying drawings are schematic diagrams of the adaptive memory module, supervised contrastive learning module and orthogonal migration module provided by the present invention, showing how to integrate output features and simultaneously optimize small sample genome variation identification tasks. DETAILED DESCRIPTION
[0076] The following will be combined with the attached embodiment of the present invention Figure 1-3, a clear and complete description of the technical solutions in the embodiments of the present invention is given. Obviously, the described embodiments are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of the present invention.
[0077] Example 1
[0078] The embodiment of the present invention discloses a genome variation detection method based on few-sample learning, which can effectively identify genome variation entities and their categories in biomedical texts, especially in a few-sample learning environment, improve the accuracy and generalization ability of the detection results, and provide strong technical support for genomic research, personalized medicine and clinical diagnosis. The method is implemented by a genome variation detection model that can be trained through few-sample learning, and the embodiment of the present invention includes: a data preprocessing module, an adaptive memory module, a supervised contrast learning module, an orthogonal migration module and a decoding module.
[0079] The genome variation detection method of the present invention is described in detail below:
[0080] S1: Preprocess the genomic variation data to be analyzed, extract and annotate the entity position and span information in the text, which specifically includes:
[0081] S11: Before processing genomic variation data for analysis, preprocessing steps are required, including word segmentation and structured annotation of relevant text content.
[0082] The first step is to identify and extract genomic variation entities in the text, and record the start and end positions of these entities in the text.
[0083] Based on the marked entity positions, the position span information of each entity is calculated and encoded. This information is recorded as span i =[start i ,end i ], where start i and end i Indicates the start index and end index of any i-th entity, and the position span information span i Details the scope of the entity within the text.
[0084] S12: Assign a specific category information C to each entity i , where C i Represents the category of the i-th entity, which may include gene type, mutation type, etc.
[0085] S13: Convert the extracted entity position span information and category information into a matrix representation. The dimension of the entity span matrix S is N×2, where N is the number of entities, and each row records the start and end positions of an entity. The dimension of the category matrix C is N×1, which is used to store the category label of each entity.
[0086] S14: For further processing, it is necessary to generate a feature embedding representation of the entity. This step combines the entity in the text with its associated context information to construct the entity's feature embedding matrix E. At the same time, the entity and its category information are preliminarily processed to provide the required input data for subsequent model training.
[0087] By preliminarily processing the entities and their category information through the above steps, it is possible to prepare input data for subsequent model training.
[0088] S2: Input the preprocessed text features into the adaptive memory module, which strengthens the model's ability to handle long-distance dependencies by dynamically storing and updating context information. Through the memory and forgetting mechanism, it helps the model retain key information in a few-sample learning environment;
[0089] S21: Input the entity position span information S and category information C extracted in S1 into the embedding layer, and use the pre-trained language model (such as BioBERT) to vectorize the text. The input text sequence is X = {x1.x2,…,x n}, where n is the length of the text, and each word x i , mapped to vector e through embedding function i :e i =Embed(x i θ BioBERT ). Among them, θ BioBERT is the pre-trained parameter of the BioBERT model, x i Represents the i-th word in the text sequence, and the generated vector representation matrix E has the dimension E=[e1,e2,…,e n ];
[0090] S22: Initialize the memory matrix M of the adaptive memory module (AMM) t ∈R m ×d , where m is the number of memory slots and E is the embedding matrix.
[0091] The embedding matrix E is input into the adaptive memory module and the context vector c is generated using the attention mechanism. t , which summarizes the information of the current input sequence. The calculation formula is c t =Attention(E,Mt ,θ attn ), where θ attn is the parameter of the attention mechanism;
[0092] S23: Use the gating mechanism to control the dynamic update of new information and memory matrix. The gating vector g t Calculated by the following formula: g t =σ(W g c t +b g ), where σ is the Sigmoid activation function, w g and b g is a learnable parameter, c t It is the attention mechanism that generates the context vector.
[0093] S24: through the gate vector g t Update the memory matrix to get a new memory matrix M t+1 :M t+1 =g t ·c t +(1-g t )·M t . Where · represents element-by-element multiplication, g t is the gate control vector, c t is the context vector generated by the attention mechanism, M t ∈R m×d is the memory matrix at time step t. This update mechanism ensures that the memory matrix dynamically integrates new context information while retaining key information, thereby improving the generalization ability of the model in few-sample scenarios.
[0094] S25: Using the updated memory matrix M t+1 Refine the input embedding E. By recontextualizing the input sequence using the updated memory, we get the refined embedding E' = Attention(E,M t+1 ,θ' attn ). Among them, θ' attn Indicates the updated memory M t+1 Parameters of the attention mechanism integrated with the input embedding E, where E represents the embedding matrix.
[0095] S3: Combined with the output of the adaptive memory module, the feature representation is optimized through supervised contrastive learning to enhance the semantic distinction between entities of different categories. By constructing positive and negative sample pairs, the similarity between entities of the same category is maximized and the similarity between entities of different categories is minimized;
[0096] S31: The feature representation E updated by the S2 adaptive memory module i'As input, E'={e'1,e'2,…,e' n} is the embedding representation optimized by the adaptive memory module, where n is the length of the text. Each feature vector e' i Corresponding to a labeled sample, the label set Y = {y1,y2,…,y n} means, where y i is the category label of the i-th sample.
[0097] S32: Construct positive and negative sample pairs. For sample e' i , define the same category label y i =y j The sample pair (e' i ,e' j ) is a positive sample pair, and a sample pair of different categories is a negative sample pair. Let the positive sample set be p(i) and the negative sample set be A(i), that is: P(i) = {j|y i =y j ,j≠i},A(i)={j|y i ≠y j}, where e' i The feature vector of the i-th sample comes from the output of the adaptive memory module, y i is the category label of the i-th sample, p(i) is the positive sample set, including all positive samples with e' i The sample index of the same category, A(i) is the negative sample set, including all samples with the same category as e' i Sample indexes of the same category;
[0098] S33: Define the supervised contrast loss function, which aims to maximize the similarity of positive sample pairs and minimize the similarity of negative sample pairs. The supervised contrast loss function is:
[0099]
[0100] Where I represents the sample set; τ is the temperature parameter, which is used to adjust the distribution of similarity and usually takes a positive value less than 1; exp is an exponential function, which is used to amplify the influence of similarity; |P(i)| usually represents the size of the positive sample set p(i), which is used to normalize the positive sample pairs;
[0101] sim() represents the cosine similarity function. In the embodiment of the present invention, the similarity between feature vectors is calculated, and the cosine similarity is used as a metric. For two samples e' i and e' j The similarity formula can be expressed as: Among them, sim(e' i ,e' j ) means e'i and e' j The cosine similarity between them is in the range of [-1,1], · is the dot product operation of the vector, ||e' i ||、||e' j || respectively represent e' i and e' j The L2 norm of is used to normalize the vector;
[0102] S34: By minimizing the supervised contrast loss L SCL , optimize the feature representation E', and generate the embedding representation E' optimized by contrastive learning SCL : E' SCL,i =Refine(e' i ,L SCL θ SCL ). Among them, θ SCL is the parameter of the supervised contrastive learning module, E' SCL is the embedding representation optimized by contrastive learning, Refine represents the process function of updating feature representation through back propagation and gradient descent, combined with L SCL For feature e' i Back propagation and gradient update are performed. This optimization process enhances the model's ability to distinguish samples of different categories, which helps to detect genomic variations more accurately in a small number of sample scenarios.
[0103] S4: Introducing the orthogonal transfer module, using the orthogonal constraint mechanism to achieve efficient cross-task knowledge transfer. This module ensures the knowledge sharing between different tasks by imposing orthogonal constraints;
[0104] S41: The feature representation E' optimized by supervised contrastive learning in S3 SCL As input, where E' SCL ={e' SCL,1 ,e' SCL,2 ,…,e' SCL,n}, these feature vectors contain entity representations after contrastive learning, aiming to improve the model's ability to distinguish different categories;
[0105] S42: Characteristic representation E' SCL Projected into a shared latent space to achieve effective transfer of cross-task knowledge. Use the orthogonal transformation matrix W for projection, and the transformed feature vector Z i The calculation formula is: i =We' SCL,i , where W is an orthogonal matrix with dimension d×d, d is the dimension of the eigenvector, satisfying WW T =I,Z i is the feature representation after projection, e' SCL,irepresents the i-th feature vector after supervised contrastive learning optimization;
[0106] In the above steps, orthogonal constraints are introduced to ensure the independence of feature space during knowledge transfer and prevent the knowledge of auxiliary tasks from interfering with the target task. The orthogonal constraints are maintained by the following regularization terms:
[0107]
[0108] Among them, ||·|| F represents the Frobenius norm; L orth is the orthogonality constraint loss, which is used to constrain W to maintain orthogonality; I is the unit matrix with dimension d×d;
[0109] In the above steps, the total loss function L is defined, and the task-specific loss L task ; and the orthogonality constraint loss L orth Combination: L = L task +λL orth , where λ is a hyperparameter that balances the two losses; L is the total loss function, L task is the task-specific loss function; L orth is an orthogonal constraint loss function. By minimizing L, knowledge transfer is effectively performed while maintaining feature independence;
[0110] S43: The feature Z after orthogonal projection i As the final representation, it is provided to the subsequent prediction layer for the classification task of genomic variation. Orthogonal projection ensures the stability and generalization ability of feature representation, which helps to improve model performance under few-sample conditions.
[0111] S5: Integrate the output features of the supervised contrastive learning module and the orthogonal transfer module, and use a joint learning strategy to simultaneously optimize the small sample entity recognition task;
[0112] S51: Input the feature representation Z processed by the orthogonal migration module into the joint learning framework. Let Z = {z1, z2, …, z n}, where z i is the feature vector of the i-th entity. These feature vectors are used for the final classification task of the entity category;
[0113] S52: Using the optimized features E' obtained from the supervised contrastive learning module SCL , further integrated into Z to comprehensively utilize the information learned from different modules. The fused feature representation F is calculated as follows: F = Z + E' SCL Where F represents the final entity feature vector, Z is the feature representation after being processed by the orthogonal migration module, and E' SCLIt is the optimized features obtained from the adaptive memory module and the supervised contrastive learning module;
[0114] S53: Input the fused feature representation F into the linear classifier W cls In the example, the Softmax function is used to predict the entity category label. Among them, W cls is the weight matrix of the classifier, b cls is the bias vector, is the predicted entity category, and F is the fused feature representation;
[0115] S54: Define the cross entropy loss function L for the classification task ent , used to optimize the accuracy of entity recognition: Among them, y i is the true label, is the predicted probability, N is the number of samples;
[0116] S55: By minimizing the loss function L ent , optimize the parameters of the entire model. Finally, the predicted genomic variation entity category is output.
[0117] The final output contains the category prediction results of all genomic variation entities. Through the comprehensive features output by the adaptive memory module, the supervised contrastive learning module, and the orthogonal migration module, the model can effectively identify genomic variation entities in the text and predict the category label of each entity and its classification probability. This joint learning strategy optimizes the expression of entity features under the condition of few samples, and improves the accuracy and generalization ability of entity recognition. With the help of the joint prediction layer, the model can accurately classify entities in genomic variation text based on the integration of multi-module feature information. The output results include the category label and corresponding confidence of each entity, providing efficient and reliable technical support for subsequent biomedical research, clinical diagnosis and genomic variation detection;
[0118] S6: Steps S1-S5 jointly construct a genome variation detection model for few-shot learning, which inputs the text to be detected and finally outputs the genome variation entity detection result;
[0119] S61: Collect text data related to the genomic variation to be detected from genomic databases (such as tmVar, BRONCO, EMU, etc.) and public scientific research literature. These data include information related to diseases, mutation types, gene mutation locations, etc. Standardize the collected text data, including removing noise, segmenting words, extracting the location information and category labels of genomic variation entities, and generating embedded vector representations to provide high-quality input data for model training;
[0120] S62: Input the collected genomic variation text into the genomic variation detection model trained in steps S1 to S5 for processing. The model uses a joint learning strategy to integrate the features of the adaptive memory module, the supervised contrastive learning module, and the orthogonal migration module to achieve the recognition and classification of genomic variation entities in the text. The model can accurately identify the variation type, variation location, and associated gene information in the text, and generate entity category labels and their classification probabilities;
[0121] S63: Output the prediction results of the genomic variation detection model, including the category label and corresponding confidence score of the genomic variation entity identified in each text. This result provides a reliable reference for biomedical research and clinical diagnosis, helping researchers and clinicians to quickly locate and analyze disease-related genomic variations.
[0122] After S6, the model is further tested and evaluated. The performance of the model is analyzed using commonly used evaluation indicators, including accuracy, recall, and F1 value, to comprehensively measure the performance of the model in entity recognition tasks. At the same time, by adjusting the model's hyperparameters (such as the relevant parameters of the pre-trained language model BioBERT, the adaptive memory module, the supervised contrastive learning module, and the orthogonal migration module), the model is optimized to improve the accuracy, robustness, and generalization ability of genomic variant entity detection.
[0123] In order to further implement the above technical solution, the model optimization is performed using the Adam optimizer, and its rules are as follows:
[0124]
[0125]
[0126] Among them, m t 、v t are the estimated values of the first-order moment and the second-order moment respectively; β1 and β2 are the attenuation factors of the first-order moment and the second-order moment, usually with values of β1 = 0.9 and β2 = 0.999; represents the gradient of the loss function L with respect to the model parameter θ; η is the learning rate, which is usually set to a small value, such as η = 1e-4; and ε is a small constant to prevent division by zero, usually set to ε = 1e-8; θ t+1 are the model parameters after optimization and update.
[0127] After each training stage, the model will update the weight parameters according to the gradient information of the loss function through the back-propagation mechanism, thereby gradually improving the performance of the model.
[0128] In the process of model performance evaluation, we use accuracy, recall and F1 value as core indicators to comprehensively measure the performance of the model in entity recognition tasks.
[0129] The accuracy is expressed as:
[0130]
[0131] The recall rate is expressed as:
[0132]
[0133] The F1 score corresponding to the F value is:
[0134]
[0135] Among them, TP represents the number of true positive examples, FP represents the number of false positive examples, and FN represents the number of false negative examples.
[0136] In this embodiment, the genomic variation text to be detected is processed by a pre-trained language model (such as BioBERT). First, the text is pre-processed by word segmentation and standardization, and the entity position information and span information in the text are extracted and converted into index form. Then, in the embedding layer, each word in the text is embedded as a vector representation to generate the word embedding matrix E of the text, where E = {e1, e2, …, e n}, n is the text length, and E is the embedded vector.
[0137] These embedding vectors are input into the adaptive memory module together with the location and span information of the entity. The adaptive memory module effectively captures the contextual information and long-distance dependencies in the text through a dynamic storage and update mechanism, and strengthens the modeling ability of key information.
[0138] The feature representation processed by the adaptive memory module is further input into the supervised contrastive learning module. This module optimizes the feature representation by constructing positive and negative sample pairs, enhances the ability to distinguish between entities of different categories, and improves the generalization performance of the model in a few-sample environment.
[0139] Through the orthogonal transfer module, the model transfers knowledge between multiple tasks and introduces orthogonal constraints to avoid information interference between different tasks, thereby further improving the robustness and detection accuracy of the model.
[0140] Finally, the output feature representations from the adaptive memory module, the supervised contrastive learning module, and the orthogonal migration module are integrated and input into the decoding layer. The decoding layer uses the joint feature representation to predict the category label of each entity through the Softmax function to generate the genomic variation detection results. The model is trained and tuned using optimization algorithms (such as AdamW) to further improve the detection accuracy and robustness of the model, and ultimately achieve efficient recognition and classification of genomic variation entities.
[0141] Example 2
[0142] Based on the above embodiments, in a specific embodiment, in order to verify the effectiveness of the model of the present invention, multiple groups of comparative experiments were conducted:
[0143] First, on public genomic variation detection datasets (such as tmVar and BRONCO&EMU), the performance of various models in the task of few-sample genomic variation detection was experimentally compared. These models include: a few-sample learning model (GDPN) combined with a Gaussian distribution prototype network, a named entity recognition method based on model-independent meta-learning (DMetaNER), a cross-domain few-sample model (MANNER) enhanced by variational memory, and a sequence annotation model (BioBERT-BiLSTM) combined with a pre-trained language model BioBERT and a bidirectional long short-term memory network (BiLSTM). The performance of the above baseline model was systematically compared and analyzed with the few-sample genomic variation detection model based on a hybrid neural network proposed in the present invention to verify the effectiveness and advantages of the model of the present invention.
[0144] The genomic variation detection model proposed in the present invention is based on the pre-trained language model BioBERT, and combines an adaptive memory module, a supervised contrastive learning module, and an orthogonal transfer module to enhance the generalization and accuracy of the model in the detection of few-sample genomic variations. Specifically, the adaptive memory module is used to dynamically store and update key contextual information, thereby improving the modeling ability of long-distance dependencies; the supervised contrastive learning module optimizes the feature representation of entities of different categories and enhances the classification and differentiation ability; the orthogonal transfer module realizes the effective transfer of cross-task knowledge through the orthogonal constraint mechanism, thereby further improving the performance and stability of the model.
[0145] To improve the performance and robustness of the model, the present invention adopts a variety of optimization strategies. The training batch size (train_batch_size) is set to 8 to balance the model training efficiency and hardware resource utilization; the evaluation batch size (eval_batch_size) is 1 to ensure high accuracy in the evaluation phase. The learning rate (learning_rate) is set to 3e-6, and a 10% warmup ratio (warmup_ratio) is used to achieve smooth training dynamics. The optimizer selects AdamW, combined with weight decay (weight_decay=1e-3) to reduce the risk of overfitting.
[0146] In terms of model design, the memory slot size (memory_size) of the adaptive memory module is set to 20 to ensure sufficient context storage capacity; the contrast loss temperature parameter (temperature) of the supervised contrastive learning module is set to 0.07 to ensure effective distinction between categories; the orthogonal constraint weight (orthogonal_weight) in the orthogonal transfer module is set to 0.1 through cross-validation to balance feature independence and sharing capabilities. The model hidden layer dimension (hidden_size) is 768, and the feedforward network dimension (feedforward_size) is set to 2048 to ensure the model's expressiveness and computational efficiency.
[0147] In the experiment, in order to prevent overfitting, the model uses dropout technology. Specifically, the dropout ratio of the attention layer (attention_dropout) and the dropout ratio of the BioBERT embedding layer (bert_dropout) are both set to 0.1, taking into account both learning ability and model robustness. In addition, in order to optimize the accuracy of detection, the relation filter threshold (relation_filter_threshold) is set to 0.4, which is obtained through multiple cross-validations to achieve the best balance between precision and recall.
[0148] All experiments were run on an Nvidia RTX 3090 GPU with 24GB of video memory, combined with the PyTorch framework to ensure efficient training and evaluation on large-scale datasets. The experimental results show that the proposed model significantly outperforms existing methods in multiple few-sample task scenarios, demonstrating its excellent ability in genomic variation detection tasks.
[0149] Several comparative models were constructed for testing, and the experimental results on the public genome tmVar dataset are shown in Table 1:
[0150] Table 1 tmVar comparison experimental results
[0151]
[0152] Several comparative models were constructed for testing, and the experimental results on the public genome BRONCO&EMU dataset are shown in Table 2:
[0153] Table 2 BRONCO & EMU comparative experimental results
[0154]
[0155]
[0156] Tables 1 and 2 show the experimental results of different models on public genomic variation detection datasets. It can be seen from the experimental data that the few-sample genomic variation detection model based on hybrid neural network proposed in the present invention has achieved significantly better performance than the baseline model in multiple task scenarios. Specifically, in the 6Shot-1Way scenario of the tmVar dataset, the model of the present invention achieved an F1 score of 75.99%, which significantly improved the detection ability under few-sample conditions compared with BioBERT-BiLSTM (64.01%) and MANNER (74.96%). In the more challenging 6Shot-3Way scenario, the model of the present invention still maintained excellent performance, with an F1 score of 47.62%, significantly surpassing DMetaNER (31.86%) and GDPN (26.09%). In the 5Shot-1Way and 5Shot-2Way tasks of the BRONCO&EMU dataset, the model of the present invention achieved F1 scores of 67.62% and 60.39%, respectively, which is also significantly better than other baseline models. Experimental results show that the proposed model has superior performance in the task of detecting genome variations with few samples, especially after combining the adaptive memory module and the supervised contrastive learning module, the model has shown excellent ability in processing long-distance dependencies and complex entity relationships. At the same time, the introduction of the orthogonal transfer module further enhances the cross-task generalization ability of the model, enabling it to effectively cope with the challenges of data scarcity and uneven label distribution.
[0157] The impact of different contrastive learning weights on model performance is shown in Table 3 in the experimental results of the tmVar dataset:
[0158] Table 3 Experimental results of tmVar with different contrastive learning weights
[0159]
[0160]
[0161] The impact of different contrastive learning weights on model performance is shown in Table 4 on the experimental results of the BRONCO&EMU dataset:
[0162] Table 4 Experimental results of BRONCO&EMU with different contrastive learning weights
[0163] Different contrast learning weights F1 score (%) F1 score (%) 5Shot-1Way 5Shot-2Way 0.10 64.59 52.39 0.15 67.62 60.39 0.20 59.61 53.03 0.25 56.65 53.79 0.30 57.48 51.86 0.40 55.26 52.62
[0164] In order to evaluate the impact of contrastive learning weights on model performance, experiments were conducted on public genomic variation detection datasets (such as tmVar and BRONCO&EMU), and the weight parameters in the contrastive learning module were adjusted to analyze its contribution to the few-shot genomic variation detection task. The contrastive learning weights control the degree of attention paid to the contrastive loss during the optimization process, thereby affecting the model's ability to distinguish features of entities of different categories.
[0165] The experimental results are shown in Tables 3 and 4: In the 6Shot-1Way scenario of the tmVar dataset, when the contrastive learning weight is set to 0.15, the F1 score of the model reaches 75.99%, which is better than the cases with weights of 0.1 (72.85%) and 0.2 (73.45%), showing the positive effect of moderate weights on model performance. Similarly, in the 5Shot-1Way scenario of the BRONCO&EMU dataset, when the contrastive learning weight is 0.15, the model achieves the best F1 score of 67.62%.
[0166] However, when the weights are set too high or too low, the model performance decreases. This suggests that too high weights may lead to over-reliance on contrastive learning loss, inhibiting the optimization of other loss terms, while too low weights cannot fully play the role of contrastive learning in feature differentiation.
[0167] In summary, the experimental results verify the significant impact of contrastive learning weights on model performance. By reasonably adjusting the weight parameters (such as 0.15), a balance can be achieved between feature discrimination ability and other optimization goals, thereby significantly improving the performance of the model in the task of detecting genome variations with a small number of samples.
[0168] The impact of the number of memory slots of different adaptive memory modules on model performance is shown in Table 5 in the experimental results of the tmVar dataset:
[0169] Table 5 Experimental results of tmVar dataset with different numbers of memory slots
[0170] Different number of memory slots F1 score (%) F1 score (%) F1 score (%) 6shot-1Way 6shot-2Way 6shot-3Way 5 73.41 47.50 20.00 10 68.50 43.45 40.81 15 74.74 51.95 36.92 20 75.99 59.87 47.61 25 62.38 40.37 44.44 30 56.01 54.48 40.00
[0171] The impact of the number of memory slots of different adaptive memory modules on model performance is shown in Table 6 in the experimental results of the BRONCO&EMU dataset:
[0172] Table 6 Experimental results of different numbers of memory slots on the BRONCO&EMU dataset
[0173]
[0174]
[0175] In order to study the impact of memory size on model performance, experiments were conducted on public genomic variation detection datasets (such as tmVar and BRONCO&EMU) to analyze the contribution of the number of memory slots (memory size) of the adaptive memory module to the task of few-sample genomic variation detection. The memory size determines the model's ability to dynamically store and update key information, thereby affecting the model's modeling effect on long-distance dependencies.
[0176] The experimental results are shown in Tables 5 and 6: In the 6Shot-1Way scenario of the tmVar dataset, when the memory size is set to 20, the model's F1 score reaches 75.99%, which is the best performance. A smaller memory size (such as 5) leads to insufficient context capture of the model, and the F1 score drops to 73.41%; while a larger memory size (such as 15) may introduce redundant information, increase computational complexity, and the performance drops to 74.74%. Similarly, in the 5Shot-1Way scenario of the BRONCO&EMU dataset, when the memory size is set to 20, the model achieves the highest F1 score of 60.39%, indicating that when dealing with more complex tasks, appropriately increasing the memory size can capture more effective information.
[0177] Experiments show that a moderate memory size is crucial to improving model performance. A smaller memory size limits the model's context memory capacity, while a larger memory size may introduce noise or computational overhead. By properly adjusting the memory size (such as 20 in the tmVar dataset and 20 in the BRONCO&EMU dataset), the adaptive memory module can more efficiently capture long-distance dependencies and significantly improve the accuracy of the few-sample genomic variation detection task.
[0178] The experimental results of the impact of different word representations on model performance on the tmVar dataset are shown in Table 7:
[0179] Table 7 Experimental results of different word representations on the tmVar dataset
[0180]
[0181]
[0182] The experimental results of the impact of different word representations on model performance on the BRONCO&EMU dataset are shown in Table 8:
[0183] Table 8 Experimental results of different word representations on the BRONCO&EMU dataset
[0184] Different word representation methods F1 score (%) F1 score (%) 5Shot-1Way 5Shot-2Way BERT 48.47 47.71 SciBERT 42.47 59.70 ALBERT 45.15 31.77 ClinicalBERT 51.26 51.26 BioBERT 67.62 60.39
[0185] In order to study the impact of different word representations on model performance, experiments were conducted on public genomic variation detection datasets (such as tmVar and BRONCO&EMU) to analyze the performance of different pre-trained language models in the task of few-sample genomic variation detection. These word representation models include: general pre-trained language models (such as BERT), pre-trained models for scientific text (SciBERT), parameter-efficient language models (ALBERT), models designed specifically for clinical text (ClinicalBERT), and models for biomedical text (BioBERT).
[0186] The experimental results are shown in Tables 7 and 8: In the 6Shot-1Way scenario of the tmVar dataset, BioBERT significantly outperforms other word representation models with an F1 score of 75.99%. This shows that BioBERT can more effectively capture contextual semantic information and fine-grained features in the biomedical field and adapt to the complexity of genomic variation detection tasks. In contrast, BERT's F1 score is 48.63%, reflecting its limitations in dealing with tasks in specific biomedical fields. SciBERT performs slightly better in the 6Shot-1Way scenario, achieving an F1 score of 51.42%, but its performance is still inferior to BioBERT due to its focus on broad semantics in the scientific field rather than specialized semantics in the biomedical field. In the 5Shot-1Way scenario of the BRONCO&EMU dataset, BioBERT also achieved the best performance with an F1 score of 67.62%. In contrast, although ClinicalBERT is optimized for clinical text, its generalization ability for diverse genomic variation expressions is insufficient, with an F1 score of 51.26%. ALBERT achieved an F1 score of 45.15% in the 5Shot-1Way scenario thanks to its efficient parameter design, and its performance has improved, but it still has limited performance when processing complex biomedical texts.
[0187] Experimental results show that the choice of word representation has a significant impact on model performance. BioBERT, with its large-scale pre-training on biomedical text, is able to capture fine-grained domain features and demonstrates excellent performance in the task of few-sample genomic variation detection. In contrast, general models and other domain-specific models have certain limitations when processing biomedical text. Therefore, choosing an appropriate word representation model (such as BioBERT) is the key to achieving high accuracy and robustness in genomic variation detection.
[0188] Example 3
[0189] Based on the above embodiments, in a specific embodiment, the present invention proposes a genome variation detection system based on few-sample learning, including: a data preprocessing module, an adaptive memory module, a supervised contrastive learning module, an orthogonal migration module and a decoding module;
[0190] Data preprocessing module, which is used to obtain and process input text related to genomic variation and perform basic text preprocessing operations. It extracts the location and span information of genomic variation entities in the text and uses the embedding layer to convert the text into a low-dimensional word vector representation of fixed dimension. Specifically, this module uses a pre-trained language model (such as BioBERT) to semantically embed each word in the text, map it into a dense vector containing semantic information, and finally generate E∈R n×d The word embedding matrix is a word embedding matrix of , where n is the length of the text and d is the dimension of the word vector. This embedding matrix is used as the input of subsequent modules to provide deep semantic representation to support subsequent analysis.
[0191] Adaptive memory module, which aims to improve the model's performance in dealing with long-distance dependencies by dynamically storing and updating key contextual information. Through the memory and forgetting mechanism, the adaptive memory module dynamically adjusts its memory state during training to retain important contextual information while forgetting irrelevant content. The input features interact with the memory unit through the attention mechanism to generate weighted contextual representations and update the memory state, thereby effectively modeling the long-distance dependencies of the input data. The output memory features provide enhanced contextual representations for subsequent modules, which helps the model understand complex textual relationships more accurately in a few-shot learning environment.
[0192] Supervised contrastive learning module, which optimizes input feature representation through contrastive learning method and improves the model's ability to distinguish different categories of genomic variant entities. Based on the construction of positive and negative sample pairs, the supervised contrastive learning module uses a similarity measurement mechanism to enhance the similarity between samples of the same category, while reducing the similarity between samples of different categories. By optimizing the contrastive loss function (such as cross entropy loss), this module significantly improves the robustness of feature representation and classification performance, especially under the condition of few samples, it can effectively enhance the ability to identify the category of genomic variant entities;
[0193] Orthogonal transfer module, which aims to achieve cross-task knowledge transfer through orthogonal constraint mechanism while avoiding negative interference between tasks. In the multi-task learning environment, this module introduces an orthogonal transformation matrix to ensure that the feature subspaces shared by different tasks remain orthogonal, thereby optimizing knowledge sharing and feature learning of specific tasks. By introducing orthogonal constraint regularization, this module promotes effective transfer and adaptability across tasks while maintaining the independence of feature spaces, further improving the generalization performance in the few-shot learning environment.
[0194] The decoding module is responsible for the final classification prediction of the features processed by each functional module to achieve the detection task of genomic variation. By fusing the output features of the supervised contrastive learning module and the orthogonal migration module, the decoding module integrates this information using the joint feature encoding layer and extracts key information through the maximum pooling operation. Subsequently, the classification layer predicts the fused comprehensive features and outputs the category label of the genomic variation. At the same time, by fine-tuning and optimizing the model parameters, this module further improves the accuracy and reliability of variation detection, thereby achieving efficient identification of genomic variation.
[0195] This system organically combines the functions of the above modules, from the preprocessing of input text to the final genome variation detection, to form a complete detection process. The data preprocessing module extracts key information, the adaptive memory module models context dependencies, the supervised contrastive learning module optimizes feature differentiation, the orthogonal transfer module realizes cross-task knowledge transfer, and finally the decoding module completes feature fusion and detection result output. The system has good robustness and adaptability, and can effectively complete genome variation detection tasks under the condition of few samples.
[0196] Each embodiment in this specification is described in a progressive manner, with the focus on describing the differences from other embodiments. Similar parts of each embodiment can be referenced to each other. For the device part disclosed in the embodiment, since it corresponds to the method part disclosed in the embodiment, the relevant content can refer to the description of the method part.
[0197] Through the detailed description of the above embodiments, those skilled in the art can implement and use the present invention. For the modification and variation of the present invention, the professional and technical personnel in the field can make adjustments according to the actual application without departing from the spirit and scope of the present invention. Therefore, the present invention is not limited to the specific embodiments listed here, but should be based on its wide applicability, in accordance with the basic principles and technical innovations of the present invention.
Claims
1. A genome variation detection method based on few-sample learning, characterized in that: The following steps are involved: Constructing a genome variation detection model for few-sample learning, wherein the genome variation detection model includes a data preprocessing module, an adaptive memory module, a supervised contrastive learning module, an orthogonal migration module, and a decoding module; The data preprocessing module is used to extract the location information and span information of the genomic variation entity in the genomic variation related text to be analyzed, and map each word in the text into a vector representation of a fixed dimension through a pre-trained language model to generate a word embedding matrix; The adaptive memory module is used to calculate and update the memory state by weighted summing the input features, and output the adaptive memory update features; The supervised contrastive learning module is used to generate contrastive learning optimization features under few-sample conditions by using contrastive learning strategies and based on positive sample pairs and negative sample pairs in adaptive memory update features; The orthogonal migration module is used to receive contrastive learning optimization features, optimize feature space through orthogonal constraints, and output orthogonal constraint optimization features; A decoding module is used to fuse the contrastive learning optimization features output by the supervised contrastive learning module and the orthogonal constraint optimization features output by the orthogonal migration module to generate the final genome variation detection results; The constructed genome variation detection model is trained, and the trained genome variation detection model is used to complete genome variation detection.
2. The genome variation detection method based on few-sample learning according to claim 1, characterized in that: The specific contents of data preprocessing in the data preprocessing module include: S11: Extract genomic variant entities in the text, mark their starting and ending positions in the text, and generate entity position span information [start i ,end i ], where start i and end i Respectively represent the starting and ending indexes of the i-th entity; S12: Based on the marked entity positions, calculate and encode the position span information of each entity; at the same time, assign relevant category information C to each entity i , where C i represents the category of the i-th entity; S13: converting the extracted entity position span information and category information into matrix representations respectively; S14: Based on the entity position span information, the entities in the text are combined with their corresponding context information to generate an entity feature embedding matrix E.
3. The genome variation detection method based on few-sample learning according to claim 2, characterized in that: The adaptive memory module outputs the adaptive memory update feature, specifically comprising the following steps: S21: Input the entity embedding matrix E into the adaptive memory module and use the attention mechanism to generate the context vector c t ; S22: Use gating mechanism to control the dynamic update of new information and memory matrix; S23: through the gate vector g t Update the memory matrix to get a new memory matrix M t+1 :M t+1 =g t ·c t +(1-g t )·M t ; where · represents element-by-element multiplication, g t is the gate control vector, c t is the context vector generated by the attention mechanism, M t ∈R m×d is the memory matrix at time step t; S24: Using the updated memory matrix M t+1 The input embedding matrix E is refined to obtain the refined embedding feature E'=Attention(E,M t+1 ,θ' attn ), and as the adaptive memory update feature output by the adaptive memory module, where θ' attn Indicates the updated memory M t+1 Parameters of the attention mechanism integrated with the input embedding E, where E represents the embedding matrix.
4. The genome variation detection method based on few-sample learning according to claim 3, characterized in that: In the supervised contrastive learning module, a contrastive learning strategy is used to generate contrastive learning optimization features under a few-sample condition based on positive sample pairs and negative sample pairs in the genomic variation data to be analyzed, specifically including: S31: The refined embedding feature E i 'As input; each feature vector e i 'Corresponding to a labeled sample, use the label set Y = {y1,y2,…,y n } means, where y i is the category label of the i-th sample; S32: For the labeled sample e i ', define the same category label y i =y j The sample pair (e' i ,e' j ) is a positive sample pair, and sample pairs of different categories are negative sample pairs; S33: Construct a supervised contrast loss function to maximize the similarity of positive sample pairs and minimize the similarity of negative sample pairs. The supervised contrast loss function is: Where I represents the sample set; sim() represents the cosine similarity function; τ is the temperature parameter, which is used to adjust the smoothness of the similarity distribution and is usually a positive value less than 1; exp is an exponential function, which is used to amplify the influence of similarity; |P(i)| usually represents the size of the positive sample set p(i), which is used to normalize the positive sample pairs; S34: By minimizing the supervised contrast loss L SCL , optimize the feature representation and generate the feature embedding representation E' optimized by contrastive learning SCL , and as the contrastive learning optimization feature output by the supervised contrastive learning module, the feature embedding representation E' SCL The specific expression is: AND' SCL,i =Refine(and' i ,THE SCL ;θ SCL ); Among them, θ SCL is the parameter of the supervised contrastive learning module, E' SCL It is the embedding representation optimized by contrastive learning, and Refine represents the function for optimizing features.
5. The genome variation detection method based on few-sample learning according to claim 4, characterized in that: The orthogonal transfer module can use the orthogonal constraint mechanism to achieve efficient transfer of cross-task knowledge, specifically including: S41: Embed the features optimized by contrastive learning into representation E' SCL As input, where E' SCL ={e' SCL,1 ,e' SCL,2 ,…,e' SCL,n }; S42: Embed the features optimized by contrastive learning into representation E' SCL Projected into the shared latent space, projected using the orthogonal transformation matrix W, the transformed feature vector Z is obtained i , the calculation formula is: Z i =We' SCL,i , where W is an orthogonal matrix with dimension d×d, d is the dimension of the eigenvector, satisfying WW T =I,Z i is the feature representation after projection, e' SCL,i represents the i-th feature vector after supervised contrastive learning optimization; S43: The feature Z after orthogonal projection i Orthogonal constrained optimization features as output of the final orthogonal migration module.
6. The genome variation detection method based on few-sample learning according to claim 1, characterized in that: The decoding module is used to fuse the contrastive learning optimization features output by the supervised contrastive learning module and the orthogonal constraint optimization features output by the orthogonal transfer module, specifically including: S51: Input the orthogonal constraint optimization feature Z obtained after processing by the orthogonal migration module into the joint learning framework; Z = {z1, z2, …, z n }, where z i is the feature vector of the i-th entity; S52: Use the joint learning framework to optimize the contrastive learning feature E' output by the supervised contrastive learning module SCL , fused into Z to comprehensively utilize the information learned from different modules. The fused feature representation F is calculated as follows: F=Z+E' SCL ; Where F represents the final entity feature vector, Z is the feature representation after being processed by the orthogonal migration module, and E' SCL It is the optimized features obtained from the adaptive memory module and the supervised contrastive learning module; S53: Input the fused feature representation F into the linear classifier W cls In the example, the Softmax function is used to predict the entity category label. Among them, W cls is the weight matrix of the classifier, b cls is the bias vector, is the predicted entity category, and F is the fused feature representation; S54: Define the cross entropy loss function L for the classification task ent , used to optimize the accuracy of entity recognition: Among them, y i is the true label, is the predicted probability, N is the number of samples; S55: By minimizing the loss function L ent , optimize the parameters of the entire model, and finally output the predicted genomic variation entity category.
7. The genome variation detection method based on few-sample learning according to claim 1, characterized in that: The text related to the genomic variation to be analyzed is obtained by the following steps: Extract text related to the genomic variants to be detected from biomedical databases and literature resources.
8. The genome variation detection method based on few-sample learning according to claim 1, characterized in that: It also includes optimization evaluation of the constructed genomic variation detection model.
9. The genome variation detection method based on few-sample learning according to claim 8, characterized in that: The AdamW optimization algorithm is used to optimize the performance of the model. The specific rules of the AdamW optimization algorithm are as follows: m t =β1m t-1 +(1-β1)▽ θ L(θ t ); v t =β2v t-1 +(1-β2)(▽ θ L(θ t )) 2 ; Among them, m t 、v t are the estimated values of the first-order moment and the second-order moment respectively; β1 and β2 are the attenuation factors of the first-order moment and the second-order moment, usually β1 = 0.9 and β2 = 0.999; ▽ θ L(θ t ) represents the gradient of the loss function L with respect to the model parameter θ; η is the learning rate, which is usually set to a small value, such as η = 1e-4; ε is a small constant to prevent division by zero, and its value is ε = 1e-8; θ t+1 are the model parameters after optimization and update.
10. A genome variation detection system based on few-sample learning, characterized in that: A genome variation detection method based on few-sample learning according to any one of claims 1 to 9 is applied, comprising: a data preprocessing module, an adaptive memory module, a supervised contrastive learning module, an orthogonal migration module and a decoding module; The data preprocessing module is used to extract the location information and span information of the genomic variation entity in the genomic variation related text to be analyzed, and map each word in the text into a vector representation of a fixed dimension through a pre-trained language model to generate a word embedding matrix; The adaptive memory module is used to calculate and update the memory state by weighted summing the input features, and output the adaptive memory update features; The supervised contrastive learning module is used to generate contrastive learning optimization features under few-sample conditions by using contrastive learning strategies and based on positive sample pairs and negative sample pairs in adaptive memory update features; The orthogonal migration module is used to receive contrastive learning optimization features, optimize feature space through orthogonal constraints, and output orthogonal constraint optimization features; The decoding module is used to fuse the output features of the supervised contrastive learning module and the orthogonal migration module to generate the final genome variation detection results.
Citation Information
Patent Citations
Transfer learning method and device based on domain pair association
CN115983375A
Multi-modal pre-training model migration method based on self-supervised learning
CN118097685A
Continual text recognition using prompt-guided knowledge distillation
US20240362937A1
Entity relation mining method based on biomedical literature
WO2021190236A1
Cited By
Genome variation filtering method and system based on deep learning
CN120748511A
A deep learning-based method and system for filtering genomic variants
CN120748511B
Virus genome data analysis and prediction system based on deep learning
CN120913637A