Gene function prediction method and system based on multi-label active learning

Through the method based on multi-label active learning, the prototype network is used to optimize gene function prediction, and problems such as scarcity of data, insufficient label correlation and functional space sparseness in the initial stage are solved, efficient and accurate gene function prediction is achieved, and experimental costs are reduced.

CN120048348APending Publication Date: 2025-05-27SHANDONG UNIV
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510119226.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-24
Publication Date
2025-05-27

AI Technical Summary

Technical Problem

The scarcity of functional annotation data in the initial stage of existing gene function prediction methods leads to inaccurate prediction uncertainty, insufficient label correlation modeling, low prediction efficiency, insufficient gene-functional pair selection diversity and biological representation, redundant experiments are verified in the model, and it is difficult to efficiently process large-scale gene functional data.

Method used

Using a multi-label active learning method, a gene-function pair training set is constructed in a multi-label scenario, a labeled and unlabeled data set is generated, and a prototype network is used for training and optimization, and an initial information score is scored based on the uncertainty of gene function prediction, label correlation and label distribution imbalance, and a gene-function pair that constitutes the initial and alternative sets is selected to optimize the prototype network to improve prediction performance.

Benefits of technology

It significantly improves the efficiency and accuracy of gene function prediction, reduces the redundancy of experimental verification, reduces biological experiment costs, improves the model's prediction performance for rare functions, and adapts to the efficient processing needs of large-scale gene function data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120048348A_ABST
    Figure CN120048348A_ABST
Patent Text Reader

Abstract

The invention provides a gene function prediction method and system based on multi-label active learning, and the method comprises the steps: selecting gene-function pairs forming an initial set from an unlabeled data set according to an initial information score quantified by the uncertainty of gene function prediction, label correlation and label distribution imbalance; according to a selection standard of feature diversity, label space diversity and sample-label pair interaction diversity, a standby set is selected in combination with the initial information score; the gene function prediction efficiency and accuracy are remarkably improved, reliable technical support is provided for large-scale gene function research, and the method is particularly suitable for complex function prediction scenes such as research fields of new gene function discovery, disease-related gene function annotation and the like and has important practical application value.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field related to artificial intelligence, and particularly relates to a gene function prediction method and system based on multi-label active learning. Background Art

[0002] The statements in this part only provide background technical information related to the present invention and do not necessarily constitute prior art.

[0003] Gene function prediction aims to predict the relationship between genes and biological functions by analyzing the characteristic information of genes (such as sequence characteristics, expression levels, etc.).

[0004] As understood by the inventors, the problems existing in current gene function prediction at least include:

[0005] I. The problem of inaccurate uncertainty estimation of function prediction caused by scarce functional annotation data in the initial stage: Traditional methods usually require a large amount of annotated data, and the annotation of gene functions often depends on biological experiments or expert knowledge. Gene function annotation data is scarce and costly; in the initial stage of gene function prediction, due to the extremely small amount of functional annotation data, the uncertainty estimation of gene functions based on model prediction in existing methods is often inaccurate. This is because the model lacks sufficient training data in the initial stage, resulting in a large deviation in function prediction results, which in turn affects the accuracy of gene-function pair selection in active learning. Traditional uncertainty estimation methods are highly dependent on model prediction and are prone to selecting suboptimal gene-function pairs in the early stage, resulting in a decline in model prediction efficiency.

[0006] II. The problem of insufficient modeling of label correlation: In gene function data, there are often complex biological correlations between different functions, such as pathway co-occurrence relationships or functional exclusion relationships. These correlations have an important impact on function prediction efficiency and model performance. However, existing methods usually model label correlation through simple statistical means and are difficult to capture deeper biological relationships. Especially in the case of scarce functional annotation data, this deficiency not only limits the model's understanding of label correlation but may also lead to redundant experimental verification and waste of experimental resources.

[0007] III. The problem of low prediction efficiency caused by the sparsity of the function space: In gene function prediction, the function distribution is usually uneven. Most genes are only related to a small number of functions, while the number of genes for some functions, that is, rare functions, is extremely small. When existing methods deal with function sparsity, they tend to bias towards high-frequency functions and ignore the learning of rare functions, which will result in poor prediction performance of the model for rare functions. In addition, traditional methods lack a comprehensive exploration of the function space, which may lead to insufficient function coverage and thus limit the overall prediction performance.

[0008] IV. Problems of insufficient diversity and biological representativeness in gene-function pair selection: In the active learning of gene function prediction, existing methods often struggle to simultaneously consider the diversity of gene expression profiles and biological representativeness when batch-selecting gene-function pairs. For example, some methods only focus on genes with relatively high uncertainty in function prediction, but these genes may be too close in the expression profile space, resulting in redundant biological information; other methods do not fully consider the representativeness of gene expression distributions, which may lead to a decline in the generalization ability of the model for specific cell type or tissue-specific expression patterns.

[0009] V. Problem of redundant experimental verification in the mode: When existing batch function prediction methods select multiple gene-function pairs, it is often difficult to avoid redundant experimental verification, resulting in increased experimental costs. For example, some methods do not fully consider the biological complementarity between gene-function pairs during batch selection, leading to overlapping information content of the selected verification objects in the expression profile or function space, thus wasting limited experimental resources.

[0010] VI. Ability to efficiently process large-scale gene function data: In practical applications, gene function prediction data usually has high-dimensional expression profile features and large-scale gene samples, which poses high requirements for the computational efficiency of active learning methods. However, existing methods often have high computational complexity when dealing with large-scale genomic data and are difficult to meet the needs of real-time prediction.

[0011] Therefore, how to solve the key problems such as inaccurate prediction, low prediction efficiency, and poor model generalization ability in gene function prediction is an issue that needs to be addressed currently. Summary of the Invention

[0012] To overcome the deficiencies of the above-mentioned prior art, the present invention provides a gene function prediction method and system based on multi-label active learning, which improves the efficiency and accuracy of gene function prediction and provides reliable technical support for large-scale gene function research.

[0013] To achieve the above object, the present invention adopts the following technical solutions:

[0014] In the first aspect, the present invention provides a gene function prediction method based on multi-label active learning, including:

[0015] Construct a training set of gene-function pairs in a multi-label scenario, and generate a labeled data set and an unlabeled data set according to the training set;

[0016] Train a prototype network using the training set to obtain a trained prototype network;

[0017] Use the trained prototype network to perform prediction processing on the gene data to be predicted to obtain the prediction result of the gene data to be predicted;

[0018] Among them, in the iterative training of the prototype network, gene-function pairs that form an initial set are selected from the unlabeled dataset according to the initial information scores quantified by the uncertainty of gene function prediction, label correlation, and label distribution imbalance.

[0019] According to the selection criteria of feature diversity, label space diversity, and sample-label pair interaction diversity, combined with the initial information scores corresponding to the gene-function pairs, gene-function pairs that form an alternative set are selected from the initial set, and the prototype network is optimized according to the selected alternative set.

[0020] In a second aspect, the present invention provides a gene function prediction system based on multi-label active learning, including:

[0021] A construction module configured to: construct a training set of gene-function pairs in a multi-label scenario, and generate a labeled dataset and an unlabeled dataset according to the training set;

[0022] A training module configured to: train a prototype network using the training set to obtain a trained prototype network; among them, in the iterative training of the prototype network, gene-function pairs that form an initial set are selected from the unlabeled dataset according to the initial information scores quantified by the uncertainty of gene function prediction, label correlation, and label distribution imbalance.

[0023] According to the selection criteria of feature diversity, label space diversity, and sample-label pair interaction diversity, combined with the initial information scores corresponding to the gene-function pairs, gene-function pairs that form an alternative set are selected from the initial set, and the prototype network is optimized according to the selected alternative set;

[0024] A prediction module configured to: perform prediction processing on the gene data to be predicted using the trained prototype network to obtain a prediction result of the gene data to be predicted.

[0025] In a third aspect, the present invention provides an electronic device, including a memory and a processor, as well as computer instructions stored on the memory and running on the processor. When the computer instructions are run by the processor, the method described in the first aspect is completed.

[0026] In a fourth aspect, the present invention provides a computer-readable storage medium for storing computer instructions. When the computer instructions are executed by a processor, the method described in the first aspect is completed.

[0027] In a fifth aspect, the present invention provides a computer program product, including a computer program. When the computer program is executed by a processor, the method described in the first aspect is implemented.

[0028] The above one or more technical solutions have the following beneficial effects:

[0029] In the present invention, gene-function pairs that make up the initial set are selected from the unlabeled dataset according to the initial information scores quantified by the uncertainty of gene function prediction, label correlation, and label distribution imbalance; furthermore, according to the selection criteria of feature diversity, label space diversity, and sample-label pair interaction diversity, and combined with the initial information scores, a backup set is selected; significantly improving the efficiency and accuracy of gene function prediction, providing reliable technical support for large-scale gene function research, and being particularly suitable for complex function prediction scenarios, such as new gene function discovery, disease-related gene function annotation and other research fields, and having important practical application value.

[0030] The advantages of the additional aspects of the present invention will be partially given in the following description, partially become obvious from the following description, or be understood through the practice of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS

[0031] The specification drawings forming a part of the present invention are used to provide a further understanding of the present invention. The schematic embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute an improper limitation to the present invention.

[0032] Figure 1 It is the overall structure diagram of the model in the multi-label active learning method in the first embodiment of the present invention;

[0033] Figure 2 It is the structural diagram of the sample-label pair selection in the first stage in the first embodiment of the present invention;

[0034] Figure 3 It is the structural diagram of the sample-label pair selection in the second stage in the first embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0035] It should be noted that the following detailed descriptions are all exemplary and are intended to provide further descriptions of the present invention. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which the present invention belongs.

[0036] It should be noted that the terms used herein are only for describing specific embodiments and are not intended to limit the exemplary embodiments according to the present invention.

[0037] Without conflict, the embodiments in the present invention and the features in the embodiments can be combined with each other.

[0038] Embodiment 1

[0039] This embodiment discloses a gene function prediction method based on multi-label active learning, including:

[0040] Construct a training set of gene-function pairs in a multi-label scenario, and divide the training set into a labeled data set and an unlabeled data set;

[0041] Use the training set to train the prototype network to obtain a trained prototype network;

[0042] Use the trained prototype network to perform prediction processing on the gene data to be predicted to obtain the prediction result of the gene data to be predicted;

[0043] Among them, in the iterative training of the prototype network, according to the initial information scores quantified by the uncertainty of gene function prediction, label correlation, and label distribution imbalance, select gene-function pairs that form the initial set from the unlabeled data set;

[0044] According to the selection criteria of feature diversity, label space diversity, and sample-label pair interaction diversity, combined with the initial information scores corresponding to the gene-function pairs, select gene-function pairs that form the alternative set from the initial set, and optimize the prototype network according to the selected alternative set.

[0045] In this embodiment, "sample" refers to gene expression data, "label" refers to a functional label, and the labels of genes can be the following specific biological function categories:

[0046] Labels related to metabolic functions: such as "glycolysis", "lipid metabolism", "amino acid metabolism", etc.

[0047] Labels related to immune functions: such as "antigen processing and presentation", "inflammatory response", "B cell activation", etc.

[0048] Labels related to signal transduction: such as "Wnt signaling pathway", "MAPK signaling pathway", "calcium ion signaling pathway", etc.

[0049] Labels related to cell functions: such as "cell cycle regulation", "apoptosis", "DNA repair", etc.

[0050] By adding the above functional labels to gene samples, it can help biological researchers better understand the roles played by genes in different biological processes.

[0051] The method of this embodiment models the deep features of gene expression data and functional labels through a prototype network, significantly improving the accuracy of gene function label prediction and the robustness of label correlation modeling. The method of this embodiment is particularly applicable to challenges such as label sparsity and incomplete functional annotation commonly existing in the field of gene function prediction. It adopts a two-stage selection strategy, comprehensively considering various factors such as uncertainty and diversity, to ensure the accurate selection of high-information and representative gene-function pairs in the early stage of active learning.

[0052] The innovation of the method of this embodiment lies in its ability to efficiently mine the key information of gene expression profiles at the initial stage when gene function annotation data is scarce, avoiding the redundant annotation and waste of experimental resources caused by inaccurate prediction of traditional methods. By accurately screening the gene-function pairs to be annotated, the active learning algorithm significantly reduces the annotation workload of biological experts, maximally saves the experimental verification cost, and at the same time rapidly improves the model performance, providing an efficient solution for large-scale gene function prediction.

[0053] The following details each part involved in a gene function prediction method based on multi-label active learning proposed in this embodiment:

[0054] Prototype network construction: To efficiently select gene-function pairs with information gain for annotation, this embodiment proposes a prototype network based on functional label prototype vectors. This network is particularly adapted to the characteristics of high-dimensional sparse expression data in gene function prediction tasks and can accurately model the relationship between gene expression patterns and functional labels in the embedding space. Different from traditional methods that only rely on the distance metric of feature distributions, this embodiment explicitly models the positive and negative distributions of gene functions through the design of functional label prototype vectors, and at the same time combines the biological similarity between functional labels, significantly improving the modeling ability of complex gene function relationships.

[0055] Specifically, the prototype network maps gene expression data from a d-dimensional feature space to a d'-dimensional embedding space, and the mapping parameters are learned by the network. Through an optimization strategy, the prototype network can embed gene samples with the same function near the positive class prototype of the function label and far from its negative class prototype, thereby effectively distinguishing gene functions in the embedding space.

[0056] To further model the complex relationships between functional labels, this embodiment proposes a design method based on dynamic functional label prototype vectors. By dynamically updating the label prototype vectors, it captures the time-varying characteristics of gene function distributions and implicitly models the biological associations between functional labels. Most traditional methods design prototype networks for multi-classification problems, where each sample can only belong to one category, ignoring the multi-label nature of gene functions and the complex associations between functional labels. In contrast, the method of this embodiment is oriented towards multi-label scenarios and uses the hetGNN network as a dynamic prototype network for multi-label tasks, capable of handling multiple functional labels of a sample simultaneously.

[0057] The prototype network plays an important role in the following aspects:

[0058] First, the prototype network maps high-dimensional gene expression data into a low-dimensional embedding space, clarifying the distance relationship between gene samples and functional labels. In the low-dimensional space, the distribution of gene samples can more clearly reflect their potential functional labels;

[0059] Second, the prototype network generates positive and negative class prototype vectors for each functional label, representing the distribution centers of genes with and without this function respectively. At the same time, it allows a sample to be associated with multiple functional labels simultaneously, avoiding the limitations of the single-category assumption in multi-classification methods and significantly enhancing the ability of gene function prediction. Through the design of the prototype network for multi-label scenarios, the separability of functional labels in the embedding space is further improved;

[0060] In addition, during the iterative process of active learning, the prototype network dynamically updates the functional prototype vectors through a dynamic optimization mechanism. This dynamic optimization mechanism ensures that the prototype network can adapt to the changing label distributions and biological data characteristics, thereby further enhancing the prediction performance and adaptability of the model.

[0061] Specifically, assume that the initial labeled dataset contains n s labeled samples, and each sample consists of its gene expression feature vector and functional label vector. Among them, the functional label vector indicates whether each gene has a specific function (1 means present, 0 means absent). For each functional label, positive and negative prototypes are defined as the embedding means of the support samples with and without this function respectively. These prototypes represent the positive and negative class distributions of the functional label in the embedding space. Through this method, the prototype network can not only capture the distribution characteristics of a single functional label but also effectively model the biological associations between functional labels, providing support for the subsequent multi-label active learning process.

[0062] The calculation formulas for the positive and negative prototypes p c and are as follows:

[0063]

[0064] Among them, n s represents the set of labeled gene samples, represents the labeled sample which is the prototype of the functional label obtained by processing through the prototype network, represents the sample and is the c-th functional label in the sample. I(·) represents the indicator function, represents the existence of this function, while represents the non-existence of this function.

[0065] The above formula calculates the prototype by taking the average of the embeddings of samples with or without the c-th label respectively. Specifically, the closer the gene sample is to the positive prototype, the higher the probability that the sample has this function, while the closer the sample is to the negative prototype, the lower the probability that the sample has this function.

[0066] To quantify this probability relationship this embodiment adopts a distance-based probability estimation model.

[0067] Specifically, by calculating the Euclidean distances between the query unlabeled instance and the positive and negative class prototypes of label c in the embedding space, and using the softmax function to convert these distances into probability values to estimate the probability that the query instance has label c. When this probability exceeds a preset threshold, it is predicted that the gene has the corresponding function.

[0068]

[0069] Among them, represents the Euclidean distances between the unlabeled instance and the positive and negative class prototypes of label c in the embedding space, and n q represents the set of unlabeled gene samples.

[0070] This probability estimation method based on dynamic prototypes significantly improves the accuracy of gene function prediction and at the same time adapts to the sparse distribution characteristics of functional labels.

[0071] To further improve the performance of the prototype network, this embodiment optimizes the network parameters by minimizing the adaptive weighted cross-entropy loss function where S(c) is the label distribution imbalance weight,

[0072]

[0073] is the prediction result of the prototype model for the sample label. This loss function ​ By dynamically adjusting the weight S(c) of the labels, the ability to prioritize the attention to sparse functional labels is enhanced, while ensuring that the learning effect of common functional labels is not affected.

[0074] Specifically, the adaptive weighted cross-entropy loss function can dynamically allocate weights according to the scarcity of the labels, enabling sparse functional labels to obtain a higher learning priority during the optimization process. Through the optimization of this loss function, the prototype network can more accurately model the relationship between gene samples and functional labels in the embedding space, pulling the embedding vector of each gene sample closer to the positive class prototype of its corresponding function and farther away from the negative class prototype, thereby forming a clearer and more reliable functional division boundary in the gene function prediction task.

[0075] By iteratively optimizing the above loss function, the prototype network of this embodiment learns the positions of the positive and negative class prototypes and the parameters θ of the prototype network; finally, by optimizing the loss function of the query samples, the optimized network parameters θ are obtained * and an embedding space that can effectively separate samples of different labels is constructed, and the optimized prototypes in this space provide a basis for subsequent uncertainty estimation techniques.

[0076]

[0077] The prototype network of this embodiment has significant advantages in the multi-label gene function prediction task. First, through the construction and dynamic optimization of the functional label prototype vectors, this embodiment can model the complex correlations of multi-functional genes, capture the deep semantic relationships between various functional labels in multi-label data, thereby effectively avoiding the limitations of traditional statistical methods in dealing with complex functional co-occurrences and associations. Second, by combining the joint modeling of positive and negative class prototypes and the cross-entropy loss function designed specifically for multi-label tasks, this embodiment is particularly outstanding in the identification of rare functions and the adaptability to complex functional label distributions, significantly improving the prediction ability for sparse labels while taking into account the accuracy of high-frequency labels. Finally, through the mechanism of dynamically updating the functional label prototype vectors, this embodiment can quickly adjust the model parameters in each active learning iteration to adapt to the dynamic changes of new functions in multi-label data, thereby achieving efficient learning of newly annotated genes and significantly improving the overall performance and generalization ability of multi-label gene function prediction.

[0078] This embodiment designs a two-stage gene-function pair selection strategy to ensure that the selected gene-function pairs can provide effective biological information and cover the overall gene expression distribution.

[0079] Stage 1: Gene-function pair selection.

[0080] 1.1. Uncertainty.

[0081] This method quantifies the prediction uncertainty of gene-function pairs by analyzing the proximity of the embedding vectors of unlabeled gene samples to the positive and negative class prototypes of functional labels. If a gene sample is similarly distant from both the positive and negative class prototypes of a certain functional label, it indicates a high uncertainty in its true function. This method is different from traditional methods that rely on classifier prediction results or distances to decision boundaries, and can capture the ambiguity of gene functions in the embedding space more meticulously, thus providing a more robust uncertainty estimate.

[0082] Specifically, the information entropy is calculated through the probability of a gene sample having a certain function, serving as a quantification index for the uncertainty U(x i ,c). First, the probability that a gene-function pair is a positive class is calculated based on the distance relationship between the sample and the positive and negative class prototypes. Then, the information entropy is calculated based on this probability. A higher information entropy indicates a higher uncertainty in function prediction. In this way, the method of this embodiment can effectively measure the uncertainty of gene function prediction, providing a more reliable basis for active learning selection.

[0083]

[0084] Among them, a higher information entropy indicates a higher uncertainty. This uncertainty measure combines the distance relationship between the positive and negative class prototypes, providing a more effective sample selection strategy for active learning.

[0085] Compared with traditional methods, the method of this embodiment has significant advantages in calculating the uncertainty of gene function prediction. Traditional methods usually measure uncertainty through the prediction confidence of classifiers or distances to decision boundaries. However, in the initial stage of active learning, due to the scarcity of functionally labeled data, the prediction results of classifiers are usually noisy, and uncertainty estimates are often inaccurate, easily leading to the selection of suboptimal gene samples. The method of this embodiment quantifies uncertainty through the distance relationship between positive and negative class prototypes in the embedding space, avoiding the drawbacks of directly relying on classifier outputs, and thus showing higher robustness in the stage of scarce gene function data.

[0086] In addition, traditional decision boundary-based methods ignore the geometric distribution of samples in the feature space and only focus on the distance of samples to the classification hyperplane. Therefore, it is difficult to accurately distinguish the sources of uncertainty. The method of this method can capture the ambiguity of samples more meticulously by combining the distance relationship between positive and negative class prototypes. For example, for samples that are close to both positive and negative class prototypes, their uncertainty will be accurately quantified as a high value and will be preferentially selected for annotation, thus improving the efficiency of active learning.

[0087] In addition, the method of this embodiment has stronger adaptability in the scenario of multifunctional gene prediction. Since genes often have multiple functions and there are complex correlation relationships between functions, the uncertainty estimation based on the prediction results of classifiers in traditional methods will be interfered by the relationships between functions, making it difficult to accurately evaluate the uncertainty of a single function. The method of this embodiment calculates the uncertainty through the local embedding geometric relationship of gene-function pairs, which can avoid the interference of function relationships on the uncertainty estimation of single functions, and thus is more suitable for the multifunctional gene prediction scenario.

[0088] Generally speaking, the uncertainty calculation method based on the prototype network proposed in this embodiment can more accurately capture uncertain samples when the labeled gene data is scarce in the initial stage, providing a more efficient and robust selection strategy for active learning.

[0089] 1.2. Label correlation.

[0090] This embodiment proposes a novel label correlation modeling mechanism that incorporates the biological correlation between gene functions into the information quantity evaluation to improve the efficiency of active learning.

[0091] This mechanism quantifies the label correlation by calculating the cosine similarity between the positive class prototype vectors of function labels, generating a label correlation matrix.

[0092] Specifically, for functions c 1 and c 2 , their correlation is quantified by the cosine similarity of the positive class prototype vectors.

[0093]

[0094] Thus, a label correlation matrix W is generated. This method uses this matrix to define a weight L(x i , c) for each gene-function pair (x i , c) to measure the information quantity of querying this pair. The basic principle of this weight mechanism is as follows: If function c is highly correlated with the known functions of the gene, then querying function c may provide less information. On the contrary, if the correlation is low, then querying this function may reveal more information about the true function of the gene.

[0095]

[0096] Among them, Y i + represents the set of labels known to be positive classes for the sample x i in the current unlabeled set, and c′ is a certain label in Y i + .

[0097] Compared with existing methods, the proposed method for calculating tag correlation in this embodiment has significant advantages in multiple aspects. Traditional methods usually rely on Gene Ontology (GO) or functional co-occurrence statistics to model tag correlation. However, these statistical methods are mainly based on global distribution and are difficult to capture the deep biological relationships between functions, especially in cases where functional annotations are incomplete or new functions are discovered, showing significant limitations. The method in this embodiment realizes the deep modeling of tag correlation through the functional embedding vectors generated by the prototype network, which can more accurately reflect the biological connections between functions. In addition, statistical methods are easily affected by noisy data or label imbalance, resulting in inaccurate correlation estimation. This method realizes the deep modeling of tag correlation through the tag embedding vectors generated by the prototype network. Specifically, the prototype network can capture the semantic features of tags in the embedding space, and by calculating the cosine similarity between positive class prototype vectors, it can more precisely quantify the correlation between tags. This approach can not only avoid the influence of noise and label imbalance but also more accurately reflect the implicit relationships between gene functions.

[0098] In addition, different from some traditional methods that directly use gene tag correlation as the basis for sample selection, this method further combines the tag correlation matrix and the query weights of sample-tag pairs to dynamically evaluate the information content of each gene-function pair. Through this dynamic weight mechanism, tag correlation is used to actually optimize the functional verification strategy. For example, in some genes, if the function to be verified has a high biological correlation with known functions, the system will preferentially select other functions with lower biological correlations for verification in order to maximize the biological information gain. This mechanism not only effectively avoids redundant functional verification experiments but also can better cover the functional space and improve the efficiency of gene function prediction.

[0099] There is also a key problem in traditional methods for tag correlation modeling, that is, they usually assume that the correlation between gene functions is static and consistent across all genes. However, in actual biological scenarios, the tag correlation of different genes may vary depending on cell type, tissue specificity, or environmental conditions. In contrast, this embodiment calculates tag correlation dynamically through the features of the embedding space, thus being able to better adapt to the context-dependence of functional relationships. This dynamic modeling method can more accurately reflect the true biological characteristics of genes, further improving the adaptability and robustness of the active learning strategy in function prediction.

[0100] The label correlation calculation method of this embodiment solves the limitations of inaccurate label correlation estimation, redundant experimental verification, and static modeling in traditional methods by introducing a prototype network and a dynamic weight mechanism. It can more effectively capture deep functional relationships in a complex biological function distribution environment and preferentially select gene-function pairs with the largest amount of information for verification, thus significantly improving the efficiency of gene function prediction and the performance of the model.

[0101] 1.3. Label distribution imbalance weight.

[0102] In the gene function prediction task, the functional label distribution is usually significantly imbalanced. Some common functions appear with a higher frequency, while some important biological functions are very rare. This imbalance may cause the model to be more inclined to predict common functions and ignore rare but potentially biologically significant functions, ultimately affecting the overall prediction performance. To solve this problem, this method proposes an optimization mechanism based on "functional distribution imbalance weight", which can improve the coverage ability of important rare functions in the active learning process by dynamically adjusting the priority of rare functions.

[0103] The label distribution imbalance weight is quantified according to the occurrence frequency of the label, and its definition is as follows:

[0104]

[0105] where Y i represents the set of true labels of sample x i in the labeled dataset, n is the total number of samples with label c in the labeled data, is the indicator function. A higher S(c) value indicates that label c is rarer.

[0106] This embodiment has outstanding advantages in improving the prediction ability of rare gene functions: by assigning higher weights to rare functions and preferentially selecting genes that may carry rare function labels for annotation, it can accumulate more rare function training samples in the early stage, helping the model accurately capture these function labels with great scientific research value; at the same time, its mechanism of dynamically adjusting function priorities can update the scarcity weights in real time according to the latest annotation results. When new functions are discovered, it can immediately emphasize the annotation and attention to them, avoiding being ignored by traditional static weight strategies; in addition, the computational cost of this method in large-scale gene function prediction is relatively low, only requiring incremental statistics based on the function distribution of the labeled data, and it is more suitable for efficient application in biological research scenarios with a large data scale and rapidly updated function labels.

[0107] By introducing the label distribution imbalance weight, this method effectively alleviates the problem of performance degradation caused by label imbalance in multi-label learning. Its remarkable performance in sparse label coverage, dynamic adjustment strategy, and computational efficiency enables this method to achieve more efficient active learning in multi-label tasks.

[0108] 1.4. Calculation of the comprehensive information content of gene-function pairs.

[0109] To comprehensively measure the query information content of gene-function pairs, this method combines the uncertainty U(x i , c) of gene function prediction, the label correlation weight L(x i , c), and the label distribution imbalance weight S(c). The final information content score φ(x i , c) is:

[0110] φ 1 (x i , c) = U(x i , c) · L(x i , c) · S(c)

[0111] Select the n 1 sample-label pairs with the highest information content scores to form the unlabeled dataset

[0112]

[0113] where rankφ 1 (x i , c)) represents the sorting result of φ 1 (x i , c). These n 1 pairs represent the current most valuable function prediction tasks.

[0114] Stage 2: Optimization of gene-function pair selection.

[0115] As Figure 3 shown, this stage refines the initial candidate batch (containing n 1 sample-label pairs) obtained in Stage 1 into a smaller and more representative batch with a size of n 2 . The batch-mode selection may introduce redundant function annotation tasks, thus increasing the cost of biological experiment verification. Therefore, by selecting gene-function pairs that are diverse at the same time, the information gain is maximized, while the redundancy is minimized in the dimensions of gene expression patterns and functional labels.

[0116] The whole process is constructed in an iterative manner First, from Select the gene-function pair with the highest information content score as the seed and initialize

[0117]

[0118] Subsequently, when adding new gene-function pairs to consider both the initial information content score and the diversity with the current gene-function pairs in it.

[0119] 2.1 Multidimensional Diversity Optimization.

[0120] Diversity is used to measure the biological difference degree between the candidate function c and the functions already included in the batch . This metric is crucial for minimizing redundant experiments and maximizing the information gain of each function verification. A higher diversity value indicates a lower biological similarity between the candidate function and the existing functions, so verifying this function may provide more new biological insights.

[0121] In active learning for gene function prediction, traditional diversity optimization methods usually only focus on the distribution of genes in the expression profile space, for example, measuring diversity by the Euclidean distance between gene expression patterns. However, this method ignores the biological relationships between functions and the interaction information of gene-function pairs, which may lead to over - bias towards specific function categories and fail to comprehensively cover the biological function space. Especially in the scenario of multifunctional gene prediction, relying solely on the single - dimensional optimization of the expression profile space may be difficult to ensure the functional diversity and biological representativeness of the selected samples.

[0122] This method proposes a multidimensional diversity optimization mechanism. By comprehensively optimizing from three dimensions: gene expression diversity, functional space diversity, and gene - function pair interaction diversity, the selected gene - function pairs have good coverage in both the gene expression profile and the functional distribution. This mechanism can capture the complex distribution characteristics of biological data more comprehensively and avoid the limitations of traditional single - dimensional optimization.

[0123] Multidimensional diversity optimization includes the following three main dimensions:

[0124] ① Feature Diversity: Measure the distribution difference of candidate genes in the embedding space, and the formula is as follows:

[0125]

[0126] Among them, represents the feature diversity of the candidate sample x i , is the current selected sample set, ∥x i - x j ∥ 2is sample x i and x j 's Euclidean distance.

[0127] ② Label space diversity: Based on the dynamically updated label correlation matrix, measure the difference between candidate functions and selected functions. The formula is as follows:

[0128]

[0129] where represents the label diversity of candidate label c, and W(c, c′) is the dynamically updated label correlation matrix.

[0130] ③ Sample-label pair interaction diversity: Combine expression profile diversity and functional space diversity to comprehensively measure the biological diversity of gene-function pairs. The formula is as follows:

[0131]

[0132] Compared with traditional methods, the multi-dimensional diversity optimization of this method shows significant advantages in terms of biological comprehensiveness, functional coverage, and prediction robustness. Traditional methods only focus on the expression distance between genes, ignoring functional relationships and gene-function interaction information. In contrast, this method optimizes diversity from three dimensions: gene expression, functional labels, and gene-function pairs, ensuring that the selected gene-function pairs comprehensively reflect the biological distribution. The multi-dimensional optimization not only improves the coverage ability of complex functional distributions, making the model more adaptable to the global distribution of gene functions, but also avoids relying on a single dimension by combining multi-dimensional information, showing stronger robustness in the scenario of high-dimensional gene expression data and significantly improving the accuracy of function prediction and the generalization ability of the model.

[0133] 2.2. Calculation of the comprehensive information content of sample-label pairs.

[0134] The initial information content score φ 1 (x i , c) is updated to incorporate the diversity factor:

[0135]

[0136] The gene-function pair with the highest score will be added to , and the process is repeated until contains n 2 gene-function pairs:

[0137]

[0138] Through this iterative approach, the comprehensively optimal gene-function pairs finally selected are ensured in terms of information content and biological diversity, providing the most valuable candidate set for subsequent experimental verification. This optimization strategy not only improves the accuracy of function prediction but also significantly reduces the cost and time investment in biological experimental verification.

[0139] By making technical improvements to address the deficiencies in existing gene function prediction technologies, positive effects in multiple aspects are achieved, mainly reflected in enhancing function prediction performance, optimizing function annotation efficiency, strengthening model robustness and biological adaptability, etc. The beneficial effects of the present invention are described as follows:

[0140] I. Improving the effectiveness of learning in the initial stage

[0141] In the initial stage of traditional gene function prediction multi-label active learning methods, due to the scarcity of function annotation data, the model often has difficulty accurately estimating the uncertainty of gene-function pairs, resulting in the selection of suboptimal gene samples. To address this issue, the present invention introduces a pre-training mechanism based on prototype networks. Through a small number of function annotation samples, the biological relationships between gene functions and the characteristics of expression profile distributions are learned, thus significantly enhancing the effectiveness of function prediction in the initial stage. Through the positive and negative class prototypes of function labels generated by the prototype network, the model can achieve a more accurate estimation of function prediction uncertainty in the initial stage. This mechanism effectively alleviates the learning bias problem caused by over-reliance on unreliable function predictions in traditional methods, ensures the quality of gene-function pair selection, and lays a more solid biological foundation for the subsequent learning of function prediction models.

[0142] II. Accurately modeling label correlations and avoiding redundant queries

[0143] Label correlation is a key factor in gene function prediction. Existing technologies usually model it through simple statistical methods (such as function co-occurrence matrices). This method is difficult to capture deep biological relationships in cases where function annotation is sparse or the distribution is complex, leading to redundant function verification experiments and waste of experimental resources. The present invention learns the biological relationships between functions through functional prototype vectors and constructs a label correlation matrix based on this, enabling the accurate modeling of correlations between functions at a higher biological semantic level. This precise label correlation modeling not only avoids repeated verification of highly biologically correlated functions but also can more effectively explore unknown regions of the function space, thereby maximizing biological information gain. This feature is particularly applicable to gene function prediction scenarios with a large number of functions and complex distributions, significantly optimizing function verification efficiency.

[0144] III. Balancing label distribution and optimizing model performance

[0145] In gene function prediction tasks, the functional distribution is often highly imbalanced, with extremely few positive samples for rare functions, resulting in weak prediction ability of the model for rare functions. Traditional function prediction methods usually give priority to gene samples related to high-frequency functions, ignoring the attention to rare functions, thus limiting the comprehensive coverage of the model in the functional space.

[0146] IV. Sample selection strategy considering both diversity and information content

[0147] The present invention introduces a "functional distribution imbalance weight calculation" mechanism, which dynamically adjusts the selection priority of gene-function pairs according to the occurrence frequency of functions, ensuring the participation of rare functions in the active learning process. This mechanism can assign higher weights to rare functions, preferentially select gene samples related to rare functions for annotation, thereby significantly improving the prediction effect of rare functions. By dynamically adjusting the function priority, this mechanism makes the coverage of the model in the functional space more balanced, while enhancing the prediction ability for low-frequency functions. Theoretically, this method can significantly improve the global performance of the model without affecting the prediction of common functions, especially performing well in scenarios where the functional distribution is extremely imbalanced.

[0148] V. Improving annotation efficiency and reducing annotation costs

[0149] Traditional gene function prediction methods often have the problem of excessive redundant verification in batch sample selection, resulting in increased experimental costs and limited biological information gain, restricting the improvement of prediction performance. This method proposes a two-stage gene-function pair selection strategy, which significantly improves the efficiency and effectiveness of sample selection by comprehensively considering uncertainty, label correlation, label distribution imbalance, diversity, and biological representativeness.

[0150] In the first stage, combining uncertainty, label correlation, and label distribution imbalance, it preferentially screens out candidate gene-function pairs with the greatest biological information gain, ensuring that the initial selected sample set contains more critical biological signals. In the second stage, through a multi-dimensional diversity optimization mechanism, it further optimizes diversity and representativeness on the basis of the initial selected samples, comprehensively measuring the diversity of gene expression profiles, the diversity of the functional space, and the interaction diversity of gene-function pairs, ensuring that the finally selected samples have good coverage in both the expression profile space and the functional space, avoiding excessive bias and repeated verification.

[0151] This strategy shows significant advantages in reducing redundant information in batch selection, while greatly improving the biological coverage and information content of samples, providing more comprehensive and rich learning information for the model. Especially in the scenario of high-dimensional sparse gene expression data, this method can effectively capture complex biological distribution characteristics, making the function prediction model more robust and generalization-capable.

[0152] 6. Enhance model robustness and generalization

[0153] The core goal of active learning for gene function prediction is to obtain the best function prediction effect with the least experimental verification cost. The present invention significantly reduces the biological experimental verification cost by optimizing the gene-function pair selection mechanism. Under the premise of ensuring the functional prediction performance, this method can achieve prediction performance comparable to or even better than traditional methods with fewer experimental verification samples. Through a two-stage selection strategy, the present invention avoids the need for experimental verification of redundant gene-function pairs and significantly improves the efficiency of functional verification. In addition, the label distribution imbalance weight mechanism further reduces the repeated verification of high-frequency functions, thereby focusing the limited experimental budget on the gene-function pairs with the largest biological information gain. This efficiency improvement is of great significance in practical application scenarios with limited biological experimental resources.

[0154] 7. Supporting the wide application of multi-label tasks

[0155] The present invention makes full use of the advantages of meta-learning. Through the pre-training of the prototype network, the model can quickly adapt to the gene expression distribution and functional structure under different biological conditions. This design significantly enhances the robustness of the model, so that it can still maintain stable functional prediction performance when facing different cell types, tissue specificity or environmental conditions. In addition, the present invention comprehensively considers label relevance and scarcity in the process of gene-function pair selection, avoids the learning bias caused by excessive focus on local biological information in traditional methods, and significantly improves the biological generalization ability of the model. This feature makes this method highly applicable in a variety of gene function prediction scenarios.

[0156] Gene function prediction tasks are widely present in biological research, including multiple levels such as molecular function, biological process and cellular components. The present invention can provide efficient and accurate solutions in various function prediction application scenarios through efficient label correlation modeling and scarcity optimization strategy.

[0157] For example:

[0158] In molecular function prediction, this method can effectively handle complex genes with multiple functions at the same time;

[0159] In the prediction of biological processes, the present invention significantly outperforms traditional methods in predicting gene functions involving multiple pathways;

[0160] In the prediction of cell components, the present invention can better capture the sparse relationship between genes and subcellular localization and provide more accurate prediction results.

[0161] These characteristics indicate that the present invention has broad application potential and can solve a variety of practical problems in gene function prediction.

[0162] VIII. Reducing Computational Resource Consumption and Environmentally Friendly

[0163] By reducing redundant verification and optimizing gene-function pair selection, the present invention significantly reduces the requirements for computational resources and experimental resources during the active learning process of function prediction. In batch mode, the present invention avoids frequent model retraining and reduces computational complexity through an efficient sample screening mechanism. This low-resource consumption characteristic not only reduces the operating cost in practical applications but also improves the efficiency of biological experiments, enabling more effective utilization of limited experimental resources.

[0164] Example 2

[0165] The purpose of this example is to provide a gene function prediction system based on multi-label active learning, including:

[0166] A construction module configured to: construct a training set of gene-function pairs in a multi-label scenario and generate a labeled data set and an unlabeled data set according to the training set;

[0167] A training module configured to: train a prototype network using the training set to obtain a trained prototype network; wherein, during the iterative training of the prototype network, gene-function pairs that make up an initial set are selected from the unlabeled data set according to the initial information scores quantified by the uncertainty of gene function prediction, label correlation, and label distribution imbalance;

[0168] According to the selection criteria of feature diversity, label space diversity, and sample-label pair interaction diversity, combined with the initial information scores corresponding to the gene-function pairs, an alternative set is selected from the initial set, and the prototype network is optimized according to the selected alternative set;

[0169] A prediction module configured to: perform prediction processing on the gene data to be predicted using the trained prototype network to obtain the prediction result of the gene data to be predicted.

[0170] In more examples, there is also provided:

[0171] An electronic device, including a memory, a processor, and computer instructions stored on the memory and running on the processor. When the computer instructions are run by the processor, the method described in Example 1 is completed. For the sake of brevity, it will not be elaborated here.

[0172] It should be understood that in this embodiment, the processor may be a central processing unit (CPU), or the processor may also be other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or any conventional processor, etc.

[0173] The memory may include a read-only memory and a random access memory, and provide instructions and data to the processor. A part of the memory may also include a non-volatile random access memory. For example, the memory may also store information about the device type.

[0174] A computer-readable storage medium for storing computer instructions, which when executed by the processor, implement the method described in Embodiment 1.

[0175] The method in Embodiment 1 can be directly implemented by a hardware processor, or implemented by a combination of hardware and software modules in the processor. The software module may be located in a mature storage medium in the art such as a random access memory, a flash memory, a read-only memory, a programmable read-only memory, or an electrically erasable programmable memory, a register, etc. This storage medium is located in the memory, and the processor reads the information in the memory and combines its hardware to complete the steps of the above method. To avoid repetition, it will not be described in detail here.

[0176] A computer program product, including a computer program, which when executed by the processor, implements the method described in Embodiment 1.

[0177] The present invention also provides at least one computer program product tangibly stored on a non-transitory computer-readable storage medium. The computer program product includes computer-executable instructions, such as instructions included in program modules, which are executed in a device on a target real or virtual processor to perform the process / method described above. Generally, program modules include routines, programs, libraries, objects, classes, components, data structures, etc. that perform specific tasks or implement specific abstract data types. In various embodiments, the functions of program modules may be combined or divided as needed. The machine-executable instructions for program modules may be executed locally or within a distributed device. In a distributed device, program modules may be located in local and remote storage media.

[0178] The computer program code for implementing the method of the present invention can be written in one or more programming languages. These computer program codes can be provided to the processors of general-purpose computers, special-purpose computers, or other programmable data processing devices, such that when the program codes are executed by the computers or other programmable data processing devices, the functions / operations specified in the flowcharts and / or block diagrams are implemented. The program codes can be executed entirely on the computer, partially on the computer, as an independent software package, partially on the computer and partially on a remote computer, or entirely on a remote computer or server.

[0179] In the context of the present invention, the computer program code or related data can be carried by any suitable carrier so that the devices, apparatuses, or processors can execute the various processes and operations described above. Examples of carriers include signals, computer-readable media, and the like. Examples of signals can include electrical, optical, radio, acoustic, or other forms of propagated signals, such as carrier waves, infrared signals, etc.

[0180] Those of ordinary skill in the art can realize that the units and algorithm steps of the examples described in conjunction with this embodiment can be implemented by electronic hardware or a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of this application.

[0181] Although the specific implementation manners of the present invention have been described above in conjunction with the accompanying drawings, it is not a limitation on the protection scope of the present invention. Those skilled in the art should understand that based on the technical solutions of the present invention, various modifications or deformations that can be made by those skilled in the art without creative efforts are still within the protection scope of the present invention.

Claims

1. A gene function prediction method based on multi-label active learning, characterized in that: include: Constructing a training set of gene-function pairs in a multi-label scenario, and generating a labeled data set and an unlabeled data set according to the training set; Using the training set to train the prototype network to obtain a trained prototype network; The trained prototype network is used to perform prediction processing on the gene data to be predicted, and the prediction result of the gene data to be predicted is obtained; wherein, in the iterative training of the prototype network, gene-function pairs constituting an initial set are selected from the unlabeled data set according to the uncertainty of gene function prediction, label relevance, and initial information scores quantified by label distribution imbalance; According to the selection criteria of feature diversity, label space diversity and sample-label pair interaction diversity, combined with the initial information scores corresponding to the gene-function pairs, gene-function pairs constituting the candidate set are selected from the initial set, and the prototype network is optimized according to the selected candidate set.

2. A gene function prediction method based on multi-label active learning as claimed in claim 1, characterized in that: According to the distance relationship between the gene sample and the label positive prototype and the label negative prototype respectively, the probability that the gene-function pair is a positive class is calculated; according to the calculated probability that the gene-function pair is a positive class, the information entropy is calculated, and the uncertainty of the gene function prediction is determined according to the information entropy; wherein, the higher the information entropy, the higher the uncertainty of the gene function prediction.

3. A gene function prediction method based on multi-label active learning as claimed in claim 1, characterized in that: Determining the label correlation is specifically as follows: calculating the cosine similarity between the functional label positive class prototype vectors to obtain a label correlation matrix, assigning corresponding weights to corresponding gene-function pairs according to the label correlation matrix, and measuring the amount of information of the corresponding gene-function pairs according to the weights.

4. A gene function prediction method based on multi-label active learning as claimed in claim 1, characterized in that: The imbalance of label distribution is quantified according to the frequency of label occurrence, and genes that may carry rare labels are annotated by dynamically adjusting the priority of rare labels.

5. A gene function prediction method based on multi-label active learning as claimed in claim 1, characterized in that: The calculation of the feature diversity, label space diversity and gene-function pair interaction diversity is specifically as follows: According to the Euclidean distance between gene samples, the distribution differences of candidate gene samples in the embedding space are measured to determine the quantitative results of feature diversity; According to the dynamically updated label correlation matrix, the difference between the candidate labels and the selected labels is measured to determine the quantitative results of the diversity of the label space; The quantified result of the feature diversity and the quantified result of the label space diversity are multiplied to obtain the quantified result of the gene-function pair interaction diversity.

6. A gene function prediction method based on multi-label active learning as claimed in claim 1 or 5, characterized in that: The initial information volume score of the gene-function pair is multiplied by the quantified result of the corresponding interaction diversity of the gene-function pair to obtain the sample-label pair comprehensive information volume of the gene-function pair, and the gene-function pair with the highest sample-label pair comprehensive information volume score is selected into the candidate set, and the update is continuously iterated until the number of gene-function pairs in the candidate set reaches the set number.

7. A gene function prediction system based on multi-label active learning, characterized in that: include: A construction module is configured to: construct a training set of gene-function pairs in a multi-label scenario, and generate a labeled data set and an unlabeled data set according to the training set; A training module is configured to: train the prototype network using the training set to obtain a trained prototype network; wherein, in the iterative training of the prototype network, select gene-function pairs constituting an initial set from the unlabeled data set according to the uncertainty of gene function prediction, label relevance, and initial information scores quantified by label distribution imbalance; According to the selection criteria of feature diversity, label space diversity and sample-label pair interaction diversity, combined with the initial information scores corresponding to the gene-function pairs, a candidate set composed of the components is selected from the initial set, and the prototype network is optimized according to the selected candidate set; The prediction module is configured to: use the trained prototype network to perform prediction processing on the gene data to be predicted to obtain the prediction result of the gene data to be predicted.

8. An electronic device, characterized in that: The method comprises a memory and a processor and computer instructions stored in the memory and executed on the processor, wherein when the computer instructions are executed by the processor, the method according to any one of claims 1 to 6 is completed.

9. A computer-readable storage medium, characterized in that: Used to store computer instructions, which, when executed by a processor, complete the method described in any one of claims 1 to 6.

10. A computer program product, characterized in that The invention comprises a computer program, which, when executed by a processor, implements the method according to any one of claims 1 to 6.

Citation Information

Cited By

  • Gene function characterization method and device

    CN122201460A