Amino acid sequence data classification method and system based on machine learning

By building a machine learning model containing multiple deep learning layers and using generative adversarial networks for training, the problems of low efficiency and insufficient accuracy of amino acid sequence classification in the prior art are solved, and more efficient and accurate classification effects are achieved.

CN120072060APending Publication Date: 2025-05-30GUANGDONG POLYTECHNIC OF IND & COMMERCE +2
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202411991842.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-31
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

The prior art has problems in the classification of amino acid sequences that are inefficient, high cost, difficult to deal with complex nonlinear relationships, and insufficient ability to generalize new sequences.

Method used

A method of classification of amino acid sequence data based on machine learning is proposed. By obtaining and standardizing amino acid sequence data, a classification model including embedding layer, feature mapping layer, position coding layer, feature extraction layer, fuzzing processing layer, feature fusion module, bidirectional long and short-term memory network, text convolutional network layer, feature aggregation network encoding layer, full connection layer and output layer is constructed, and a classification model is trained using the generative adversarial network, optimizer and loss function.

Benefits of technology

It significantly improves the classification accuracy and efficiency of amino acid sequence data, enhances the generalization ability of the model and captures complex patterns.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120072060A_ABST
    Figure CN120072060A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of amino acid sequence classification, and discloses an amino acid sequence data classification method and system based on machine learning, and the method comprises the steps: obtaining amino acid sequence data, and dividing the amino acid sequence data into a training set and a test set; performing standardization and biological feature coding processing on the training set in sequence to obtain standardized amino acid sequence data and biological features; the method comprises the following steps: constructing a classification model for classifying amino acid sequence data, and a generative adversarial network, an optimizer and a loss function which are used for training and optimizing the classification model, training the classification model until a training round number is reached, and obtaining a trained classification model; and inputting the amino acid sequence data of the test set into the trained classification model for classification processing to obtain an amino acid sequence data classification result. According to the method, efficient and accurate classification of the amino acid sequence data is realized, so that the classification accuracy of the multifunctional therapeutic peptide is improved, and the classification efficiency is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of amino acid sequence classification, and more specifically, to a method and system for classifying amino acid sequence data based on machine learning. Background Art

[0002] The classification of amino acid sequence data is an important task in the field of bioinformatics. By extracting features from amino acids in protein sequences, biological properties such as the function and structure of proteins can be predicted and classified. With the rapid development of genomics and proteomics, the classification of amino acid sequences has been widely applied in fields such as drug research and development, disease prediction, and gene function annotation. In particular, multifunctional therapeutic peptides, as a class of short-chain amino acid sequences with multiple biological activities, exhibit functions such as antibacterial, anti-inflammatory, and anti-tumor, and are widely used in the fields of agriculture, medicine, and microbial research. The development of efficient classification and prediction methods for them is of great significance.

[0003] Currently, traditional amino acid sequence classification methods mainly rely on biochemical experiments and rule-based sequence alignment techniques. Experimental methods directly measure the biological activity data of multifunctional therapeutic peptides through cell culture, animal model testing, etc., but usually take a long time and are costly, making it difficult to meet the needs of large-scale research. In addition, sequence alignment methods are usually based on linear matching techniques, making it difficult to effectively handle complex non-linear relationships and lacking high-throughput screening capabilities. In recent years, machine learning and deep learning technologies have gradually been applied to amino acid sequence classification. Through models such as convolutional neural networks (CNNs) and recurrent neural networks (RNNs), potential features in sequences are extracted to achieve automated classification prediction and improve efficiency. However, these methods still have limitations in feature extraction and model construction.

[0004] The main problems of current technologies in amino acid sequence classification include the following points: First, traditional experimental methods are costly and time-consuming and are not suitable for large-scale research; rule-based methods based on sequence alignment lack the ability to model complex non-linear relationships. Second, although deep learning methods can capture complex sequence patterns, their training process is highly dependent on large-scale labeled data and their generalization ability for new sequences is insufficient, which may lead to a decrease in the reliability of classification results. At the same time, existing deep learning models perform limitedly in dealing with long-distance dependence relationships, integrating multi-source information, and coping with the ambiguity and continuity in feature expression, affecting classification accuracy and flexibility. Therefore, there are still significant deficiencies in efficiency and accuracy in the existing technologies, and further optimization and improvement are urgently needed. Summary of the Invention

[0005] In order to improve the efficiency and accuracy of the multifunctional therapeutic peptide classification technology, the present invention proposes the following technical solutions: In the first aspect, the present invention proposes a method for classifying amino acid sequence data based on machine learning, including: Obtain the amino acid sequence data of multifunctional therapeutic peptides with functional labels, and divide the amino acid sequence data into a training set and a test set.

[0006] Successively perform normalization and biometric encoding processing on the amino acid sequence data in the training set to obtain normalized amino acid sequence data and biometric features.

[0007] Construct a classification model for classifying amino acid sequence data, as well as a generative adversarial network, an optimizer, and a loss function for training and optimizing the classification model.

[0008] Based on the optimizer and the loss function, use the normalized amino acid sequence data, biometric features, and generative adversarial network to train the classification model until the number of training rounds is reached, and obtain a trained classification model.

[0009] Input the amino acid sequence data of the test set into the trained classification model for classification processing to obtain the classification result of the amino acid sequence data.

[0010] As a preferred technical solution, the normalization processing of the amino acid sequence data includes: Use a standard amino acid character set containing 20 natural amino acids and 1 unknown amino acid to map the characters of each amino acid sequence data to a unique integer index to generate a numerical amino acid sequence.

[0011] Set the maximum sequence length. For the numerical amino acid sequences shorter than the maximum sequence length, perform zero-padding at the end. For the numerical amino acid sequences longer than the maximum sequence length, intercept the front sequence to the maximum sequence length.

[0012] Convert the functional label of each amino acid sequence into a binary vector, where each component of the binary vector represents a function. If the amino acid sequence has this function, the corresponding component takes the value of 1, otherwise it takes the value of 0.

[0013] As a preferred technical solution, the classification model includes an embedding layer, a feature mapping layer, a position encoding layer, a feature extraction layer, a fuzzification processing layer, a feature fusion module, a bidirectional long short-term memory network, a text convolution network layer, a feature aggregation network encoding layer, a fully connected layer, and an output layer.

[0014] The embedding layer is used to map the numerical amino acid sequence into a high-dimensional vector space to obtain a first feature representation.

[0015] The feature mapping layer is used to perform feature mapping on the biological features of the amino acid sequence to obtain a second feature representation with the same dimension as the first feature representation.

[0016] The position encoding layer is used to generate position encoding and add the position encoding to the first feature representation to obtain a third feature representation containing position information.

[0017] The feature extraction layer is used to perform feature learning on the third feature representation in parallel using multiple attention heads to generate a fourth feature representation for each position under different attention heads.

[0018] The fuzzification processing layer is used to perform fuzzification processing on the second feature representation to obtain a fifth feature representation.

[0019] The feature fusion module is used to perform fusion processing on the first feature representation, the fourth feature representation, and the fifth feature representation to obtain a sixth feature representation containing comprehensive features.

[0020] The bidirectional long short-term memory network is used to perform bidirectional feature extraction on the sixth feature representation to obtain a seventh feature representation containing forward and backward information.

[0021] The text convolutional network is used to perform a convolutional operation on the seventh feature representation to extract local features of the sequence and obtain an eighth feature representation.

[0022] The feature aggregation network encoding layer is used to perform aggregation encoding on the eighth feature representation to integrate multi-level feature information and obtain a ninth feature representation.

[0023] The fully connected layer is used to perform layer-by-layer dimensionality reduction processing on the ninth feature representation to extract high-order features and obtain a tenth feature representation.

[0024] The output layer is used to perform classification output on the tenth feature representation, map the classification result to the interval from 0 to 1, and obtain predicted values consistent with the number of functional categories.

[0025] As a preferred technical solution, based on the optimizer and the loss function, the classification model is trained using standardized amino acid sequence data, biological features, and a generative adversarial network, including: Using the standardized amino acid sequence data and biological features as real data to train the generative adversarial network to obtain a trained generative adversarial network.

[0026] Using the trained generative adversarial network to generate amino acid sequence data as fake data.

[0027] Performing merging processing on the real data and the fake data to obtain an enhanced training set.

[0028] Input the enhanced training set into the classification model and calculate the classification loss based on the loss function.

[0029] Perform backpropagation according to the classification loss, update the parameters of the classification model, and dynamically adjust the learning rate during the training process of the classification model based on the optimizer until the number of training rounds is reached to obtain the trained classification model.

[0030] As a preferred technical solution, the generative adversarial network includes a generator and a discriminator. Training the generative adversarial network using the standardized amino acid sequence data and biological features as real data includes: Convert the standardized amino acid sequence data and biological features as real data into one-hot representation and input them into the discriminator, and calculate the discrimination loss of the real data by calculating the output probability of the real data.

[0031] Sample random noise from the standard normal distribution, synthesize amino acid sequence data as fake data based on the random noise through the generator, input the fake data into the discriminator, and calculate the discrimination loss of the fake data by calculating the output probability of the fake data.

[0032] Update the parameters of the discriminator and the generator through backpropagation according to the discrimination loss of the real data and the discrimination loss of the fake data.

[0033] As a preferred technical solution, the optimizer uses the Adam optimizer and combines it with a cosine learning rate scheduler to dynamically adjust the learning rate in stages through cosine annealing and warm-up annealing.

[0034] Among them, the expression of the cosine learning rate scheduler is as follows:

[0035] In the formula, represents the warm-up annealing stage, represents the cosine annealing stage, represents the learning rate at the start of warm-up annealing, represents the current step number, represents the number of steps in the warm-up annealing stage, represents the initial learning rate, represents the final learning rate, represents the number of steps used for cosine annealing, represents the total number of update steps.

[0036] As a preferred technical solution, the expression of the loss function of the classification model is as follows:

[0037] Among them, N is the total number of samples, represents the weight coefficient of the positive class samples, represents the predicted value of the positive class samples after clipping, represents the true label, represents the predicted value of the negative class after clipping, represents the weight adjustment parameter based on the category.

[0038] As a preferred technical solution, the biometric features include one or more of amino acid index features, pseudo amino acid composition features, amino acid physicochemical property features, amino acid pair substitution score features, or amino acid composition features.

[0039] As a preferred technical solution, after obtaining the trained classification model, the method further includes evaluating the classification model using evaluation metrics.

[0040] The evaluation metrics include Aiming, Coverage, Accuracy, Aboslute_True, and Aboslute_False, and the expressions of each evaluation metric are as follows:

[0041] In the formula, is used to measure the correct proportion in the classification labels, and n is the number of samples. represents the true positive count of the th sample, that is, the number of labels that the classification model predicts as positive and are actually positive. represents the false positive count of the

[0042]

[0043] In the formula, is used to measure the proportion of the labels correctly classified by the classification model among all the true labels, represents the false negative count of the

[0044]

[0045] In the formula, is used to measure the proportion of the correctly classified labels among all the relevant labels, including the correctly classified, misclassified, and unclassified true labels.

[0046]

[0047] Wherein, Absolute_True is used to measure the proportion of the predicted labels of the classification model that are exactly the same as the true labels. represents the predicted label vector of the th sample, is an indicator function used to check whether the predicted labels and the true labels exactly match.

[0048]

[0049] In the formula, is used to measure the proportion of misclassification of the classification model in the multi-label classification task. represents the number of labels for each sample. represents the h th predicted label vector of the th sample for the h th label.

[0050] In the second aspect, the present invention also proposes a machine learning-based amino acid sequence data classification system, which is applied to the machine learning-based amino acid sequence data classification method described in any one of the schemes in the first aspect, and includes: An acquisition module for acquiring the amino acid sequence data of multifunctional therapeutic peptides with functional labels and dividing the amino acid sequence data into a training set and a test set.

[0051] A preprocessing module for sequentially performing normalization and biometric encoding processing on the amino acid sequence data in the training set to obtain normalized amino acid sequence data and biometrics.

[0052] A construction module for constructing a classification model for amino acid sequence data classification, as well as a generative adversarial network, an optimizer, and a loss function for training and optimizing the classification model.

[0053] A training module for training the classification model based on the optimizer and the loss function, using the normalized amino acid sequence data, biometrics, and the generative adversarial network until the number of training rounds is reached, to obtain a trained classification model.

[0054] A classification and testing module for inputting the amino acid sequence data of the test set into the trained classification model for classification processing to obtain the amino acid sequence data classification result.

[0055] The beneficial effects of the present invention at least include: The present invention provides a good data foundation for subsequent processing by obtaining the amino acid sequence data of multifunctional therapeutic peptides and dividing them into training sets and test sets. Through standardization and biometric encoding of the training set data, the key features of the amino acid sequences are effectively extracted, providing high-quality inputs for the training of classification models. By constructing a classification model and introducing a generative adversarial network, an optimizer, and a loss function, a powerful machine learning framework is formed, which can effectively capture the complex patterns of amino acid sequences. Using the optimizer and the loss function, combined with the generative adversarial network, the classification model is trained, significantly improving the generalization ability and classification accuracy of the classification model. Finally, by inputting the test set data into the trained classification model for classification, efficient classification of new amino acid sequence data is achieved. This method not only improves the accuracy of multifunctional therapeutic peptide classification but also greatly enhances the classification efficiency. By integrating multiple machine learning techniques, the present invention can analyze amino acid sequence data more comprehensively, providing reliable technical support for the function prediction of multifunctional therapeutic peptides, helping to accelerate the new drug R & D process, reduce laboratory verification costs, and providing innovative methodological support for research in the field of bioinformatics. BRIEF DESCRIPTION OF THE DRAWINGS

[0056] Figure 1 FIG. is a schematic flowchart of a method for classifying amino acid sequence data based on machine learning provided by an embodiment of the present application.

[0057] Figure 2 FIG. is a data processing flowchart of a classification model provided by an embodiment of the present application.

[0058] Figure 3 FIG. is a training flowchart of a generative adversarial network provided by an embodiment of the present application.

[0059] Figure 4 FIG. is a graph showing the change of the learning rate of a classification model with the training steps provided by an embodiment of the present application.

[0060] Figure 5 FIG. is an architecture diagram of a system for classifying amino acid sequence data based on machine learning provided by an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0061] The following will describe the embodiments of the present invention with reference to the accompanying drawings and preferred technical solutions. Those skilled in the art can easily understand the other advantages and effects of the present invention from the content disclosed in this specification. The present invention can also be implemented or applied through other different specific embodiments, and various details in this specification can also be modified or changed based on different viewpoints and applications without departing from the spirit of the present invention. It should be understood that the preferred technical solutions are only for explaining the present invention and not for limiting the protection scope of the present invention.

[0062] It should be noted that the illustrations provided in the following embodiments only schematically illustrate the basic concept of the present invention. Therefore, only the components related to the present invention are shown in the drawings, rather than being drawn according to the number, shape, and size of the components in actual implementation. The type, quantity, and proportion of each component in actual implementation can be arbitrarily changed, and the component layout type may also be more complex.

[0063] In the following description, a large number of details are explored to provide a more thorough explanation of the embodiments of the present invention. However, it is obvious to those skilled in the art that the embodiments of the present invention can be implemented without these specific details. In other embodiments, well-known structures and devices are shown in the form of block diagrams rather than in detail to avoid making the embodiments of the present invention difficult to understand.

[0064] Embodiment 1 This embodiment proposes a method for classifying amino acid sequence data based on machine learning, as Figure 1 shown, Figure 1 is a schematic flow diagram of a method for classifying amino acid sequence data based on machine learning provided in this embodiment. The method includes the following steps: S1: Obtain amino acid sequence data of multifunctional therapeutic peptides with functional labels, and divide the amino acid sequence data into a training set and a test set.

[0065] The dataset used in this example is the multifunctional therapeutic peptide dataset used by Henghui Fan et al. The dataset is divided into a training set and a test set, where there are 7899 training data and 1975 test data to ensure that the classification model can be effectively trained and evaluated. Each data sample consists of an amino acid sequence and a corresponding functional label. The amino acid sequence is represented by amino acid characters, and the functional label is a multi-dimensional vector indicating whether the peptide has specific biological functions, such as antibacterial, antiviral, and anticancer.

[0066] The dataset contains 21 functional types of peptides, including: anti-angiogenic peptides, antimicrobial peptides (ABP), anticancer peptides (ACP), antidiabetic peptides (ADP), anti-endotoxin peptides, antifungal peptides (AFP), anti-HIV peptides, anti-hypertensive peptides (AHP), anti-inflammatory peptides, anti-MRSA peptides, anti-parasite peptides, anti-tuberculosis peptides (ATP), antiviral peptides (AVP), blood-brain barrier penetrating peptides, biofilm inhibitory peptides, cell penetrating peptides, dipeptidyl peptidase IV peptides (DPPIP), quorum sensing inhibitory peptides, surface binding peptides, and tumor homing peptides.

[0067] S2: Sequentially perform normalization and biometric encoding processing on the amino acid sequence data in the training set to obtain normalized amino acid sequence data and biometric features.

[0068] S3: Construct a classification model for amino acid sequence data classification, as well as a generative adversarial network, an optimizer, and a loss function for training and optimizing the classification model.

[0069] S4: Based on the optimizer and the loss function, use the standardized amino acid sequence data, biological features, and the generative adversarial network to train the classification model until the number of training rounds is reached, and obtain a trained classification model.

[0070] S5: Input the amino acid sequence data of the test set into the trained classification model for classification processing to obtain the classification result of the amino acid sequence data.

[0071] It can be understood that in this embodiment, by obtaining the amino acid sequence data of multifunctional therapeutic peptides and dividing them into a training set and a test set, a good data foundation is provided for subsequent processing. The standardization and biological feature encoding processing of the training set data effectively extract the key features of the amino acid sequence, providing high-quality input for the training of the classification model. Constructing a classification model and introducing a generative adversarial network, an optimizer, and a loss function forms a powerful machine learning framework that can effectively capture the complex patterns of amino acid sequences. Using the optimizer and the loss function, combined with the generative adversarial network to train the classification model, significantly improves the generalization ability and classification accuracy of the classification model. Finally, by inputting the test set data into the trained classification model for classification, efficient classification of new amino acid sequence data is achieved. This method not only improves the accuracy of multifunctional therapeutic peptide classification but also greatly enhances the classification efficiency. By integrating multiple machine learning technologies, this embodiment can more comprehensively analyze amino acid sequence data, providing reliable technical support for the function prediction of multifunctional therapeutic peptides, helping to accelerate the new drug R & D process, reduce laboratory verification costs, and providing innovative methodological support for research in the field of bioinformatics.

[0072] Embodiment 2 This embodiment makes improvements on the basis of the machine learning-based amino acid sequence data classification method proposed in Embodiment 1.

[0073] In this embodiment, the standardization processing of amino acid sequence data includes: Read the amino acid sequence data from a text file, parse and store it in memory for subsequent data processing operations. Ensure that each amino acid sequence corresponds one-to-one with its corresponding functional label during the reading process to avoid data misalignment and maintain the integrity and consistency of the data set. In this way, ensure that the classification model can correctly associate each amino acid sequence with the corresponding functional label during training and prediction.

[0074] Using a standard amino acid character set containing 20 natural amino acids and 1 unknown amino acid, map the characters of each amino acid sequence data to a unique integer index to generate a numerical amino acid sequence.

[0075] In the specific implementation process, define a standard amino acid character set, including 20 natural amino acids and a symbol "X" representing unknown or special amino acids. Therefore, the final character set is "XACDEFGHIKLMNPQRSTVWY", a total of 21 characters. Map each amino acid character to a unique integer index to form a character-to-index mapping table for converting amino acid characters into numerical representations, facilitating the processing of the classification model. During the encoding process, convert each amino acid character in the amino acid sequence into the corresponding integer index in sequence to obtain a numerical sequence representation, which is convenient for the input of the classification model.

[0076] Set the maximum sequence length. For numerical amino acid sequences shorter than the maximum sequence length, pad with zeros at the end. For numerical amino acid sequences longer than the maximum sequence length, truncate the front part of the sequence to the maximum sequence length.

[0077] As an exemplary illustration, in this embodiment, a maximum sequence length of 50 is set to cover the length requirements of most amino acid sequences. Unifying the length of the input sequences can adapt to the input requirements of the classification model. For sequences shorter than the maximum length, pad with zero values at the end until the sequence length reaches the maximum length. Padding with zero values facilitates the classification model not to be interfered by the short sequence length when processing variable-length sequences, ensuring the consistency of the data format. For sequences longer than the maximum length, intercept the first 50 characters from the beginning of the sequence to ensure that the length of the intercepted sequence is equal to the maximum length, thereby retaining the key information in the sequence and ensuring a unified sequence length when inputting into the classification model.

[0078] Convert the functional labels of each peptide into numerical forms to form a multi-label binary vector. Each component of the binary vector corresponds to a functional type, with values of 0 or 1. If the value is 0, it means the peptide does not have this function; if the value is 1, it means the peptide has this function. For example, the label vector of a peptide with antibacterial function is 1 in the antibacterial dimension and 0 for the remaining functions. In this way, standardize the label data into numerical forms to provide a multi-functional parsable label format for the classification model, which helps the training and prediction of the multi-label classification model.

[0079] In this embodiment, the biological features include amino acid index features, pseudo amino acid composition features, amino acid physicochemical property features, amino acid pair substitution score features, and amino acid composition features.

[0080] By introducing the above five biological feature encodings into the amino acid sequence, the numerical representation of the sequence not only contains character information but also rich biological meanings: Amino acid index (AAI) is a set of indicators describing the physicochemical properties of amino acids, covering various physicochemical properties such as hydrophobicity, polarity, and molecular volume, with a total of 531 indicators. Different amino acids have their respective numerical representations under each indicator. In this embodiment, each amino acid in each amino acid sequence is converted into its corresponding AAI value to form a matrix reflecting the overall physicochemical characteristics of the sequence, thereby enhancing the recognition ability of the classification model for the physicochemical properties of amino acid sequences.

[0081] Pseudo amino acid composition (PAAC) encoding combines amino acid composition and sequence information and can reflect the sequence correlation between amino acids. For each amino acid sequence, in this embodiment, the frequencies of amino acids in the sequence and the correlation parameters between adjacent amino acids are calculated to capture the global information and sequence dependence of the sequence, so that the classification model can more comprehensively learn the functional characteristics of peptides.

[0082] Physicochemical property encoding (PC6) includes six main physicochemical properties: hydrophobicity, molecular volume, polarity, isoelectric point, acid dissociation constant, and non-covalent interaction. For each amino acid sequence, in this embodiment, the six physicochemical property characteristics of each amino acid are summarized, and the statistical values of the sequence on these six main properties are calculated to capture the key physicochemical information of the sequence, enabling the classification model to identify the functional characteristics of peptides.

[0083] Amino acid pair substitution score matrix (BLOSUM62) is used to measure the substitution score between amino acids, reflecting the evolutionary relationship and functional similarity of amino acids. For each amino acid sequence, in this embodiment, the substitution score for amino acid pairs is calculated according to the BLOSUM62 matrix to obtain the evolutionary information of the amino acid sequence, enabling the classification model to learn the functional association of amino acid sequences.

[0084] Amino acid composition (AAC) encoding forms a 20-dimensional feature vector by counting the frequencies of each type of amino acid in the sequence. For each amino acid sequence, in this embodiment, the occurrence times of each amino acid in the sequence are calculated and normalized to relative frequencies to obtain the global information of amino acid composition, which helps the classification model to identify the functional patterns of different amino acid sequences.

[0085] For each biometric encoding, sequentially read the feature data from each feature file and remove the blank lines. Extract the amino acid identifiers (single-letter codes, such as A, C, D, etc.) in the first column and store them as the amino acid attribute list cha. Subsequently, define the index matrix to store the numerical features of each amino acid under different attributes. When reading each line of data, remove the null values and convert the attribute columns to floating-point numbers, and fill them into the index matrix.

[0086] Convert the index matrix into a Numpy array to generate a feature dictionary, where the keys are the single-letter codes of the amino acids and the values are the feature vectors of the amino acids in the feature matrix. To handle unknown or non-standard amino acid symbols (such as X), configure a vector of all zeros for the unknown amino acids in the dictionary to ensure the consistency of the input to the classification model.

[0087] Encode each amino acid sequence using the above feature dictionary: traverse each amino acid character and convert it into the corresponding feature vector. If the character does not appear in the dictionary, replace it with a vector of all zeros. Combine all the feature vectors into a feature matrix of a fixed length for the classification model to process.

[0088] To ensure that the output sequence lengths are consistent, uniformly pad or truncate all the encoded sequences to the maximum length max_len (defined as 50 in this embodiment). If the sequence length is less than 50, pad with zero vectors at the end. If the length exceeds 50, retain the first 50 feature vectors. The shape of each processed sequence matrix is (50, number of features).

[0089] Convert all the processed sequence embedding data into a Numpy array and further convert it into a PyTorch tensor with the data type set to floating-point. This design facilitates directly inputting the biometric information into the classification model to help the classification model better understand the biological information of the amino acid sequences during the training process.

[0090] In this embodiment, as Figure 2 shown, the classification model includes an embedding layer, a feature mapping layer, a positional encoding layer, a feature extraction layer, a fuzzification processing layer, a feature fusion module, a bidirectional long short-term memory network, a text convolutional network layer, a feature aggregation network encoding layer, a fully connected layer, and an output layer.

[0091] The embedding layer is used to map the numericalized amino acid sequences into a high-dimensional vector space to obtain the first feature representation.

[0092] In the specific implementation process, the embedding layer maps the numerical representation of the amino acid sequence (usually an integer-encoded amino acid sequence) to a high-dimensional vector space, capturing the potential features and sequence patterns of the amino acids. Through this process, the amino acid sequence is transformed from its original numerical form into a continuous space representation that can be processed by a deep learning classification model. The output of the embedding layer is a three-dimensional tensor with dimensions [number of batch sequences, maximum sequence length, number of embedding features], where the number of embedding features is the set embedding dimension.

[0093] The feature mapping layer is used to perform feature mapping on the biological features of the amino acid sequence to obtain a second feature representation with the same dimension as the first feature representation.

[0094] In the specific implementation process, the feature mapping layer maps biological features of different dimensions to a unified vector space through a linear transformation (a fully connected layer is used in this embodiment). This process fuses various biological features of the amino acid sequence (such as AAI, PAAC, etc.) with the embedding features and converts them into representations of the same dimension, facilitating subsequent feature fusion.

[0095] It can be understood that in this embodiment, by introducing multiple additional feature mapping layers, different types of biological features are mapped into the embedding space. These features include amino acid index (AAI), pseudo amino acid composition (PAAC), amino acid pair substitution score matrix (BLOSUM62), amino acid composition (AAC), and based on six predefined physicochemical properties (PC6), etc. These additional features provide rich domain knowledge and context information, enabling the classification model to better understand the biological significance in the amino acid sequence, helping the classification model capture more information during feature extraction, and thus improving the classification effect and the generalization ability of the classification model.

[0096] The position encoding layer is used to generate position encoding and add the position encoding to the first feature representation to obtain a third feature representation containing position information.

[0097] In the specific implementation process, the position encoding introduces the position information of the amino acids in the amino acid sequence to help the classification model understand the sequential relationship of the amino acids in the sequence. Since the self-attention mechanism lacks an inherent understanding of position, position encoding is needed to supplement this information. In this embodiment, sine and cosine functions are used to generate position encoding vectors, which are added to the embedding vectors to obtain a sequence representation containing position information. For the pos th element in the sequence.

[0098] The calculation formula for the position encoding of even dimensions is:

[0099] The calculation formula for odd-dimensional position encoding is as follows:

[0100] where represents the position of the element in the sequence (position index, starting from 0), is the dimension index of the embedding vector, is the total dimension of the embedding vector, i.e., num_hiddens.

[0101] The feature extraction layer is used to perform feature learning on the third feature representation in parallel using multiple attention heads, generating a fourth feature representation for each position under different attention heads.

[0102] In the specific implementation process, the position-encoded embedding vector serves as the input to the multi-head self-attention mechanism. The multi-head self-attention mechanism processes the input sequence through multiple parallel "heads", and each head learns the dependencies between different positions in the sequence. This enables the classification model to capture long-range dependencies in the sequence, focus on key amino acids and sequence segments, and thereby enhance the classification model's understanding of the sequence function. Each self-attention head calculates the attention weights for different positions in the sequence to weight the input sequence, thereby obtaining a weighted output feature. This mechanism enables the classification model to focus on the amino acids and sequence patterns that are most critical for functional classification when performing classification tasks.

[0103] The fuzzification processing layer is used to perform fuzzification processing on the second feature representation to obtain a fifth feature representation.

[0104] In the specific implementation process, this embodiment introduces the fuzzification layer of the fuzzy neural network in the feature processing link to enhance the classification model's ability to process uncertainty and fuzzy information. The fuzzification layer uses the Gaussian membership function to convert the numerical value of the input feature into a membership degree, and this process can effectively simulate the uncertainty in biological features and improve the expression of features.

[0105]

[0106] In the formula, is the th Gaussian membership degree, is the output after linear transformation, is the center of the th Gaussian function, is the th width of the Gaussian function.

[0107] It can be understood that, in order to enhance the processing ability of the classification model for complex features, the fuzzification layer of the fuzzy neural network is introduced in this embodiment, and a Gaussian membership function is used to perform fuzzification processing on various external features. The fuzzification processing converts the feature values into membership degrees, enhancing the flexibility and robustness of the feature representation, enabling the classification model to better process uncertain and continuous data features. At the same time, through the fuzzification processing, the training time of the classification model is shortened by half, significantly improving the training efficiency and stability, and further optimizing the quality of feature expression.

[0108] The feature fusion module is used to perform fusion processing on the first feature representation, the fourth feature representation, and the fifth feature representation to obtain a sixth feature representation containing comprehensive features.

[0109] In the specific implementation process, the first feature representation, the fourth feature representation, and the fifth feature representation are fused. Through weighted or concatenated methods, these features from different sources are combined into a unified feature vector. This process can integrate various types of feature information, providing a more comprehensive feature input for the subsequent network, thereby improving the classification performance.

[0110] The bidirectional long short-term memory network is used to perform bidirectional feature extraction on the sixth feature representation to obtain a seventh feature representation containing forward and backward information.

[0111] In the specific implementation process, in order to capture long-distance dependency relationships in the sequence, the bidirectional long short-term memory network (BiLSTM) is introduced in this embodiment. BiLSTM can simultaneously process the front and back information of the sequence through two LSTM layers in the forward and reverse directions, and model the global features in the sequence. The features processed by BiLSTM are more abundant, which can enhance the classification model's understanding of the global information of the amino acid sequence. The input dimension of BiLSTM is 256, and the hidden layer dimension is 128.

[0112] It can be understood that after feature fusion, the bidirectional long short-term memory network (BiLSTM) is introduced in this embodiment, which can effectively capture the dependencies between the context before and after in the sequence. BiLSTM enhances the understanding of the global sequence information by considering both the forward and reverse information of the sequence at the same time. This is particularly important in tasks that require understanding context relationships, which can improve the classification model's ability to capture long-distance dependencies and enhance the accuracy of the classification model in amino acid sequence classification.

[0113] The text convolutional network is used to perform a convolutional operation on the seventh feature representation to extract local features of the sequence and obtain an eighth feature representation.

[0114] In the specific implementation process, the TextCNN layer is used to extract local features of the amino acid sequence. This layer contains three one-dimensional convolutional layers with kernel sizes of 4, 6, and 8, which are used to extract local features at different scales. After each convolutional layer, the ReLU activation function is applied to increase the nonlinear expression ability of the classification model. After the convolutional layer, the maximum pooling operation is performed to reduce the feature dimension and retain the most important information. Finally, the features extracted by the three convolutional layers are spliced ​​together to form a vector containing multi-level local features.

[0115] It is understandable that this embodiment organically combines fuzzy neural networks, bidirectional long short-term memory networks (BiLSTM) and text convolutional networks (TextCNN), so that the classification model can take advantage of the advantages of three different networks to obtain more comprehensive information from variable-length amino acid sequences. Fuzzy neural networks enhance the classification model's ability to process fuzzy data, BiLSTM improves the classification model's ability to understand global dependencies, and TextCNN can extract local features of sequences. The combination of multiple networks not only improves the stability of the classification model, but also improves the accuracy of classification.

[0116] The feature aggregation network coding layer is used to perform aggregation coding on the eighth feature representation, integrate multi-level feature information, and obtain a ninth feature representation.

[0117] In the specific implementation process, the feature aggregation network (FAN) encoding layer further improves the quality of feature representation and the classification ability of the classification model by aggregating the multi-level features processed by BiLSTM and CNN. The FAN encoding layer can effectively integrate feature information from different levels of classification models, enhance the representation ability of the classification model, and improve its recognition accuracy for peptide function classification tasks.

[0118] The fully connected layer is used to perform layer-by-layer dimensionality reduction processing on the ninth feature representation, extract high-order features, and obtain a tenth feature representation.

[0119] In the specific implementation process, in the fully connected layer, the features are processed layer by layer through multiple fully connected networks. Each layer reduces the feature dimension and gradually extracts high-order features. The ReLU activation function is applied after each fully connected layer to enhance the nonlinear expression ability of the network, so that the classification model can learn more complex feature representations.

[0120] The output layer is used to classify and output the tenth feature representation, map the classification result to a range of 0 to 1, and obtain a prediction value consistent with the number of functional categories.

[0121] In the specific implementation process, in the output layer, the last fully connected layer outputs nodes with the same number as the number of functional categories, representing the prediction results of each peptide function. Then, these prediction results are mapped to between 0 and 1 through the Sigmoid function to obtain the probability values of each functional category. Since the number of categories set in this embodiment is 21, the output layer contains 21 nodes, corresponding to the 21 functions that the amino acid sequence may possess respectively.

[0122] It can be understood that in this embodiment, the word embedding, attention encoding, and various additional feature representations are comprehensively fused through the feature fusion module. The feature fusion module fuses features from different sources by introducing learnable weight parameters and linear transformation to ensure the maximum retention of information. In this way, the classification model can make full use of information from different sources, effectively improving the overall feature expression ability and comprehensive performance, and showing better effects when processing diverse data.

[0123] By introducing an additional feature embedding layer, a fuzzification processing layer, a feature fusion layer, and a bidirectional long short-term memory network (BiLSTM) layer, this embodiment makes the classification model structure more modular and flexible. The modular design facilitates future expansion and adjustment in different tasks, improving the adaptability and maintainability of the classification model. For example, new feature processing modules can be easily added or existing modules can be replaced to adapt to different data types and task requirements, and this design significantly improves the scalability and generality of the classification model.

[0124] In this embodiment, through the bidirectional long short-term memory network (BiLSTM) and the feature fusion layer, it is possible to flexibly process input data of different lengths and different dimensions. For example, BiLSTM can process variable-length sequence data, while the feature fusion layer can integrate features of different dimensions. Such a design not only improves the adaptability and generality of the classification model, enabling it to be applicable to a wider range of data types and task scenarios, but also enhances the ability of the classification model to process inputs of different lengths, making it suitable for complex scenarios in practical applications.

[0125] In this embodiment, based on the optimizer and the loss function, the classification model is trained using the standardized amino acid sequence data, biological features, and the generative adversarial network, including: Using the standardized amino acid sequence data and biological features as real data to train the generative adversarial network to obtain a trained generative adversarial network.

[0126] Using the trained generative adversarial network to generate amino acid sequence data as fake data.

[0127] Performing a merging process on the real data and the fake data to obtain an enhanced training set.

[0128] Input the enhanced training set into the classification model, and calculate the classification loss based on the loss function.

[0129] Perform backpropagation according to the classification loss, update the parameters of the classification model, and dynamically adjust the learning rate during the training process of the classification model based on the optimizer until the number of training rounds is reached to obtain a trained classification model.

[0130] In this embodiment, as Figure 3 shown, the generative adversarial network includes a generator and a discriminator. Using the standardized amino acid sequence data and biological features as real data to train the generative adversarial network includes: Train the discriminator: Convert the standardized amino acid sequence data and biological features as real data into one-hot representation and input them into the discriminator. Calculate the discriminant loss d_real_loss of the real data by calculating the output probability real_validity of the real data.

[0131] Sample random noise z from the standard normal distribution. Synthesize amino acid sequence data as fake data gen_data through the generator based on the random noise z. Input the fake data gen_data into the discriminator, and calculate the discriminant loss d_fake_loss of the fake data by calculating the output probability fake_validity of the fake data.

[0132] Calculate the total discriminator loss (d_real_loss + d_fake_loss) / 2, and update the parameters of the discriminator through backpropagation according to the total discriminator loss (d_real_loss + d_fake_loss) / 2.

[0133] Train the generator: Sample random noise z from the standard normal distribution. Synthesize amino acid sequence data as fake data gen_data through the generator based on the random noise z. Input the fake data gen_data into the discriminator, and calculate the output probability fake_validity of the fake data. Since the goal of the generator is to make the discriminator think that the generated data is real, the loss of the discriminator at this time is adversarial_loss(validity, valid). adversarial_loss(validity, valid) is a method of calling the loss function. The loss function used in this solution to calculate the loss of the generative adversarial network is the binary cross-entropy loss function. Update the generator parameters through backpropagation according to the loss adversarial_loss(validity, valid).

[0134] It is understandable that in this embodiment, a generative adversarial network (GAN) is integrated into the classification model to further improve the performance and robustness of the functional classification of multifunctional therapeutic peptides. Through the generation and adversarial mechanisms of GAN, the generalization ability and classification accuracy of the classification model are enhanced. GAN can generate fake data similar to the real data distribution, thereby alleviating the problems of insufficient data and class imbalance, and at the same time optimizing the feature representation of the classification model through the adversarial training mechanism. This innovative combination improves the application value and scalability of the classification model in fields such as drug design and bioinformatics research, making the classification model show significant advantages in the classification task of multifunctional therapeutic peptides.

[0135] In this embodiment, the optimizer uses the Adam optimizer combined with a cosine learning rate scheduler, and dynamically adjusts the learning rate in stages through cosine annealing and warm-up annealing according to the current training progress. Gradually reducing the learning rate to improve the convergence effect of the classification model. The cosine annealing strategy adjusts the learning rate periodically according to the progress of the training rounds. The warm-up strategy gradually increases the learning rate at the beginning of training to stabilize the training.

[0136] Among them, the expression of the cosine learning rate scheduler is as follows:

[0137] In the formula, represents the warm-up annealing stage, represents the cosine annealing stage, represents the learning rate at the start of warm-up annealing, represents the current step number, represents the number of steps in the warm-up annealing stage, represents the initial learning rate, represents the final learning rate, represents the number of steps for cosine annealing, represents the total number of update steps.

[0138] Such as Figure 4As shown, the classification model in this embodiment adopts a strategy that combines learning rate warm-up and cosine annealing. In the initial stage, the learning rate gradually increases to the peak value (about 0.0018), which is called learning rate warm-up and helps the model to converge quickly during the initial training. After reaching the maximum learning rate, the model starts to gradually decrease the learning rate, showing a trend of cosine annealing until it approaches zero. This learning rate scheduling strategy effectively balances the rapid convergence in the initial stage of training and the stable optimization in the later stage. The relatively high learning rate in the initial stage accelerates the model's exploration of the global optimal solution, while the gradually decreasing learning rate in the later stage helps the model to make fine adjustments in the local area, reduces oscillations during the training process, and thus improves the stability and generalization ability of the model. Finally, this strategy significantly improves the performance of the model and reduces the risk of overfitting, enabling the model to show better results in the classification task of multifunctional therapeutic peptides.

[0139] In this embodiment, to solve the problems of class imbalance and difficult-to-classify samples, this solution adopts a custom MarginalFocalDiceLoss function. This loss function combines the margin loss (LDAM) and the focal dice loss (FocalDiceLoss). This design aims to enhance the classification model's ability to identify difficult-to-classify samples, especially when dealing with class-imbalanced data, and improve the generalization ability and classification performance of the classification model.

[0140] To assign different importance to different classes in the case of class imbalance, the calculation of class weights is crucial. The specific calculation method is as follows:

[0141] where, represents the weight of the majority-class label samples, represents the weight of the minority-class label samples, represents the number of majority-class label samples, represents the number of minority-class label samples, represents the total number of samples (after majority-class and minority-class samples).

[0142] The margin loss (LDAM, Large Margin Adaptive Loss) is a loss function used to adjust the class imbalance problem. It penalizes difficult-to-classify samples by introducing margin adjustments. The purpose of LDAM is to make the classification model more sensitive to these samples by increasing the loss of difficult-to-classify samples. The loss calculation formula for margin adjustment is:

[0143] In the formula, represents the predicted value of the classification model, Represents the marginal adjustment value for minority class samples, Represents the true label value, Represents the marginal adjustment value for majority class samples.

[0144] FocalDiceLoss combines Focal Loss and Dice coefficient to better handle class imbalance and difficult-to-classify samples. The calculation of Focal Dice Loss consists of two parts: the positive class sample part and the negative class part.

[0145] For positive class samples (usually representing the target class), its loss is calculated using Focal Loss, and its expression is as follows:

[0146] In the formula, Represents the predicted value of the positive class sample after clipping, Represents the true label (taking values of 1 or 0), Represents the original predicted value of the positive class sample, Represents the clipping parameter of the positive class sample.

[0147] For negative classes, the calculation of Focal Loss is similar, but usually the classes are weighted to avoid the situation where too many negative class samples lead to ineffective training of the classification model. Its expression is as follows:

[0148] Where, Represents the predicted value of the negative class sample after clipping, Represents the original predicted value of the negative class sample, Represents the clipping parameter of the negative class sample.

[0149] Finally, the total loss of FocalDiceLoss is the weighted sum of the losses of the positive and negative class parts, and its calculation formula is:

[0150] In the formula, Is the weight coefficient of the positive class sample.

[0151] Combining LDAM and FocalDiceLoss, the resulting weighted loss function can be expressed as:

[0152] Finally, the expression of the loss function of the classification model can be expressed as:

[0153] It is understandable that this embodiment proposes a new composite loss function that combines the advantages of the Focal Loss, Dice Loss, and LDAM (Label-Distribution-Aware Margin Loss). This loss function is specifically designed to handle the problem of class imbalance and is applicable to the functional classification task of multifunctional therapeutic peptides. This composite loss function can help the classification model focus on difficult-to-classify samples, improve the classification accuracy, and also show good performance and stability in the segmentation task. This loss function is not only applicable to the classification task of multifunctional therapeutic peptides but can also be extended to other sequence or image classification tasks with class imbalance.

[0154] Where, N is the total number of samples, represents the weight coefficient of the positive class samples, represents the predicted value of the positive class samples after clipping, represents the true label (taking values of 1 or 0), represents the predicted value of the negative class after clipping, represents the weight adjustment parameter based on the class.

[0155] In this embodiment, after obtaining the trained classification model, the method further includes evaluating the classification model using evaluation metrics.

[0156] The evaluation metrics include Aiming, Coverage, Accuracy, Aboslute_True, and Aboslute_False, and the expressions of each evaluation metric are as follows:

[0157] In the formula, is used to measure the correct proportion in the classification labels, and n is the number of samples. represents the number of true positives of the th sample, that is, the number of labels that the classification model predicts as positive and are actually positive. represents the number of false positives of the th sample, that is, the number of labels that the classification model predicts as positive and are actually negative.

[0158]

[0159] In the formula, is used to measure the proportion of the labels correctly classified by the classification model among all the true labels, represents the number of false negatives of the th sample, that is, the number of labels that the classification model predicts as negative and are actually positive.

[0160]

[0161] In the formula, is used to measure the proportion of correctly classified labels among all relevant labels, including correctly classified, misclassified, and unclassified true labels.

[0162]

[0163] In the formula, Aboslute_True is used to measure the proportion of the predicted labels of the classification model that are exactly the same as the true labels. represents the predicted label vector of the th sample, is an indicator function used to check whether the predicted label and the true label exactly match. If they are exactly the same, it returns 1; otherwise, it returns 0.

[0164]

[0165] In the formula, is used to measure the proportion of misclassified labels in the multi-label classification task of the classification model. represents the number of labels for each sample. represents the h th predicted label vector of the th h label in the

[0166] Table 1 Performance evaluation results of different classification models on the test set

[0167] As can be seen from Table 1, the classification model of the present invention is superior to other comparative classification models in all indicators, especially showing obvious advantages in the three main indicators of Aiming, Coverage, and Accuracy. Specifically, the Aiming value of the classification model of the present invention is 0.719, the Coverage value is 0.711, and the Accuracy value is 0.673, all of which are significantly higher than those of other classification models. This indicates that the classification model of the present invention has stronger accuracy and comprehensiveness in the classification task of multifunctional therapeutic peptides, can better cover the target categories and reduce misclassification. In addition, in the two indicators of Absolute True and AbsoluteFalse, the classification model of the present invention also shows a high correct rate and a low error rate, further verifying the generalization ability and robustness of the classification model.

[0168] Overall, the excellent performance of the classification model in this solution on the test set indicates that it has significant advantages in dealing with the functional classification task of multifunctional therapeutic peptides. These results show that this solution can not only improve the classification accuracy but also effectively address the problem of class imbalance, with high practical application value and promotion potential.

[0169] To facilitate the reproduction, comparison, and subsequent analysis of the classification model, this solution sets up functions for saving the classification model, recording results, and visualizing the training process. The parameters of the trained classification model are saved to a specified path for subsequent loading and use. When saving the classification model, it includes the complete network structure and parameter information to ensure seamless reproduction of the training status of the classification model in different environments.

[0170] Information such as the evaluation metrics and training time of the classification model is recorded in a CSV file to establish a detailed training log. This log record can be conveniently used for result reproduction, classification model performance comparison, and subsequent parameter tuning and improvement.

[0171] Draw the learning rate change curve during the training process of the classification model and the curve graph of the evaluation metrics changing with the training steps. Through these visualization charts, the training process and performance change trend of the classification model can be intuitively observed, facilitating the diagnosis and optimization of the training process of the classification model.

[0172] Example 3 As Figure 2 shown, this example proposes a machine learning-based amino acid sequence data classification system applied to the machine learning-based amino acid sequence data classification method as described in the above example, including: an acquisition module 100, a preprocessing module 200, a construction module 300, a training module 400, and a classification and testing module 500.

[0173] Among them, the acquisition module 100 is used to acquire the amino acid sequence data of multifunctional therapeutic peptides with functional labels and divide the amino acid sequence data into a training set and a test set. The preprocessing module 200 is used to sequentially perform standardization and biometric encoding processing on the amino acid sequence data in the training set to obtain standardized amino acid sequence data and biometric features. The construction module 300 is used to construct a classification model for amino acid sequence data classification, as well as a generative adversarial network, an optimizer, and a loss function for training and optimizing the classification model. The training module 400 is used to train the classification model based on the optimizer and loss function, using the standardized amino acid sequence data, biometric features, and generative adversarial network until the number of training epochs is reached to obtain a trained classification model. The classification and testing module 500 is used to input the amino acid sequence data of the test set into the trained classification model for classification processing to obtain the classification result of the amino acid sequence data.

[0174] It should be noted that the foregoing explanation of the embodiments of the amino acid sequence data classification method based on machine learning is also applicable to the amino acid sequence data classification system based on machine learning in this embodiment, and will not be elaborated here.

[0175] In the description of this specification, the descriptions with reference to the terms "one embodiment", "some embodiments", "example", "specific example", or "some examples", etc. mean that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present application. In this specification, the schematic expressions of the above terms are not necessarily directed to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described can be combined in any one or N embodiments or examples in a suitable manner. In addition, without conflict, those skilled in the art can combine and combine the different embodiments or examples described in this specification and the features of different embodiments or examples.

[0176] In addition, the terms "first" and "second" are only used for descriptive purposes and cannot be understood as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include at least one of the features. In the description of the present application, the meaning of "N" is at least two, such as two, three, etc., unless otherwise specifically defined.

[0177] Any process or method description in the flowchart or described in other ways herein can be understood to represent a module, segment, or part of code including one or more executable instructions for implementing a customized logical function or process. The scope of the preferred embodiments of the present application includes additional implementations, where the functions can be executed in a substantially simultaneous manner or in a reverse order according to the functions involved, rather than in the order shown or discussed, which should be understood by those skilled in the art of the embodiments of the present application.

[0178] It should be understood that each part of the present application can be implemented by hardware, software, firmware, or a combination thereof. In the above embodiments, N steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented by hardware, as in another embodiment, any one of the following well-known technologies in the art or a combination of them can be used: discrete logic circuits with logic gate circuits for implementing logical functions on data signals, application-specific integrated circuits with appropriate combinational logic gate circuits, programmable gate arrays, field programmable gate arrays, etc.

[0179] Those of ordinary skill in the art can understand that all or part of the steps carried out in the methods of the above embodiments can be completed by instructing relevant hardware through a program. The program can be stored in a computer-readable storage medium. When the program is executed, it includes one or a combination of the steps of the method embodiments.

[0180] Obviously, the above embodiments of the present invention are merely examples for clearly illustrating the present invention, rather than limitations on the implementation manners of the present invention. For those of ordinary skill in the art, other different forms of changes or modifications can be made based on the above description. It is not necessary and impossible to enumerate all the implementation manners here. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the claims of the present invention.

Claims

1. A method for classifying amino acid sequence data based on machine learning, characterized in that: include: Acquiring amino acid sequence data of a multifunctional therapeutic peptide having a functional tag, and dividing the amino acid sequence data into a training set and a test set; The amino acid sequence data in the training set are sequentially standardized and biometrically encoded to obtain standardized amino acid sequence data and biometrics; Constructing a classification model for amino acid sequence data classification, and a generative adversarial network, an optimizer, and a loss function for training and optimizing the classification model; Based on the optimizer and the loss function, the classification model is trained using standardized amino acid sequence data, biological features, and a generative adversarial network until the number of training rounds is reached to obtain a trained classification model; The amino acid sequence data of the test set is input into the trained classification model for classification processing to obtain the classification result of the amino acid sequence data.

2. The method for classifying amino acid sequence data based on machine learning according to claim 1, characterized in that: Standardization of amino acid sequence data, including: Using a standard amino acid character set containing 20 natural amino acids and 1 unknown amino acid, the characters of each amino acid sequence data are mapped to a unique integer index to generate a numerical amino acid sequence; Set the maximum sequence length. For numerical amino acid sequences that are shorter than the maximum sequence length, zero padding is performed at the end. For numerical amino acid sequences that exceed the maximum sequence length, the front sequence is truncated to the maximum sequence length. The functional label of each amino acid sequence is converted into a binary vector, wherein each component of the binary vector represents a function, and if the amino acid sequence has the function, the corresponding component takes the value of 1, otherwise it takes the value of 0.

3. The method for classifying amino acid sequence data based on machine learning according to claim 2, characterized in that: The classification model includes an embedding layer, a feature mapping layer, a position encoding layer, a feature extraction layer, a fuzzification processing layer, a feature fusion module, a bidirectional long short-term memory network, a text convolutional network layer, a feature aggregation network encoding layer, a fully connected layer and an output layer; The embedding layer is used to map the digitized amino acid sequence to a high-dimensional vector space to obtain a first feature representation; The feature mapping layer is used to perform feature mapping on the biological features of the amino acid sequence to obtain a second feature representation with the same dimension as the first feature representation; The position encoding layer is used to generate a position code, and add the position code to the first feature representation to obtain a third feature representation containing position information; The feature extraction layer is used to perform feature learning on the third feature representation in parallel using multiple attention heads to generate a fourth feature representation of each position under different attention heads; The fuzzification processing layer is used to perform fuzzification processing on the second feature representation to obtain a fifth feature representation; The feature fusion module is used to fuse the first feature representation, the fourth feature representation and the fifth feature representation to obtain a sixth feature representation containing comprehensive features; The bidirectional long short-term memory network is used to perform bidirectional feature extraction on the sixth feature representation to obtain a seventh feature representation containing forward and backward information; The text convolution network is used to perform a convolution operation on the seventh feature representation to extract local features of the sequence to obtain an eighth feature representation; The feature aggregation network coding layer is used to aggregate and encode the eighth feature representation, integrate multi-level feature information, and obtain a ninth feature representation; The fully connected layer is used to perform layer-by-layer dimensionality reduction processing on the ninth feature representation, extract high-order features, and obtain a tenth feature representation; The output layer is used to classify and output the tenth feature representation, map the classification result to a range of 0 to 1, and obtain a prediction value consistent with the number of functional categories.

4. The method for classifying amino acid sequence data based on machine learning according to claim 2, characterized in that: Based on the optimizer and the loss function, the classification model is trained using standardized amino acid sequence data, biological features and a generative adversarial network, including: Using standardized amino acid sequence data and biological features as real data to train the generative adversarial network to obtain a trained generative adversarial network; Use the trained generative adversarial network to generate amino acid sequence data as fake data; The real data and the fake data are combined to obtain an enhanced training set; Inputting the enhanced training set into the classification model, and calculating the classification loss based on the loss function; Back propagation is performed according to the classification loss to update the parameters of the classification model. Based on the optimizer, the learning rate is dynamically adjusted during the training process of the classification model until the number of training rounds is reached to obtain a trained classification model.

5. The method for classifying amino acid sequence data based on machine learning according to claim 4, characterized in that: The generative adversarial network includes a generator and a discriminator, and the generative adversarial network is trained using standardized amino acid sequence data and biological features as real data, including: The standardized amino acid sequence data and biological features are converted into one-hot representation as real data and then input into the discriminator. The discriminant loss of the real data is calculated by calculating the output probability of the real data. Sampling random noise from a standard normal distribution, synthesizing amino acid sequence data as false data through a generator based on the random noise, inputting the false data into a discriminator, and calculating the discrimination loss of the false data by calculating the output probability of the false data; According to the discriminative loss of real data and the discriminative loss of fake data, the parameters of the discriminator and generator are updated through back propagation.

6. The method for classifying amino acid sequence data based on machine learning according to claim 4, characterized in that: The optimizer uses the Adam optimizer combined with the cosine learning rate scheduler to dynamically adjust the learning rate in stages through cosine annealing and warm-up annealing; Among them, the expression of the cosine learning rate scheduler is as follows: In the formula, represents the preheating annealing stage, represents the cosine annealing stage, represents the learning rate at the beginning of preheat annealing, Indicates the current number of steps. represents the number of steps in the preheating annealing stage, represents the initial learning rate, represents the final learning rate, represents the number of steps used for cosine annealing, Indicates the total number of update steps.

7. The method for classifying amino acid sequence data based on machine learning according to claim 1, characterized in that: The expression of the loss function of the classification model is as follows: in, N is the total number of samples, represents the weight coefficient of the positive sample, represents the predicted value of the positive sample after clipping, represents the true label, represents the negative class prediction value after clipping, Represents the class-based weight adjustment parameter.

8. The method for classifying amino acid sequence data based on machine learning according to claim 1, characterized in that: The biological feature includes one or more of an amino acid index feature, a pseudo amino acid composition feature, an amino acid physicochemical property feature, an amino acid pair substitution score feature or an amino acid composition feature.

9. The method for classifying amino acid sequence data based on machine learning according to claim 1, characterized in that: After obtaining the trained classification model, the method further includes evaluating the classification model using an evaluation index; The evaluation indicators include Aiming, Coverage, Accuracy, Aboslute_True and Aboslute_False. The expressions of each evaluation indicator are as follows: In the formula, Used to measure the correct proportion of classification labels, n is the number of samples; Indicates The true number of samples, that is, the number of labels predicted by the classification model to be positive and actually positive; Indicates The number of false positives for a sample, that is, the number of labels that the classification model predicts to be positive but is actually negative; In the formula, It is used to measure the proportion of labels correctly classified by the classification model among all true labels. Indicates The number of false negatives for a sample, that is, the number of labels that the classification model predicts to be negative but is actually positive; In the formula, It is used to measure the proportion of correctly classified labels among all relevant labels, including correctly classified, misclassified, and unclassified true labels; In the formula, Aboslute_True is used to measure the proportion of the predicted labels of the classification model that are completely consistent with the true labels. Indicates The predicted label vector of samples, Indicates The true label vector of samples, is an indicator function used to check whether the predicted label and the true label are completely matched; In the formula, Used to measure the proportion of misclassification of the classification model in multi-label classification tasks. represents the number of labels for each sample, Indicates In the sample h The predicted label vector of labels, Indicates In the sample h The true label vector of each label.

10. An amino acid sequence data classification system based on machine learning, characterized in that: include: An acquisition module, used to acquire amino acid sequence data of a multifunctional therapeutic peptide having a functional tag, and divide the amino acid sequence data into a training set and a test set; A preprocessing module is used to perform standardization and biometric encoding processing on the amino acid sequence data in the training set in turn to obtain standardized amino acid sequence data and biometrics; A construction module for constructing a classification model for amino acid sequence data classification, and a generative adversarial network, an optimizer, and a loss function for training and optimizing the classification model; A training module, for training the classification model based on the optimizer and the loss function using standardized amino acid sequence data, biological features and a generative adversarial network until the number of training rounds is reached to obtain a trained classification model; The classification test module is used to input the amino acid sequence data of the test set into the trained classification model for classification processing to obtain the classification results of the amino acid sequence data.

Citation Information

Cited By

  • Information source coding matrix data processing method based on modified conjugate gradient algorithm

    CN121690229A