A method for nucleic acid sequence feature mining based on autoregressive large model
Through the nucleic acid sequence feature mining method based on the autoregressive large model, the problem of insufficient feature extraction of large-scale nucleic acid sequence data sets is solved, efficient and in-depth feature extraction is achieved, and support for downstream nucleic acid analysis tasks.
Patent Information
- Application Number
- CN202410877180.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-07-02
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2044-07-02
AI Technical Summary
When processing large-scale nucleic acid sequence data sets, the feature extraction is insufficient, the model generalization ability is weak, and it is difficult to mine potential sequence patterns from large-scale data sets.
The nucleic acid sequence feature mining method based on the autoregressive large model is adopted. By obtaining large-scale nucleic acid sequence data sets for pre-processing, k-mer fragments are summarized based on the frequency statistics method, the nucleic acid sequence is encoded, and an autoregressive converter model is constructed for unsupervised training, and the processed nucleic acid sequence features are output.
This method can generate nucleic acid sequence embedding features containing rich semantic information from a large-scale nucleic acid sequence dataset, optimize the data processing process, better understand and integrate sequence features, and reduce dependence on professional knowledge.
Smart Images

Figure CN119049566B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the combination of biochemical science and computer science, relates to the field of artificial intelligence technology, and particularly to an unsupervised nucleic acid sequence feature mining method based on an autoregressive large model. Background Art
[0002] Traditional biological analysis methods study nucleic acid sequence characteristics through gene sequencing and comparative analysis, which has the defects of low efficiency and incomplete information coverage. Therefore, characterizing nucleic acid sequences and effectively extracting nucleic acid features is a crucial topic in the study of modern biology under informatization. The extracted high-dimensional nucleic acid sequence features help to diagnose diseases more accurately, predict genetic tendencies and understand complex biological processes. The current problems include limited data processing capabilities, insufficient feature extraction and weak model generalization capabilities. With the development of artificial intelligence machine learning and deep learning related fields, deep learning methods represented by statistics and neural networks are increasingly applied to the field of biology to mine potential patterns of nucleic acid sequences from a large number of sequence samples and extract high-dimensional features of nucleic acid sequences for downstream tasks based on machine learning to solve the problems of insufficient information mining from a large number of sample perspectives and insufficient data feature extraction in existing biological and chemical methods. The method based on deep neural network can use a large number of nucleic acid sequence features to find the universal paradigm of sequences across sample sets, combine the global information and local information of sequences from the perspective of sequences, and generate nucleic acid features with high semantic information, thereby overcoming the shortcomings and deficiencies of traditional experiments.
[0003] Machine learning-based methods and deep learning neural network-based methods are common methods for extracting features from nucleic acid sequence datasets. Machine learning-based methods usually rely on manually designed features and combine statistical learning models to analyze and interpret data, such as support vector machines (SVMs), random forests, and logistic regression. However, these methods usually require domain experts to perform feature engineering and have limited ability to process nonlinear complex patterns, and cannot obtain high-dimensional features that contain rich recognition patterns. Deep learning neural network-based methods use multi-layer neural networks to automatically learn feature representations from data without manual feature engineering, such as transformer-based bidirectional representation encoders (BERT) and graph neural network-based node feature update models. However, these models generally have problems such as insufficient capture of long-range features of nucleic acid long sequences and overfitting in training. In addition, previous models generally focus on specific types of nucleic acid sequence datasets or multi-type small-scale nucleic acid sequence datasets. There is insufficient feature mining for multi-type and large nucleic acid datasets, and it is difficult to mine potential sequence patterns from large-scale datasets, thereby promoting downstream task research. Summary of the invention
[0004] The purpose of the present invention is to provide a nucleic acid sequence feature mining method based on an autoregressive large model to accurately, efficiently and deeply extract features from large-scale nucleic acid sequence data sets, thereby creating conditions for downstream nucleic acid development.
[0005] In order to achieve the above object, the present invention provides the following technical solution: a method for mining nucleic acid sequence features based on an autoregressive large model, comprising:
[0006] Acquire a large-scale nucleic acid sequence data set, and preprocess the nucleic acid sequence data set;
[0007] Induction of k-mer fragments of nucleic acids based on frequency statistics; where k represents the length of the fragment;
[0008] Using the k-mer fragment to encode a nucleic acid sequence;
[0009] An autoregressive transformer model is constructed, and the encoded nucleic acid sequence is input for unsupervised training; the autoregressive transformer model includes a multi-layer coupled position information embedding network, a cross-domain feature extraction network, a monomer feature integration network and a feature output network; wherein the cross-domain feature extraction network is used to implement an adaptive attention mechanism and a multi-head attention mechanism;
[0010] Output the processed nucleic acid sequence features.
[0011] Furthermore, the acquisition of a large-scale nucleic acid sequence data set includes collecting data from multiple different sources, such as public bioinformatics databases (NCBI, Ensembl), clinical sample libraries of medical institutions and scientific research institutions.
[0012] Furthermore, the preprocessing includes removing low-quality sequences, standardizing and removing duplicate sequences;
[0013] The removal of low-quality sequences comprises: performing a quality score on the sequence according to the quality score of each base in the obtained nucleic acid sequence; setting a threshold and retaining sequences with a quality score greater than or equal to the threshold;
[0014] The standardization process is to arrange the sequences in a uniform direction;
[0015] The deduplication sequence comprises: calculating the hash value of each sequence in the data set, and deleting the sequence with the same hash value.
[0016] Furthermore, the method of summarizing k-mer fragments of nucleic acids based on frequency statistics includes the following sub-steps:
[0017] Scan the entire nucleic acid sequence data set and calculate the frequency of occurrence of all possible k-mer fragments in the sequence;
[0018] Setting a threshold to filter out k-mer fragments whose occurrence frequency is greater than or equal to the threshold;
[0019] The k-mer fragments obtained by screening are recursively merged to generate complete k-mer fragments.
[0020] Further, using the k-mer fragment to encode a nucleic acid sequence comprises the following sub-steps:
[0021] Map each recursively merged k-mer fragment to a unique integer identifier M(S p:p+k-1 ), generating a sequence of integers; where M(x) is the mapping function, S p:p+k-1 is the k-mer fragment corresponding to position p in the nucleic acid sequence S;
[0022] Two additional integer identifiers are introduced to indicate the start and end of the sequence, respectively, to generate the encoded sequence.
[0023] Furthermore, the location information is embedded in the network to perform the following sub-steps:
[0024] Based on the specific position of each k-mer fragment in the sequence, generate the corresponding position vector
[0025] The integer identifier M(S p:p+k-1 ) and the corresponding position vector Combine to generate the corresponding enhanced feature vector
[0026]
[0027] For each enhanced feature vector Perform weighted merging.
[0028] Furthermore, the position vector It is obtained by the following formula:
[0029]
[0030] Where j represents the dimension index in the position vector and d is the total number of dimensions in the position vector.
[0031] Furthermore, the following sub-steps are performed in the cross-domain feature extraction network:
[0032] Based on the adaptive attention mechanism, the importance of each element in the sequence for the overall sequence analysis is calculated and the attention weight α is assigned i :
[0033]
[0034] e i =f(x i )
[0035] Among them, e i is the element x i The energy score of , f is a learnable function;
[0036] Based on the multi-head attention mechanism, k independent attention modules are created to capture different features respectively; each module h uses the same adaptive attention mechanism to perform weighted fusion on each element according to the attention weight and output the corresponding vector o h ;
[0037] Afterwards, merge and output a comprehensive feature vector O:
[0038] O=Concat(o1,o2,…,o k )·W
[0039] Among them, Concat is the output o of the module h h concatenated into a long vector, W is a learnable weight matrix.
[0040] Furthermore, the following sub-steps are performed in the monomer feature integration network:
[0041] Using a fully connected layer, linearly transform the comprehensive feature vector O;
[0042] Through the activation function, the linearly transformed comprehensive feature vector is nonlinearly transformed.
[0043] Furthermore, the feature output network performs the following sub-steps:
[0044] The feature mapping layer is used to map the comprehensive feature vector output by the monomer feature integration network and convert it into a coding form corresponding to the original nucleic acid sequence;
[0045] A normalization layer is used to normalize the mapped comprehensive feature vector and output the processed nucleic acid sequence.
[0046] Furthermore, the unsupervised training comprises the following sub-steps:
[0047] Each time a sequence is received, the next k-mer in the sequence is predicted; the loss function between the predicted k-mer and the actual k-mer is compared, and the internal parameters are adjusted with the goal of minimizing the loss function.
[0048] The beneficial effects of the present invention are as follows:
[0049] (1) By using the autoregressive transformer model for unsupervised training, this method is able to generate nucleic acid sequence embedding features containing rich semantic information from large-scale nucleic acid sequence datasets.
[0050] (2) Through automated preprocessing and efficient encoding, the data processing pipeline is optimized, enabling it to cope with large-scale nucleic acid sequence data sets from different sources.
[0051] (3) By introducing the position information embedding network and the cross-domain feature extraction network into the autoregressive model, the model can better understand and integrate sequence features.
[0052] (4) The model’s unsupervised training method relies on the information of the sequence itself and does not require external labels, reducing dependence on professional knowledge and the complexity of preprocessing.
[0053] (5) The molecular features obtained through pre-training methods can provide support for downstream tasks and subsequent applications such as disease prediction and classification, and drug development. BRIEF DESCRIPTION OF THE DRAWINGS
[0054] Figure 1 A full flowchart for model training based on autoregressive large models;
[0055] Figure 2 Diagram of the method for constructing nucleic acid sequence features based on autoregressive large models. DETAILED DESCRIPTION
[0056] Traditional biological analysis methods study nucleic acid sequence characteristics through gene sequencing and comparative analysis, which has the defects of low efficiency and incomplete information coverage. With the development of computer technology and statistics, machine learning methods have been introduced to analyze nucleic acid sequence characteristics, but there is a problem that domain experts are required to perform feature engineering, and high-dimensional features containing rich recognition patterns cannot be obtained. Existing models based on deep learning neural networks tend to focus on smaller-scale nucleic acid data sets, and there is insufficient mining of long sequence features, which cannot effectively deal with large-scale data sets.
[0057] In order to fully mine the potential patterns of nucleic acid sequence data sets and comprehensively utilize large-scale data sets to extract nucleic acid features for downstream nucleic acid analysis tasks, the present invention proposes a nucleic acid sequence feature mining method based on an autoregressive large model. The one-dimensional nucleic acid sequence is deeply analyzed by unsupervised learning. First, the frequency statistics method is used to identify and select effective k-mer nucleotide fragments. These fragments are used as the basis for segmentation, and then feature training is performed under the framework of the autoregressive transformer model to generate dense sequence embedding features containing rich semantic information. The efficiency of feature extraction is optimized, providing support for subsequent applications such as disease prediction and classification, and drug development.
[0058] The principle of the present invention is to recursively merge k-mer fragments of nucleic acid sequences based on statistical laws, map nucleic acid sequences into digital sequences using k-mer fragments, input them into a large autoregressive model for effective learning, and generate high-dimensional nucleic acid features with strong semantic information. The method mainly consists of: k-mer fragment generation method, unsupervised autoregressive model, composed of position information embedding network, cross-domain feature extraction network and monomer feature integration network, feature output network, and the technical methods used are as follows:
[0059] First, large-scale nucleic acid sequence data sets from different sources are collected, including public bioinformatics databases and clinical sample libraries of medical and scientific research institutions. After data collection, necessary preprocessing is performed to summarize the k-mer fragments of nucleic acids based on frequency statistics to form complete k-mer fragments. The k-mer fragments are then used to encode nucleic acid sequences. The digital encoding form of the sequence is sequentially extracted through an autoregressive unsupervised large model consisting of a position information embedding network, a cross-domain feature extraction network, a monomer feature integration network, and a feature output network. Finally, high-dimensional nucleic acid sequence features are obtained to provide support for downstream nucleic acid analysis tasks.
[0060] The large-scale nucleic acid sequence data sets include: clinical sample libraries from public bioinformatics databases such as NCBI, Ensembl, and medical and scientific research institutions. These data sets cover a wide range of nucleic acid sequence information collected from multiple different sources, ensuring the diversity and representativeness of the data.
[0061] The preprocessing includes: performing quality control on the collected nucleic acid sequence data, removing low-quality sequences, performing necessary standardization processing and removing duplicate sequences, performing necessary format conversion, archiving and cataloging the sequence data, and establishing a data index.
[0062] The quality control of nucleic acid sequence data includes: preliminary evaluation of the raw data, screening out sequences that may have poor quality due to experimental conditions or technical problems. Checking the noise ratio in the sequence, identifying and eliminating sequences containing a large number of unknown bases. In addition, for inconsistent read lengths, trimming or padding will be performed to achieve a consistent length to ensure the consistency of each sequence.
[0063] The standardization of nucleic acid sequence data includes: unifying the base representation of all sequences into the same format, unifying the letter expression format and expression specification, and unifying the direction of the sequence to ensure that all sequences are analyzed and processed in the same direction.
[0064] The removal of duplicate sequences from nucleic acid sequence data includes: identifying and removing identical sequences, which may be caused by errors in experimental sampling or data processing. Checking the degree of match between each sequence and all other sequences in the data set one by one, identifying completely matching sequences, and retaining only one copy of them.
[0065] The necessary format conversion of the nucleic acid sequence data includes: converting nucleic acid sequence files obtained from multiple data sets into different formats, unifying them into the same file format, and adding specific tags and identifiers according to different format characteristics.
[0066] The archiving and cataloguing of nucleic acid sequence data includes: establishing an index for the data so that specific sequences can be quickly retrieved and accessed in subsequent analysis.
[0067] The method of summarizing the k-mer fragments of nucleic acids based on the frequency statistics method includes: scanning the entire data set. In this step, the system will traverse each nucleic acid sequence and count all possible nucleotide pairs or longer fragments in the sequence. The length of these fragments can start from two nucleotide pairs and gradually expand to longer sequences. The system will record the number of times each k-mer fragment appears in the entire data set, thereby calculating the frequency of occurrence of each fragment. The frequencies of all possible k-mer fragments are recorded, and high-frequency k-mer fragments are selected according to a pre-set threshold. It is used to determine which k-mer fragments whose frequencies meet or exceed this standard should be considered important or representative. After selecting these high-frequency fragments, the system will perform a recursive merging operation, that is, merging multiple adjacent high-frequency k-mer fragments into longer sequence fragments, and this process continues until no further merging is possible. The k-mer fragments finally formed will cover different lengths and structures, and these fragments will be used in subsequent encoding steps.
[0068] The process of encoding a nucleic acid sequence using k-mer fragments includes mapping each k-mer fragment to a unique integer identifier. This mapping process involves creating a correspondence from k-mer fragments to integers, ensuring that each different k-mer fragment corresponds to a unique integer. Two additional integer identifiers are introduced to indicate the beginning and end of the sequence. These two identifiers are special integers that are used to specifically mark the starting point and end point of the sequence. During the encoding process, each nucleic acid sequence will be converted into a series of integers.
[0069] The position information embedding network includes: using a position embedding algorithm based on sine and cosine functions to calculate the position features of different positions of the input nucleic acid features, and fusing them with the input nucleic acid features to achieve alignment of nucleic acid sequence features and position features. During the encoding process, the position of each k-mer and its position vector combination in the entire sequence encoding is consistent with the position of the k-mer in the original sequence.
[0070] The detailed process of the cross-domain feature extraction network first involves the use of an adaptive attention mechanism. The attention mechanism is used to automatically adjust the focus of processing according to the specific content of the nucleic acid sequence data. At the same time, the network implements a multi-head attention mechanism. The system performs the operation of the attention mechanism in parallel and processes multiple information streams in parallel. The system creates multiple independent attention modules, each of which is responsible for capturing different aspects of the nucleic acid sequence data. These modules work in parallel at runtime, and each module focuses on different parts of the sequence and different types of information. Each attention module conducts an in-depth analysis of the part of the sequence it focuses on, extracts the key biological features of that part, and integrates the local and global information of the nucleic acid sequence. These features are then fed into an integration module, in which information from different attention modules is integrated.
[0071] The monomer feature integration network includes: receiving features obtained from the cross-domain feature extraction network, and the monomer feature integration network is first processed through a series of nonlinear transformation layers. These layers mainly include fully connected layers, and each fully connected layer is usually followed by an activation function. The input feature data is linearly transformed, that is, the sequence feature data is weighted and operated, and the features obtained from the cross-domain feature extraction network are further integrated and optimized through a series of nonlinear transformation layers (fully connected layers and activation functions). The repeated superposition of fully connected layers and activation functions realizes the in-depth processing of input features layer by layer.
[0072] The feature output network includes: converting the processed sequence features into the final output format, which are usually high-dimensional feature vectors optimized for specific bioinformatics tasks (such as gene expression prediction, variation detection, etc.). The final feature mapping layer and normalization layer are included to ensure the stability and availability of the output features.
[0073] The present invention provides a nucleic acid sequence feature mining method based on an autoregressive large model, which extracts sequence embedding features with high semantic information through unsupervised learning on a large-scale nucleic acid sequence data set, and is used for downstream tasks such as disease prediction and drug development.
[0074] In order to better illustrate the above content, the present invention is specifically described through the following steps:
[0075] First, appropriate data is selected from a variety of data sources and preliminarily screened through a series of scoring mechanisms to ensure the quality and relevance of the data. Then, the selected sequence data is standardized and duplicate sequences are removed to ensure data consistency. The characteristics of the sequence data are analyzed by the k-mer frequency statistical method, and recursive merging is performed based on high-frequency k-mers to form longer sequence fragments. Subsequently, the k-mer fragments in the sequence are encoded as integer identifiers, and special identifiers are added to mark the beginning and end of the sequence. In order to more accurately describe the sequence characteristics, a method combining position vectors and k-mer encoding is introduced, and feature vectors are integrated through weighted merging techniques. In terms of feature extraction, adaptive attention mechanisms and multi-head attention mechanisms are used to capture key information in the sequence, and the nonlinear processing capabilities of the model are enhanced through fully connected layers and activation functions. Finally, feature mapping layers and normalization are used to form feature vectors suitable for specific bioinformatics tasks. Model training adopts an autoregressive-based framework to iteratively optimize model parameters by predicting the next k-mer in the sequence, and uses cross-entropy loss functions and gradient descent methods to update parameters to improve prediction accuracy and model generalization ability. The specific steps of the above method are:
[0076] (1) Construction process of large-scale nucleic acid sequence datasets
[0077] The first step is to select and access the data source
[0078] Considering the diversity and coverage of data, we selected public bioinformatics databases and clinical sample libraries of professional medical and research institutions. Specific data sources include NCBI GenBank, Ensembl and others such as 1000GenomesProject. These databases provide a wide range of biological sequence data, including but not limited to nucleic acid sequences of humans, animals, plants and microorganisms. When selecting data sets, the following data set selection weight formula is adopted:
[0079]
[0080] Where D i represents the amount of data selected from data set i, Dtotal is the total amount of data, Q iTo select a dataset, the data quality score of dataset i is set to the interval between 0 and 1 to take into account factors such as sequencing quality and annotation completeness. i R represents the relevance score between dataset i and the target task. i represents the data richness score of dataset i to reflect the size and diversity of the dataset. The denominator ∑(Q j ×C j ×R j ) is used for the sum of all dataset scores and is used for normalization.
[0081] Step 2 Preliminary screening of sequence data
[0082] After selecting the data source, perform preliminary data screening. This step involves defining screening criteria, including sequence length, quality score, and release date. The weight formula for preliminary screening of sequence data is as follows:
[0083]
[0084] Where S i represents the comprehensive score of sequence i, which is used to determine whether to retain the sequence, W L Represents the sequence length weight, reflecting the importance of sequence length in screening. L Represents the sequence length score, and controls the sequence length within an appropriate range. (Length-MinLength) / (MaxLength-MinLength) is used as the definition method in the present invention. W Q Indicates the sequence quality weight, reflecting the importance of sequence quality in screening, S Q represents the sequence quality score. The quality score is directly used in the present invention. D represents the release date weight, reflecting the importance of the release date in the screening, S D Represents the release date score, which adjusts the data based on the serial release date.
[0085] Step 3: Data integration and standardization
[0086] Integrate and standardize data from different sources. The integration process includes unifying data formats, merging identical or duplicate records, and building a unified data set. Standardization includes unifying sequence naming conventions and adjusting sequence directionality to ensure that all sequences are in the 5' to 3' direction. In addition, the sequence metadata, including species name and sampling location, is formatted to ensure consistency and completeness of the information.
[0087] (2) Data cleaning process (preprocessing) based on nucleic acid sequence data set
[0088] The first step is to remove low-quality sequences
[0089] This step uses the quality scoring function Q i The quality score of the obtained nucleic acid sequence is as follows:
[0090]
[0091] Where Q i is the quality score of sequence i, n is the number of bases in sequence i, and q j is the quality score of the jth base. This score is provided by the sequencing platform and reflects the probability of sequencing errors. Set the threshold Q min The value should be above 0.9 to ensure the high quality of the sequence. i ≥Q min , sequence i will be retained, otherwise it will be eliminated.
[0092] The second step is standardization and removal of duplicate sequences
[0093] In nucleic acid sequence datasets, different data sources may use different formats or standards, so standardization is required. First, the directionality of all sequences is unified to ensure that each sequence is arranged from 5' to 3'. In addition, the naming and metadata of the sequences (such as species name, sampling location, etc.) are uniformly formatted to facilitate subsequent data management and analysis.
[0094] Use hash table technology to identify and delete duplicate sequences. The specific method is as follows:
[0095]
[0096] Where S i is the sequence i, l is the sequence length, c k is the kth character in the sequence, α k is a weight coefficient, here we take k 2 , to reduce the probability of collision. By comparing the hash values of the sequences, duplicate sequences can be quickly identified and deleted, that is, sequences with consistent hash values can be deleted.
[0097] Step 3 Further screening and classification of sequence data
[0098] The formats of data sets that save sequences in different formats in multi-source data sets are unified and unified into FASTA format as the target format. During the format conversion process, each sequence file is first parsed for the specific structure and content of its original format. For the FASTQ format, the nucleic acid sequence and quality score are extracted, the quality score is discarded, and only the sequence is saved. For the GenBank format, the nucleic acid sequence and its annotation information are extracted, and the annotation information is converted into a label form and attached to the sequence name. All sequences are converted into a unified FASTA format, in which each sequence consists of an identifier and pure sequence data. The sequence data after format conversion is archived and cataloged, and the data index is established so that specific sequences can be quickly retrieved and accessed in subsequent analysis.
[0099] (3) Constructing k-mer process based on frequency statistics of nucleic acid sequence data sets
[0100] First step: frequency statistics of k-mer fragments
[0101] The system first needs to traverse the entire nucleic acid sequence data set and scan each sequence. The purpose of the scan is to count all possible k-mer fragments (k represents the fragment length, k is variable), starting from the smallest length, such as a dinucleotide pair (i.e. k = 2), and gradually increasing the length until the set maximum length.
[0102] For each possible length k, the system will count the number of occurrences of each k-mer. Let S be a specific nucleic acid sequence with a length of n, then the number of possible k-mers in the sequence is n-k+1. The statistics for each k-mer can be expressed as:
[0103]
[0104] Among them, C k (x) represents the number of times k-mer x appears in sequence S, 1(S i:i+k-1 =x) is the indicator function, S i:i+k-1 Represents the substring of sequence S from position i to i+k-1.
[0105] By performing the above statistics on all sequences, we can get the total number of occurrences of each k-mer in the entire data set. Then, calculate the frequency of each k-mer:
[0106]
[0107] The second step is to select high-frequency k-mer fragments
[0108] After counting the frequencies of all k-mers, the next step is to filter out high-frequency k-mer fragments based on a preset threshold. If the threshold is set to θ, then all k-mer fragments that satisfy fk (x) ≥ θk-mers are regarded as high-frequency fragments with high representativeness and importance.
[0109] The third step is to recursively merge high-frequency k-mer segments
[0110] After selecting the high-frequency k-mer segments, a recursive merging operation is performed to form longer sequence segments. This step is performed by checking whether adjacent high-frequency k-mer segments can be merged.
[0111] The merging rule is defined as follows: if two high-frequency k-mer fragments x and y frequently appear adjacent to each other in the dataset, they are merged into a new fragment xy. This process is performed recursively, and a threshold is set to determine whether a merge is needed until there are no more merges possible. In this way, the resulting k-mer fragments will cover different lengths and structures.
[0112] (4) Specific steps for encoding nucleic acid sequences
[0113] The first step is to replace the k-mer fragments with integer identifiers
[0114] In this step, the system replaces each k-mer segment identified in the nucleic acid sequence with a predefined integer identifier. This process first involves building a mapping table that maps each unique k-mer segment to a unique integer. Let the mapping function be M, which accepts a k-mer segment x as input and returns an integer identifier M(x).
[0115] For each position i in the nucleic acid sequence S, the corresponding k-mer is S i:i+k-i , the replacement operation can be defined as:
[0116] S′ i =M(S i:i+k-1 )
[0117] Among them, S′ i is the i-th element in the replaced sequence. This operation converts the original nucleic acid sequence S into a sequence S' consisting of integers.
[0118] Step 2: Mark the start and end of the sequence and save it
[0119] In order to identify the beginning and end of the sequence in subsequent processing, the system adds two additional integer identifiers at the beginning and end of the integer sequence S'. The identifier of the beginning of the sequence is set to α and the identifier of the end is set to ω. Therefore, the encoded sequence S" can be expressed as:
[0120] S″=[α,S′1,S′2,…,S′ n,ω]
[0121] Here, n is the length of the replaced sequence S'. Adding these special identifiers helps to clearly define the boundaries of the sequence during sequence processing and analysis. Finally, the encoded sequence S'' is saved to a database or file for subsequent data processing and analysis.
[0122] (5) Constructing features that carry positional information in the input nucleic acid sequence
[0123] The first step is to generate a position vector and combine it with the k-mer code
[0124] In this step, the system first generates a position vector for each k-mer segment in the sequence. The position vector is generated based on the specific position of the k-mer in the sequence, ensuring that k-mers at different positions correspond to different position vectors. Suppose a k-mer in the sequence is at position p, then its position vector It can be calculated by the following sine and cosine functions:
[0125]
[0126] Where j represents the dimension index in the position vector and d is the total number of dimensions of the position vector. This calculation method ensures that the vector of each position is unique and the vectors of adjacent positions are continuous in the high-dimensional space.
[0127] The system then associates the encoding of each k-mer (such as the integer identifier M(·) in the previous step) with its corresponding position vector Combine to form an enhanced feature vector
[0128]
[0129] The type of k-mer and its position in the sequence are considered simultaneously.
[0130] The second step is to weight the position vector and k-mer encoding.
[0131] Each enhanced feature vector Next, we will weight them according to specific weights to strengthen their importance and uniqueness in the sequence. Let the weight vector be Then the weighted eigenvector It can be expressed as:
[0132]
[0133] The weighted vector will be further merged into a single sequence feature vector Used to represent the characteristics of the entire sequence. The merging operation is implemented by vector superposition, and the formula is:
[0134]
[0135] n is the total length of the sequence, k is the length of the k-mer, and n-k+1 represents the total number of k-mers in the sequence. In this way, the weighted features of each k-mer and its position information are effectively fused to form a comprehensive vector that fully reflects the characteristics of the sequence.
[0136] (6) Specific steps for constructing cross-domain features of input nucleic acid sequences
[0137] The first step is to implement the adaptive attention mechanism
[0138] In this step, the system uses an adaptive attention mechanism to dynamically adjust the processing focus of nucleic acid sequence data. This mechanism assigns attention weights by calculating the importance of each sequence element to the overall sequence analysis. Assume that the sequence consists of n elements, each element is represented by x i , where i ranges from 1 to n. The attention weight α i Calculated by the following formula:
[0139]
[0140] Among them, e i is the element x i The energy score is usually calculated by a learnable function f, that is, e i =f(x i ). This weight distribution ensures that a high level of attention is paid to important elements of the sequence.
[0141] The second step is to implement the multi-head attention mechanism
[0142] In order to be able to process multiple information streams in parallel and capture different aspects of the sequence data, the system implements a multi-head attention mechanism. Under this mechanism, the system creates k independent attention modules, each responsible for analyzing different parts of the sequence and different types of information. Each module h uses the same adaptive attention mechanism, but operates with an independent set of parameters, allowing the modules to focus on capturing different features. The output o of each module h The calculation is as follows:
[0143]
[0144] Among them, α h,i is the hth head in element x i This allows each head to focus on different parts of the sequence.
[0145] The third step is to integrate the output of the multi-head attention module
[0146] In this step, the output information from different attention modules needs to be integrated into a unified output for further analysis. The integration process usually involves merging the outputs of all heads and processing them through an integration module. The integrated output O can be calculated as follows:
[0147] O=Concat(o1,o2,...,o k )·W
[0148] Among them, Concat means connecting the outputs of all heads into a long vector, and W is a learnable weight matrix used to convert the connected output into the final sequence feature representation. This step ensures that the information focused on by different attention modules is effectively integrated to form a comprehensive feature vector that fully reflects the characteristics of the nucleic acid sequence.
[0149] (7) Constructing the characteristics of nucleic acid monomers after integration
[0150] The first step is to apply a fully connected layer for linear transformation
[0151] The single feature integration network receives the output feature vectors from the cross-domain feature extraction network and performs a preliminary linear transformation on these features. In this step, the input feature vector v is first processed through a series of fully connected layers. Each fully connected layer can be regarded as a weighted summation operation of the input features, and its mathematical expression is as follows:
[0152] z (l) =W (l) v (l-1) +b (l)
[0153] Among them, W (l) and b (l) Represent the weight matrix and bias vector of the lth layer, v (l-1) is the output feature vector of the previous layer, z (l) is the linear transformation output of the current layer.
[0154] The second step is to enhance the nonlinear characteristics through activation function
[0155] After the linear transformation output of the fully connected layer, the activation function is applied to nonlinearly transform these outputs. This step allows the network to capture more complex data patterns by introducing nonlinear operations. The application of the activation function can be expressed as:
[0156] v (l) =φ(z (l) )
[0157] Here, φ(·) represents the activation function, using the ReLU function, which applies max(0,x) to each element, where x is the input element. Through such nonlinear transformation, the output v of each layer (l) Not only does it retain the linear combination of the previous layer features, but it also adds nonlinear processing, allowing the entire network to learn more complex and abstract feature representations.
[0158] (8) Output integrated nucleic acid features
[0159] The first step is to apply the feature map layer
[0160] The first key component of the feature output network is the feature mapping layer, whose main function is to convert the comprehensive feature vector processed by the previous network layer into a format suitable for specific bioinformatics tasks. This layer usually uses a fully connected layer to achieve the final mapping of the features, and its mathematical expression is as follows:
[0161] y=W o v+b o
[0162] Among them, W o and b o Represent the weight matrix and bias vector of the feature mapping layer, v is the feature vector passed from the previous network layer, and y is the feature vector after mapping. This step ensures that the feature vector can be adjusted and optimized to match specific bioinformatics analysis tasks, such as gene expression prediction or variation detection.
[0163] The second step is to normalize the eigenvector
[0164] In order to ensure the stability and availability of the output features, the feature vector will be normalized next. The purpose of the normalization layer is to adjust the scale of the feature data to meet the needs of subsequent processing and reduce the instability factors in model training. Normalization can be achieved in many ways, one of the common methods is batch normalization, which is applied well in the present invention, and its mathematical expression is:
[0165]
[0166] Here, y′ is the output of the feature mapping layer, μ and σ are the mean and standard deviation of y′, respectively, and γ and β are learnable parameters used to further adjust the normalized data. Such processing not only keeps the feature vector consistent in different data batches, but also helps to accelerate the model training process and improve the generalization ability of the model.
[0167] In addition, other normalization methods can be used, such as:
[0168] 1. The normalization layer performs normalization on the feature dimension.
[0169]
[0170] in, The output of the feature mapping layer, μ and σ are the mean and standard deviation of the feature dimension respectively.
[0171] 2. Instance Normalization
[0172]
[0173] Among them, x ij is the jth feature of the i-th nucleotide. i and σ i are the mean and standard deviation of the th instance respectively. γ and β are learnable parameters.
[0174] (9) Optimize model training methods
[0175] In the present invention, the training method of the model is based on an autoregressive framework, and the model parameters are iteratively learned and optimized by predicting the next k-mer in the sequence. Specifically, the model receives a sequence fragment in each training step and then predicts the subsequent k-mer of the fragment. The goal of the model is to minimize the difference between the predicted k-mer and the actual k-mer. This method does not rely on external labels, but learns through the information of the sequence itself.
[0176] Assume that the input sequence fragment of the model at time point t is x t , the goal of the model is to predict the next k-merk t+1 The output of the model is Yes t+1 The model adjusts its internal parameter Θ by minimizing the loss function L between the predicted output and the actual k-mer. The loss function usually uses the cross entropy loss (used in this example), which is expressed as:
[0177]
[0178] Where N is the total number of k-mers in the sequence, is the actual i-th element of the next k-mer, is the i-th element of the next k-mer predicted by the model.
[0179] During the training process, the model parameters θ are updated by the gradient descent method, and the update formula is:
[0180]
[0181] Where η is the learning rate, is the gradient of the loss function L with respect to the parameter Θ.
[0182] The above description is only an embodiment of the present specification and is not intended to limit the present specification. For those skilled in the art, the present specification may have various changes and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present specification shall be included in the scope of the claims of the present specification.
Claims
1. A method for mining nucleic acid sequence features based on an autoregressive macromodel, characterized in that: include: Acquire a large-scale nucleic acid sequence data set, and preprocess the nucleic acid sequence data set; Inducing k-mer fragments of nucleic acid based on frequency statistics; wherein k represents the fragment length, starting from the smallest length, and gradually increasing the length until the set maximum length; the inducing k-mer fragments of nucleic acid based on frequency statistics includes the following sub-steps: scanning the entire nucleic acid sequence data set, calculating the occurrence frequency of all possible k-mer fragments in the sequence; setting a threshold, screening k-mer fragments with an occurrence frequency greater than or equal to the threshold; recursively merging the screened k-mer fragments; the recursive merging is specifically: if two screened k-mer fragments x and y frequently appear adjacent to each other in the data set, merging them into a new fragment xy, setting a threshold to determine whether merging is required, until there is no more possibility of merging; Using the k-mer fragment to encode a nucleic acid sequence; An autoregressive transformer model is constructed, and the encoded nucleic acid sequence is input for unsupervised training; the autoregressive transformer model includes a multi-layer coupled position information embedding network, a cross-domain feature extraction network, a monomer feature integration network and a feature output network; wherein the cross-domain feature extraction network is used to implement an adaptive attention mechanism and a multi-head attention mechanism; Output the processed nucleic acid sequence features.
2. The method according to claim 1, characterized in that The preprocessing includes removing low-quality sequences, standardizing and removing duplicate sequences; The removal of low-quality sequences comprises: performing a quality score on the sequence according to the quality score of each base in the obtained nucleic acid sequence; setting a threshold and retaining sequences with a quality score greater than or equal to the threshold; The standardization process is to arrange the sequences in a uniform direction; The deduplication sequence comprises: calculating the hash value of each sequence in the data set, and deleting the sequence with the same hash value.
3. The method according to claim 1, characterized in that Using the k-mer fragment to encode a nucleic acid sequence comprises the following sub-steps: Map each recursively merged k-mer fragment to a unique integer identifier , generates a sequence of integers, where is the mapping function, Nucleic acid sequence Position in The corresponding k-mer fragment; Two additional integer identifiers are introduced to indicate the start and end of the sequence, respectively, to generate the encoded sequence.
4. The method according to claim 3, characterized in that The location information is embedded in the network and the following sub-steps are performed: Based on the specific position of each k-mer fragment in the sequence, generate the corresponding position vector : The integer identifier of each k-mer fragment The corresponding position vector Combine to generate the corresponding enhanced feature vector : ; For each enhanced feature vector Perform weighted merging.
5. The method according to claim 4, characterized in that The position vector It is obtained by the following formula: ; in, represents the dimension index in the position vector, is the total number of dimensions of the position vector.
6. The method according to claim 1, characterized in that The following sub-steps are performed in the cross-domain feature extraction network: Based on the adaptive attention mechanism, the importance of each element in the sequence for the overall sequence analysis is calculated and the attention weight is assigned. : ; ; in, Is an element The energy fraction, is a learnable function; Based on the multi-head attention mechanism, m independent attention modules are created to capture different features respectively; each module h uses the same adaptive attention mechanism to perform weighted fusion on each element according to the attention weight and output the corresponding vector ; After that, merge and output a comprehensive feature vector : ; in, The output of the module h is Concatenate into a long vector, is a learnable weight matrix.
7. The method according to claim 6, characterized in that The following sub-steps are performed in the monomer feature integration network: A fully connected layer is used to transform the comprehensive feature vector Perform linear transformation; Through the activation function, the linearly transformed comprehensive feature vector is nonlinearly transformed.
8. The method according to claim 1, characterized in that The following sub-steps are performed in the feature output network: The feature mapping layer is used to map the comprehensive feature vector output by the monomer feature integration network and convert it into a coding form corresponding to the original nucleic acid sequence; A normalization layer is used to normalize the mapped comprehensive feature vector and output the processed nucleic acid sequence.
9. The method according to claim 1, characterized in that: The unsupervised training includes the following sub-steps: Each time a sequence is received, the next k-mer in the sequence is predicted; the loss function between the predicted k-mer and the actual k-mer is compared, and the internal parameters are adjusted with the goal of minimizing the loss function.
Citation Information
Patent Citations
RNA-protein binding site prediction method and system based on self-attention mechanism
CN114023376A
DNA synthesis difficulty prediction system and application thereof
CN116312783A
RNA-protein interaction prediction method and apparatus, and medium and electronic device
WO2023044927A1