Method and device for processing methylation feature coding model

By extending the DNA base set and using the MosaicBERT model to construct a methylation feature coding model, the cumbersome and cost-effective BS-seq method is solved, and efficient methylation sequence analysis and target source recognition are achieved.

CN120299532AActive Publication Date: 2025-07-11PEKING UNIVERSITY CHENGDU ACADEMY FOR ADVANCED INTERDISCIPLINARY BIOTECHNOLOGIES +1
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202510354822.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-25
Publication Date
2025-07-11
Estimated Expiration
2045-03-25

AI Technical Summary

Technical Problem

Existing DNA methylation analysis methods such as BS-seq are cumbersome and costly, resulting in long task processing cycles and high cost, making it difficult to efficiently perform methylation sequence analysis and target source identification.

Method used

The quaternary base set of DNA [A,T,C,G] is extended to the five-base set [A,T,C,G,M], and the methylated feature coding model is constructed using the MosaicBERT model. Combined with the BPE algorithm and vocabulary learning, a methylated feature coding and sequence decoding model is constructed. The model is optimized through pre-training and fine-tuning to handle the methylated sequence reconstruction and target source recognition task.

Benefits of technology

Simplify task processing steps, shorten processing cycles, reduce costs, and improve real-time and efficiency of task processing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120299532A_ABST
    Figure CN120299532A_ABST
Patent Text Reader

Abstract

The embodiment of the invention relates to a methylation feature coding model processing method and device. The method comprises the steps that a BS-seq read set is collected; performing read splicing according to the read set to obtain a first sequence, performing five-element sequence conversion on the first sequence to obtain a second sequence, and performing vocabulary learning based on the second sequence; constructing a methylation feature coding model, a sequence decoding model and a target source prediction model; the methylation feature coding model and the sequence decoding model form a pre-training model framework, and the methylation feature coding model and the target source prediction model form a first task model framework; training the pre-training model framework and the first task model framework; and after the training is finished, processing the methylation sequence reconstruction task based on the pre-training model framework, and processing the DNA source identification task based on the first task model framework. According to the invention, the task processing real-time performance can be improved, and the task processing cost is reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of data processing, and particularly relates to a processing method and device for a methylation feature encoding model. Background Art

[0002] Epigenomics is an important field for studying genomic chemical modifications and their regulatory mechanisms. Among them, DNA methylation is a key form of epigenetic modification, usually existing in the form of 5-methylcytosine (5-mC) at CpG sites. Here, the CpG site refers to the base sequence in a certain region of DNA that appears in the form of cytosine base (C) followed by guanine base (G), where "CpG" is the abbreviation of "C-phosphodiester bond p-G". DNA methylation plays an important role in biological processes such as gene expression regulation, cell differentiation, and genomic imprinting. Its abnormal patterns (such as abnormal methylation positions) are often closely related to various diseases. The range of diseases mentioned here involves a series of tumor (also known as cancer) diseases, cellular leukemia, nervous system diseases, immune system diseases, etc. Therefore, accurately predicting the overall methylation state of DNA sequences is of great significance for basic disease research.

[0003] The recognized gold standard sequencing method in the field of DNA methylation is bisulfite sequencing (BS-seq). Through the BS-seq method, the methylation sites corresponding to 5-mC in the DNA sequence can be determined. The conventional processing process of the BS-seq method is roughly divided into two steps: 1) First, a DNA sample that has completed the initial sequencing is exposed to a bisulfite treatment environment. In this treatment process, the unmethylated cytosine C in the current sample is converted into uracil (U), and uracil U will be replaced by thymine (T) in the subsequent DNA amplification process. The originally methylated cytosine C, that is, 5-methylcytosine 5-mC, on the sample will not be converted during this treatment process; 2) Then, the processed DNA sample is sequenced by high-throughput sequencing technology to obtain a new DNA sequence. By comparing the two sequences before and after the treatment, the methylation status (yes / no) of each base can be obtained, and thus the overall methylation status of the current DNA sample can be obtained. After obtaining the methylation information of the DNA sample based on the BS-seq method, a series of bioinformatics recognition tasks related to basic disease research can be carried out, such as identifying the sample source based on the methylation information of the DNA sample. The sample source mentioned here can be a type of tissue (such as heart, liver, spleen, lung, stomach, duodenum, colon, etc.) or a type of cell (such as T cells, NK cells, monocytes, macrophages, granulocytes, and B cells, etc.). However, in the actual process, we found that because the steps of the BS-seq method are relatively cumbersome and the usage costs of various experimental equipment and sequencing equipment during the process are also relatively high, if the methylation sequence analysis task and bioinformatics recognition task of the DNA sample are processed based on the conventional BS-seq method, problems such as long task processing cycle and high processing cost will be faced. Summary of the Invention

[0004] The object of the present invention is to provide a processing method, device, electronic device and computer-readable storage medium for a methylation feature encoding model in view of the defects of the prior art. The present invention sets a corresponding base marker M for 5-mC in the DNA sequence, and adjusts the four-base set [A, T, C, G] of DNA to a five-base set [A, T, C, G, M]; and performs big data collection on bisulfite sequencing reads of a certain species, and performs processing such as read splicing and methylation status marking based on the collected BS-seq read set to obtain a base sequence, and performs vocabulary learning based on the BPE algorithm and this base sequence; then constructs an encoder model with the MosaicBERT model as the core, denoted as the methylation feature encoding model, stores the learned vocabulary in the vocabulary storage module inside the encoding model, constructs a decoder model for methylation sequence reconstruction, denoted as the sequence decoding model, and constructs a corresponding binary classification prediction model for a class of target source types, denoted as the target source prediction model; then, a pre-training model framework is composed of the methylation feature encoding model and the sequence decoding model, and a first task model framework is composed of the methylation feature encoding model and the target source prediction model; then, first train the pre-training model framework based on the above base sequence, and then train the first task model framework; finally, process the methylation sequence reconstruction task of the DNA sequence based on the pre-training model framework, and process a class of DNA sequence source identification tasks based on the first task model framework. The methylation feature encoding model provided by the present invention can learn the methylation features of the DNA sequence. The pre-training model framework with this methylation feature encoding model as the core encoder can be used to process methylation sequence analysis tasks, and the first task model framework with this methylation feature encoding model as the core encoder can be used to process a class of target source identification tasks; based on the methylation feature encoding model provided by the present invention and the corresponding two task model frameworks to process methylation sequence analysis tasks and target source identification tasks, not only can the task processing steps be simplified, the task processing cycle be shortened, the real-time performance of task processing be improved, but also the processing cost of the task can be effectively reduced.

[0005] To achieve the above object, a first aspect of an embodiment of the present invention provides a processing method for a methylation feature encoding model, the method comprising:

[0006] Set a corresponding extended marker for 5-methylcytosine in the DNA sequence, denoted as the base marker M, and adjust the four-base set [A, T, C, G] of DNA to the corresponding five-base set [A, T, C, G, M]; and collect big data of bisulfite sequencing reads of the first species to form a corresponding BS-seq read set; and perform read splicing according to the BS-seq read set to obtain a corresponding first base sequence; and perform a five-element sequence conversion on the first base sequence based on the five-base set [A, T, C, G, M] and the methylation status markers corresponding to each base marker C on the first base sequence to obtain a corresponding second base sequence; and perform vocabulary learning based on the BPE algorithm and the second base sequence to obtain a corresponding first vocabulary;

[0007] Construct a methylation feature encoding model with the MosaicBERT model as the core; and store the first vocabulary in the vocabulary storage module of the methylation feature encoding model; and construct a sequence decoding model for methylation sequence reconstruction; and construct a corresponding binary classification prediction model based on a class of target source types, denoted as the target source prediction model; the target source type is specifically a class of tissue types or a class of cell types;

[0008] The pre-training model framework is composed of the methylation feature encoding model and the sequence decoding model; and perform the first framework training on the pre-training model framework according to the second base sequence to obtain the corresponding pre-training parameters of the encoding model;

[0009] The first task model framework is composed of the methylation feature encoding model and the target source prediction model; and construct a model data set based on the BS-seq read set and the target source type; and perform the second framework training on the first task model framework based on the pre-training parameters of the encoding model and the model data set;

[0010] After the first and second framework trainings are completed, process the methylation sequence reconstruction task of the DNA sequence based on the pre-training model framework; and process a class of DNA sequence source identification tasks corresponding to the target source type based on the first task model framework.

[0011] Preferably, the base markers A, T, C, and G are the markers for adenine, thymine, cytosine, and guanine bases respectively;

[0012] The BS-seq read set consists of multiple BS-seq read records; each of the BS-seq read records includes at least one BS-seq read and a corresponding read sample source; each of the BS-seq reads is a DNA sequence obtained by DNA sequencing after bisulfite sequencing treatment of a DNA sample of the first species, and is sorted by the base markers A, T, C, G; each base marker C on each of the BS-seq reads has a corresponding methylation status marker; the methylation status marker is a binary status marker, including two states: unmethylated and methylated; the read sample source is a source type code corresponding to the current BS-seq read, and each source type code corresponds to a specific source type; the type range of the source type consists of multiple tissue types and / or multiple cell types of the first species;

[0013] The first base sequence is a marker sequence sorted by the base markers A, T, C, G; the second base sequence is a marker sequence sorted by the base markers A, T, C, G, M;

[0014] The first vocabulary consists of multiple first sub-word records; the first sub-word record includes a first sub-word code, a first sub-word text, and a first sub-word frequency; the first sub-word text is a string sorted by one or more base markers from the five-base set [A, T, C, G, M]; the first sub-word frequency is an integer value; all the first sub-word texts in the first vocabulary are different;

[0015] The model data set includes multiple first data records; the first data record includes a first training sequence and a first label classification result; the first label classification result includes the target source type and non-target source types.

[0016] Preferably, the big data collection of the bisulfite sequencing reads of the first species to form a corresponding BS-seq read set specifically includes:

[0017] Data collection of the bisulfite sequencing reads of the DNA sequence of the first species is performed through multiple preset data collection channels to obtain multiple BS-seq reads; the multiple preset data collections include at least one or more publicly available BS-seq read data sets, and publicly available technical documents with BS-seq read information of the first species;

[0018] And set the type range of the corresponding source type based on the sample source information corresponding to all the collected BS-seq reads; and set a corresponding source type code for each source type within the type range of the source type; and use the source type code corresponding to each BS-seq read as the corresponding read segment sample source; and form a corresponding BS-seq read segment record from each BS-seq read and the corresponding read segment sample source; and perform error read segment deletion and duplicate read segment removal processing on all the obtained BS-seq read segment records;

[0019] And form the BS-seq read segment set from all the remaining BS-seq read segment records.

[0020] Preferably, the obtaining of the corresponding first base sequence by splicing the read segments according to the BS-seq read segment set specifically includes:

[0021] Sequentially splice all the BS-seq reads in the BS-seq read segment set to obtain the corresponding first base sequence.

[0022] Preferably, the obtaining of the corresponding second base sequence by performing five-element sequence conversion on the first base sequence based on the five-base set [A, T, C, G, M] and the methylation status marker corresponding to each base marker C on the first base sequence specifically includes:

[0023] Perform sequence replication on the first base sequence to obtain a corresponding first replication sequence; and modify the base marker C with the status value of methylated for each current methylation status marker on the first replication sequence to the corresponding base marker M; and use the first replication sequence with the marker modification completed as the corresponding second base sequence.

[0024] Preferably, the obtaining of the corresponding first vocabulary by performing vocabulary learning based on the BPE algorithm and the second base sequence specifically includes:

[0025] Step 61, set an empty vocabulary as the corresponding first vocabulary; and initialize the first vocabulary; and set the record total number threshold for the first vocabulary; and use the second base sequence as a corresponding first string;

[0026] Among them, the initialized first vocabulary consists of six initial first sub-word records; the first sub-word texts of the six initial first sub-word records are the character 'A', the character 'T', the character 'C', the character 'G', the character 'M', and the string 'CG' in sequence; the first sub-word frequencies of the six initial first sub-word records are all initialized to 1;

[0027] Step 62: Take each of the first sub-word texts in the first vocabulary as a corresponding basic sub-word, and form a corresponding current sub-word set from all the obtained basic sub-words; and perform a sub-word sequence splitting on the first string according to the sub-word sequence splitting method of the BPE algorithm to obtain a corresponding current sub-word sequence;

[0028] Wherein, the current sub-word sequence is formed by sequentially sorting a plurality of the basic sub-words; the basic sub-word pair is formed by sorting the front and rear two basic sub-words;

[0029] Step 63: According to the sub-word pair combination method of the BPE algorithm, form a corresponding basic sub-word pair from every two adjacent basic sub-words in the current sub-word sequence; and perform clustering on the same basic sub-word pairs to obtain a plurality of clustered word pair sets; and take the total number of the basic sub-word pairs in each of the clustered word pair sets as a corresponding first word pair count, and take the largest of the first word pair counts as a corresponding high-frequency word pair count; and identify the high-frequency word pair count; if the high-frequency word pair count is 1, go to step 68; if the high-frequency word pair count is greater than 1, take the total number of the clustered word pair sets corresponding to the high-frequency word pair count as a corresponding high-frequency word pair total, and identify the high-frequency word pair total, if the high-frequency word pair total is 1, go to step 64, if the high-frequency word pair total is greater than 1, go to step 65;

[0030] Wherein, each of the clustered word pair sets is composed of one or more basic sub-word pairs, and all the basic sub-word pairs in each of the clustered word pair sets are the same;

[0031] Step 64: Take the basic sub-word pair corresponding to the only one clustered word pair set corresponding to the high-frequency word pair total as a corresponding final selected sub-word pair; and go to step 66;

[0032] Step 65: Denote all the clustered word pair sets corresponding to the high-frequency word pair total as corresponding candidate sets; and take the basic sub-word pairs corresponding to each of the candidate sets as corresponding candidate sub-word pairs; and denote the character 'C', the character 'M' and the string 'CG' as methylation-related characters; and denote each of the candidate sub-word pairs containing the methylation-related characters as a preferred sub-word pair; and identify the total number of the preferred sub-word pairs; if the total number of the preferred sub-word pairs is 0, arbitrarily select one from all the candidate sub-word pairs as a corresponding final selected sub-word pair; if the total number of the preferred sub-word pairs is greater than 0, arbitrarily select one from all the preferred sub-word pairs as a corresponding final selected sub-word pair;

[0033] Step 66, add one of the first sub-word records in the first vocabulary as the corresponding currently added record, and set the first sub-word text of the currently added record as the corresponding final selected sub-word pair; after the record setting is completed, use each first sub-word text in the current first vocabulary as a new base sub-word, and form the latest current sub-word set from all the latest obtained base sub-words; and according to the sub-word sequence splitting method of the BPE algorithm, split the first string according to the current sub-word set to obtain the latest current sub-word sequence; and count the total number of each base sub-word that appears in the current sub-word sequence and use the statistical result as the corresponding first word frequency, and set the first word frequency corresponding to each base sub-word that does not appear in the current sub-word sequence in the current sub-word set to 0; and update the first word frequency corresponding to each first sub-word text in the current first vocabulary to the corresponding first word frequency; and after the update is completed, delete the first sub-word records in the first vocabulary with a first word frequency of 0.

[0034] Step 67, count the total number of the first sub-word records in the current first vocabulary to obtain the corresponding current record total; and identify whether the current record total is less than the record total threshold; if so, return to Step 63; if not, go to Step 68.

[0035] Step 68, perform a re-sorting on all the first sub-word records in the latest first vocabulary in descending order of the first word frequency; and based on a preset sub-word encoding rule, perform encoding settings on the first sub-word encodings corresponding to each first sub-word text in the sorted first vocabulary; and output the first vocabulary after the sub-word encoding settings are completed as the vocabulary learning result of this time.

[0036] Preferably, the methylation feature encoding model is used to perform feature encoding on the first sequence X input to the model and output the corresponding first encoding vector Y; the first sequence X is composed of multiple first tokens x i sorted in sequence; each first token x i is a type of base token in the five-base set [A, T, C, G, M]; 1 ≤ token index i ≤ L, where L is the sequence length of the first sequence X; the first encoding vector Y is composed of multiple first segmented feature vectors y j ; 1 ≤ segmentation index j ≤ W, where W is the sequence length of the first segmented sequence S corresponding to the first sequence X.

[0037] The methylation feature encoding model consists of a sequence tokenizer, the vocabulary storage module, and the MosaicBERT model; the input end of the sequence tokenizer is connected to the model input end of the methylation feature encoding model, and the output end is connected to the input end of the MosaicBERT model; the sequence tokenizer is also connected to the vocabulary storage module; the output end of the MosaicBERT model is connected to the model output end of the methylation feature encoding model;

[0038] The sequence tokenizer is used to split the first sequence X into a first sub-word sequence composed of a plurality of the first sub-word texts according to the sub-word sequence splitting method of the BPE algorithm, based on the first vocabulary stored in the vocabulary storage module; and add a preset classification sub-word text 'CLS' before the first sub-word text of the first sub-word sequence; and take the sequence length of the first sub-word sequence after the addition as the corresponding sequence length W; and take each sub-word text of the first sub-word sequence as a corresponding first token s j , and form a corresponding first token sequence S from all the first tokens s j ; and based on the first sub-word encoding corresponding to each of the first sub-word texts in the first vocabulary and the classification sub-word encoding corresponding to the classification sub-word text 'CLS', set the corresponding encoding of each of the first tokens s j in the first token sequence S to obtain corresponding first token encodings, and form a corresponding first token encoding sequence from all the obtained first token encodings; and identify the built-in mask rule type; if the mask rule type is no mask, no mask modification is performed on the first token encoding sequence; if the mask rule type is a type of random mask, randomly mask the first token encodings in the first token encoding sequence according to the built-in type of random mask ratio and the preset token mask encoding, and ensure that the ratio of the total number of masked first token encodings to the sequence length of the first token encoding sequence matches the type of random mask ratio; if the mask rule type is a second type of random mask, then for the first tokens s jDenote it as methylation-related word segments, and denote the first word segment codes corresponding to each of the methylation-related word segments in the first word segment coding sequence as methylation-related codes. Then, perform random masking on the methylation-related codes in the first word segment coding sequence according to the built-in two-category random masking ratio and the word segment masking code, and ensure that the ratio of the total number of the masked methylation-related codes to the total number of the methylation-related codes in the first word segment coding sequence matches the two-category random masking ratio. And based on the model input vector embedding coding rule of the MosaicBERT model, perform embedding coding processing on the first word segment coding sequence to obtain the corresponding first embedding coding vector E and send it to the MosaicBERT model; the first word segment sequence S consists of W first word segments s j and the first first word segment s j=1 is the classification sub-word text 'CLS'. Each of the first word segments s 2≤j≤W from the second to the last first word segment s j corresponds to one of the first sub-word texts in the first vocabulary; the first embedding coding vector E consists of W first word segment embedding coding vectors e j and the first word segment embedding coding vector e j corresponds one-to-one with the first word segment s j ;

[0039] The MosaicBERT model is used to perform feature coding processing according to the first embedding coding vector E and output the corresponding first coding vector Y.

[0040] Preferably, the sequence decoding model is used to reconstruct the methylation sequence according to the first coding vector Y input to the model and output the corresponding first decoding sequence Z; the first decoding sequence Z is sequentially sorted by L second markers z i ; each of the second markers z i is a base marker of a type in the five-base set [A, T, C, G, M]; the second marker z i corresponds one-to-one with the first marker x i ;

[0041] The sequence decoding model consists of a first mapping layer, a first linear layer, and a first Softmax layer; the first mapping layer is implemented based on a fully connected network; the first linear layer is formed based on another fully connected network;

[0042] The input end of the first mapping layer is connected to the model input end of the sequence decoding model, and the output end is connected to the input end of the first linear layer; the output end of the first linear layer is connected to the input end of the first Softmax layer; the output end of the first Softmax layer is connected to the model output end of the sequence decoding model;

[0043] The first mapping layer is used to extract the first word segmentation feature vectors y from the 2nd to the Wth of the first encoding vector Y 2≤j≤W to form a first extraction vector A with a length of W - 1; and convert the first extraction vector A into a first mapping vector M with a length of L through linear transformation and send it to the first linear layer; the first extraction vector A is composed of W - 1 first extraction sub-vectors a k where 1 ≤ vector index k ≤ (W - 1), and each of the first extraction sub-vectors a k matches the corresponding first word segmentation feature vector y j=k+1 ; the first mapping vector M is composed of L first sub-mapping vectors m i ;

[0044] The first linear layer is used to perform a fully connected calculation based on the first mapping vector M to obtain a corresponding first scoring vector B and send it to the first Softmax layer; the first scoring vector B is composed of L first scoring values b i where each of the first scoring values b i is a real number;

[0045] The first Softmax layer is used to perform a five-class probability prediction based on the first scoring vector B to obtain a corresponding first prediction vector U; and use the base marker corresponding to the maximum first prediction probability in each of the first sub-vectors u in the first prediction vector U i as the corresponding second marker z i ; and form the corresponding first decoding sequence Z by sorting all the obtained second markers z i sequentially; the first prediction vector U is composed of L of the first sub-vectors u i where each of the first sub-vectors u i is composed of five first prediction probabilities, and each first prediction probability corresponds to one type of base marker in the five-base set [A, T, C, G, M].

[0046] Preferably, the target source prediction model is used to perform a binary classification prediction based on the first encoding vector Y input to the model and output a corresponding first prediction result R; the first prediction result R is a binary prediction result, consisting of two prediction results: the target source type and the non-target source type;

[0047] The target source prediction model consists of a first extraction module, a second linear layer, and a second Softmax layer; the second linear layer is based on another fully connected network;

[0048] The input end of the first extraction module is connected to the model input end of the target source prediction model, and the output end is connected to the input end of the second linear layer; the output end of the second linear layer is connected to the input end of the second Softmax layer; the output end of the second Softmax layer is connected to the model output end of the target source prediction model;

[0049] The first extraction module is used to extract the first word segmentation feature vector y j=1 at the first position in the first encoded vector Y as the corresponding second extraction vector C and send it to the second linear layer;

[0050] The second linear layer is used to perform a fully connected calculation according to the second extraction vector C to obtain a corresponding second scoring vector D and send it to the second Softmax layer; the second scoring vector D consists of two second scores d1 and d2;

[0051] The second Softmax layer is used to perform binary classification probability prediction according to the second scores d1 and d2 of the second scoring vector D to obtain a corresponding second prediction vector V; and identify the two second prediction probabilities v1 and v2 of the second prediction vector V; if the second prediction probability v1 is larger, set the corresponding first prediction result R as the target source type; if the second prediction probability v2 is larger, set the corresponding first prediction result R as the non-target source type; the second prediction vector V consists of two second prediction probabilities v1 and v2, where the second prediction probability v1 is the true value probability and the second prediction probability v2 is the false value probability.

[0052] Preferably, the first framework training of the pre-trained model framework according to the second base sequence to obtain the corresponding encoding model pre-training parameters specifically includes:

[0053] Step 1001, use the pre-trained model framework as the corresponding current model framework; and use the methylation feature encoding model and the sequence decoding model of the current model framework as the corresponding current encoding model and current decoding model; and set two positive integers as the corresponding first and second stage training times thresholds; and set two proportionality parameters with values between 0 and 1 as the corresponding first and second proportionality parameters; and set the first and second random masking ratios built in the sequence tokenizer of the current encoding model as the corresponding first and second proportionality parameters; and set two counters initialized to 0 as the corresponding first and second counters;

[0054] Wherein, the first proportion parameter < the second proportion parameter, 10% < the first proportion parameter < 20%, 90% < the second proportion parameter < 100%;

[0055] Step 1002: taking the second base sequence as the corresponding first sequence X; performing label vector conversion on the current first sequence X to obtain a corresponding label vector PG; and performing smooth label vector conversion on the label vector PG to obtain a corresponding smooth label vector PS;

[0056] Among them, the label vector PG consists of L sub-vectors pg i Composition; the subvector pg i and the first marker x of the first sequence X i One-to-one correspondence; each of the subvectors pg i are all one-hot encoded vectors of length 5, consisting of 5 bases one-hot encoded pgA i , pgT i , pgC i , pgG i , pgM i Composition; each of the sub-vectors pg i The five base unique-hot codes of pg correspond to the five types of base markers in the five-element base set [A, T, C, G, M] one by one; each of the subvectors pg i Among the five base one-hot codes, only one is 1 and the other four are 0. The base marker corresponding to the base one-hot code of 1 is the same as the current subvector pg i The corresponding first marker x i match;

[0057] The smoothed label vector PS consists of L sub-vectors ps i Composition; the sub-vector ps i The subvector pg of the label vector PG i One-to-one correspondence; each of the sub-vectors ps i The length of the vectors is 5, and psA is encoded by 5 vectors i ,psT i ,psC i , psG i 、psM i Composition; each of the sub-vectors ps i The five vector codes of correspond one by one to the five types of base markers in the five-element base set [A, T, C, G, M] respectively;

[0058] The subvector pg i With the subvector ps iThe conversion relationship is as follows:

[0059]

[0060] α is a preset smoothing parameter with a value between 0 and 1; I is a preset all-1 vector with a length of 5; the first condition is that the sub-vector pg i the corresponding first marker x i is the base marker C or M; the second condition is that the sub-vector pg i the corresponding first marker x i is the base marker A, T, or G;

[0061] Step 1003, set the masking rule type built into the sequence tokenizer of the current encoding model to a type of random masking;

[0062] Step 1004, input the current first sequence X into the sequence tokenizer of the current encoding model for processing to obtain the corresponding first embedding encoding vector E;

[0063] Step 1005, input the first embedding encoding vector E into the MosaicBERT model of the current encoding model for processing to obtain the corresponding first encoding vector Y; and input the first encoding vector Y into the current decoding model for processing to obtain the corresponding first decoding sequence Z; and use each first sub-vector u of the first prediction vector U generated during the current processing of the current decoding model i as a corresponding sub-vector pr i and form the corresponding prediction vector PR from all the obtained sub-vectors pr i ;

[0064] Among them, the prediction vector PR is composed of L sub-vectors pr i ; the sub-vector pr i corresponds one-to-one with the second marker z of the first decoding sequence Z i ; the five first prediction probabilities of the sub-vector pr i are denoted as prA i 、prT i 、prC i 、prG i 、prM i and correspond one-to-one with five types of base markers in the five-base set [A, T, C, G, M] respectively;

[0065] Step 1006, input the prediction vector PR and the smoothing label vector PS into a preset first model loss function L M1 for calculation to obtain the corresponding first loss value;

[0066] Among them, the first model loss function L M1 is:

[0067]

[0068] Step 1007: Identify whether the first loss value meets a preset first loss value range; if it meets, increment the first counter by 1 and go to step 1008; if it does not meet, based on a preset first model optimizer, modulate the full model parameters of the current encoding model and the full model parameters of the current decoding model in the direction of minimizing the first model loss function L M1 to reach the minimum value, and return to step 1005 at the end of this round of modulation;

[0069] Among them, the first model optimizer includes at least an Adam optimizer and an SGD optimizer;

[0070] Step 1008: Identify whether the first counter exceeds the one-stage training times threshold; if it does not exceed, return to step 1004 to continue training; if it exceeds, go to step 1009;

[0071] Step 1009: Set the masking rule type built in the sequence tokenizer of the current encoding model to two-category random masking;

[0072] Step 1010: Input the current first sequence X into the sequence tokenizer of the current encoding model for processing to obtain the corresponding first embedded encoding vector E;

[0073] Step 1011: Input the first embedded encoding vector E into the MosaicBERT model of the current encoding model for processing to obtain the corresponding first encoding vector Y; and input the first encoding vector Y into the current decoding model for processing to obtain the corresponding first decoded sequence Z; and use each first sub-vector u of the first prediction vector U generated during the current processing of the current decoding model i as a corresponding sub-vector pr i and form the corresponding prediction vector PR from all the obtained sub-vectors pr i ;

[0074] Step 1012: Input the prediction vector PR and the smoothed label vector PS into the first model loss function L M1 for calculation to obtain the corresponding second loss value;

[0075] Step 1013: Identify whether the second loss value meets a preset second loss value range; if it meets, increment the second counter by 1 and proceed to Step 1014; if it does not meet, based on the first model optimizer, perform one round of modulation on the full model parameters of the current encoding model and the full model parameters of the current decoding model in the direction of minimizing the first model loss function L M1 until it reaches the minimum value, and at the end of this round of modulation, return to Step 1011;

[0076] Step 1014: Identify whether the second counter exceeds the two-stage training times threshold; if it does not exceed, return to Step 1010 to continue training; if it exceeds, proceed to Step 1015;

[0077] Step 1015: Confirm that the first framework training is completed, and save the current model parameters of the current encoding model as the corresponding pre-training parameters of the encoding model; and solidify the framework model parameters of the pre-training model framework.

[0078] Preferably, constructing the model dataset based on the BS-seq read segment set and the target source type specifically includes:

[0079] Compose a corresponding positive sample set from the BS-seq read segment records in the BS-seq read segment set whose read segment sample sources match the target source type, and compose a corresponding negative sample set from the BS-seq read segment records in the BS-seq read segment set whose read segment sample sources do not match the target source type;

[0080] And use the BS-seq read segments of each BS-seq read segment record in the positive sample set as a corresponding first training sequence, and set a first label classification result specifically set as the target source type; and compose a corresponding first data record from the first training sequence and the first label classification result corresponding to each BS-seq read segment record in the positive sample set;

[0081] And use the BS-seq read segments of each BS-seq read segment record in the negative sample set as a corresponding first training sequence, and set a first label classification result specifically set as the non-target source type; and compose a corresponding first data record from the first training sequence and the first label classification result corresponding to each BS-seq read segment record in the negative sample set;

[0082] And compose the corresponding model dataset from all the obtained first data records.

[0083] Preferably, the second framework training of the first task model framework based on the pre-trained parameters of the encoding model and the model data set specifically includes:

[0084] Step 1201, take the first task model framework as the corresponding current model framework; and take the methylation feature encoding model and the target source prediction model of the current model framework as the corresponding current encoding model and current prediction model; and set the masking rule type built in the sequence tokenizer of the current encoding model to no masking;

[0085] Step 1202, fix the model parameters of the methylation feature encoding model as the corresponding pre-trained parameters of the encoding model; and based on the LoRA fine-tuning mechanism, load a corresponding LoRA adapter on one or more specified modules of the methylation feature encoding model; and form a corresponding adapter parameter set from the adapter parameters of all the LoRA adapters;

[0086] Among them, the fixed model parameters of each specified module are denoted as the pre-fine-tuning parameters W0, and the adapter parameters of the LoRA adapter corresponding to each specified module are composed of the matrix parameters of two low-rank matrices A and B; each LoRA adapter is used to fine-tune the pre-fine-tuning parameters W0 of the corresponding specified module to obtain the corresponding post-fine-tuning parameters W1; the conversion relationship between the post-fine-tuning parameters W1 and the pre-fine-tuning parameters W0 is: W1X in = W0X in + BAX in , X in is the input vector of the specified module, and W0X in is the output vector obtained by processing the input vector X in without adding the corresponding LoRA adapter for the current specified module, and W1X in is the output vector obtained by processing the input vector X in after adding the corresponding LoRA adapter for the current specified module;

[0087] Step 1203, divide the model data set into two sub-data sets according to a preset first splitting ratio and denote them as the corresponding first training set and first evaluation set;

[0088] Among them, both the first training set and the first evaluation set are composed of multiple first data records; the ratio of the total number of records in the first training set to the total number of records in the first evaluation set satisfies the first splitting ratio; the positive and negative sample ratios in the first training set and the first evaluation set are the same;

[0089] Step 1204: Use the first first data record of the first training set as the corresponding current training record;

[0090] Step 1205: Set a corresponding label vector TG based on the first label classification result of the current training record;

[0091] Among them, the label vector TG consists of two one-hot encodings tg1 and tg2; the one-hot encodings tg1 and tg2 correspond to the target source type and the non-target source type respectively; if the first label classification result of the current training record is the target source type, then the one-hot encoding tg1 is 1 and the one-hot encoding tg2 is 0; if the first label classification result of the current training record is the non-target source type, then the one-hot encoding tg1 is 0 and the one-hot encoding tg2 is 1;

[0092] Step 1206: Input the first training sequence of the current training record as the corresponding first sequence X into the current model framework for processing to obtain the corresponding first prediction result R; and set a corresponding prediction vector TR based on the second prediction vector V generated by the current prediction model during this processing;

[0093] Among them, the prediction vector TR consists of two vector encodings tr1 and tr2; the vector encoding tr1 is the second prediction probability v1 of the second prediction vector V, and the vector encoding tr2 is the second prediction probability v2 of the second prediction vector V;

[0094] Step 1207: Input the prediction vector TR and the label vector TG into the preset second model loss function L M2 for calculation to obtain the corresponding third loss value;

[0095] Among them, the second model loss function L M2 is implemented based on the L1 loss function, the L2 loss function or the cross-entropy loss function;

[0096] Step 1208: Identify whether the third loss value meets the preset third loss value range; if the third loss value meets the third loss value range, identify whether the current training record is the last first data record of the first training set. If so, go to Step 1209. If not, use the next first data record of the first training set as the new current training record and return to Step 1205 to continue training; if the third loss value does not meet the third loss value range, based on the preset second model optimizer, make the second model loss function L M2The direction reaching the minimum value performs one round of modulation on the adapter parameter set and the full model parameters of the current prediction model, and returns to step 1206 at the end of this round of modulation;

[0097] Among them, the second model optimizer at least includes an Adam optimizer and an SGD optimizer;

[0098] Step 1209, perform one round of traversal on all the first data records in the first evaluation set; and during this round of traversal, use the currently traversed first data record as the corresponding current evaluation record; and set a corresponding label vector TG based on the first label classification result of the current evaluation record; and use the first training sequence of the current evaluation record as the corresponding first sequence X to be input into the current model framework for processing to obtain the corresponding first prediction result R, and set a corresponding prediction vector TR based on the second prediction vector V generated by the current prediction model during this processing; and form a corresponding prediction-label pair from the label vector TG and the prediction vector TR corresponding to the current evaluation record; and at the end of this round of traversal, bring all the obtained prediction-label pairs into a preset first model evaluation function for calculation to obtain the corresponding first evaluation value; and calculate the accuracy, precision, recall rate, and F1 score based on all the obtained prediction-label pairs to obtain the corresponding first accuracy, first precision, first recall rate, and first F1 score;

[0099] Among them, the first model evaluation function is implemented based on the MSE function or the RMSE function;

[0100] Step 1210, identify whether the first evaluation value, the first accuracy, the first precision, the first recall rate, and the first F1 score all meet the corresponding first evaluation value range, first accuracy range, first precision range, first recall rate range, and first F1 score range; if all meet, go to step 1211; if at least one of the first evaluation value, the first accuracy, the first precision, the first recall rate, and the first F1 score does not meet the corresponding range, return to step 1203 to continue training;

[0101] Step 1211, confirm that the training of the second framework is completed, and save the pre-trained parameters of the encoding model and the latest adapter parameter set as the corresponding fine-tuned encoding model parameter set; and solidify the framework model parameters of the first task model framework.

[0102] Preferably, the methylation sequence reconstruction task of processing DNA sequences based on the pre-trained model framework specifically includes:

[0103] The DNA sequence input by the user is denoted as the corresponding first DNA sequence; and the first DNA sequence is used as the corresponding first sequence X and input into the pre-trained model framework for processing to obtain the corresponding first decoded sequence Z, and the current first decoded sequence Z is used as the corresponding first reconstructed sequence and fed back to the current user; wherein, the first DNA sequence is a sequence of base markers sorted from the base markers in the four-base set [A, T, C, G]; the first reconstructed sequence is a sequence of base markers sorted from the base markers in the five-base set [A, T, C, G, M].

[0104] Preferably, the processing of a DNA sequence source identification task corresponding to the target source type based on the first task model framework specifically includes:

[0105] The DNA sequence input by the user is denoted as the corresponding second DNA sequence; and the second DNA sequence is used as the corresponding first sequence X and input into the first task model framework for processing to obtain the corresponding first prediction result R, and the current first prediction result R is used as the corresponding first identification result and fed back to the current user; wherein, the second DNA sequence is a sequence of base markers sorted from the base markers in the four-base set [A, T, C, G] or the five-base set [A, T, C, G, M].

[0106] A second aspect of the embodiments of the present invention provides an apparatus for implementing the processing method of the methylation feature encoding model described in the first aspect above. The apparatus includes: a preprocessing module, a model construction module, a first training module, a second training module, and a model application module;

[0107] The preprocessing module is used to set a corresponding extended marker for 5-methylcytosine in the DNA sequence, denoted as the base marker M, and adjust the four-base set [A, T, C, G] of DNA to the corresponding five-base set [A, T, C, G, M]; and perform big data collection on the bisulfite sequencing reads of the first species to form a corresponding BS-seq read set; and perform read splicing according to the BS-seq read set to obtain a corresponding first base sequence; and perform five-element sequence conversion on the first base sequence based on the five-base set [A, T, C, G, M] and the methylation state markers corresponding to each base marker C on the first base sequence to obtain a corresponding second base sequence; and perform vocabulary learning based on the BPE algorithm and the second base sequence to obtain a corresponding first vocabulary;

[0108] The model construction module is used to construct a methylation feature encoding model with the MosaicBERT model as the core; store the first vocabulary into the vocabulary storage module of the methylation feature encoding model; construct a sequence decoding model for methylation sequence reconstruction; and construct a corresponding binary classification prediction model based on a type of target source type, denoted as the target source prediction model; the target source type is specifically a type of tissue type or a type of cell type;

[0109] The first training module is used to form a pre-training model framework by the methylation feature encoding model and the sequence decoding model; and perform first-frame training on the pre-training model framework according to the second base sequence to obtain corresponding encoding model pre-training parameters;

[0110] The second training module is used to form a first task model framework by the methylation feature encoding model and the target source prediction model; construct a model data set based on the BS-seq read set and the target source type; and perform second-frame training on the first task model framework based on the encoding model pre-training parameters and the model data set;

[0111] The model application module is used to, after the end of the first and second frame trainings, process the methylation sequence reconstruction task of the DNA sequence based on the pre-training model framework; and process a DNA sequence source identification task corresponding to the target source type based on the first task model framework.

[0112] The third aspect of the embodiments of the present invention provides an electronic device, including: a memory, a processor, and a transceiver;

[0113] The processor is used to be coupled with the memory, read and execute the instructions in the memory to implement the method steps described in the first aspect above;

[0114] The transceiver is coupled with the processor, and the processor controls the transceiver to perform message sending and receiving.

[0115] The fourth aspect of the embodiments of the present invention provides a computer-readable storage medium, which stores computer instructions. When the computer instructions are executed by a computer, the computer is caused to execute the instructions of the method described in the first aspect above.

[0116] An embodiment of the present invention provides a processing method, device, electronic device, and computer-readable storage medium for a methylation feature encoding model. As can be seen from the above, in the embodiment of the present invention, a corresponding base marker M is set for 5-mC in the DNA sequence, and the four-base set [A, T, C, G] of DNA is adjusted to a five-base set [A, T, C, G, M]; and big data collection is performed on the bisulfite sequencing reads of a certain species, and based on the collected set of BS-seq reads, operations such as read splicing and methylation status marking are performed to obtain a base sequence, and vocabulary learning is performed based on the BPE algorithm and this base sequence; then, an encoder model denoted as a methylation feature encoding model is constructed with the MosaicBERT model as the core, the learned vocabulary is stored in the vocabulary storage module inside the encoding model, a decoder model denoted as a sequence decoding model for reconstructing the methylation sequence is constructed, and a corresponding binary classification prediction model denoted as a target source prediction model is constructed based on a type of target source type; then, a pre-training model framework is composed of the methylation feature encoding model and the sequence decoding model, and a first task model framework is composed of the methylation feature encoding model and the target source prediction model; then, first, the pre-training model framework is trained based on the above base sequence, and then the first task model framework is trained; finally, the methylation sequence reconstruction task of the DNA sequence is processed based on the pre-training model framework, and a type of DNA sequence source identification task is processed based on the first task model framework. The methylation feature encoding model given in the embodiment of the present invention can learn the methylation features of the DNA sequence. The pre-training model framework with this methylation feature encoding model as the core encoder can be used to process the methylation sequence analysis task, and the first task model framework with this methylation feature encoding model as the core encoder can be used to process a type of target source identification task; based on the methylation feature encoding model given in the embodiment of the present invention and the corresponding two task model frameworks to process the methylation sequence analysis task and the target source identification task, not only simplifies the task processing steps, shortens the task processing cycle, improves the real-time performance of the task processing, but also effectively reduces the task processing cost. BRIEF DESCRIPTION OF THE DRAWINGS

[0117] Figure 1 It is a schematic diagram of a processing method for a methylation feature encoding model provided in Embodiment 1 of the present invention;

[0118] Figure 2 It is a module structure diagram of the pre-training model framework, the first task model framework, the methylation feature encoding model, the sequence decoding model, and the target source prediction model provided in Embodiment 1 of the present invention;

[0119] Figure 3 It is a module structure diagram of a processing device for a methylation feature encoding model provided in Embodiment 2 of the present invention;

[0120] Figure 4 This is a schematic structural diagram of an electronic device provided in Embodiment 3 of the present invention. Specific implementation manners

[0121] In order to make the objectives, technical solutions and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0122] Embodiment 1 of the present invention provides a processing method for a methylation feature coding model, as Figure 1 shown in the schematic diagram of the processing method for a methylation feature coding model provided in Embodiment 1 of the present invention, the method mainly includes the following steps:

[0123] Step 1, set a corresponding extended marker for 5-methylcytosine in the DNA sequence, denoted as the base marker M, and adjust the four-base set [A, T, C, G] of DNA to the corresponding five-base set [A, T, C, G, M]; perform big data collection on the bisulfite sequencing reads of the first species to form a corresponding BS-seq read set; perform read splicing based on the BS-seq read set to obtain a corresponding first base sequence; perform five-element sequence conversion on the first base sequence based on the five-base set [A, T, C, G, M] and the methylation status markers corresponding to each base marker C on the first base sequence to obtain a corresponding second base sequence; perform vocabulary learning based on the BPE algorithm and the second base sequence to obtain a corresponding first vocabulary;

[0124] Specifically including: Step 11, set a corresponding extended marker for 5-methylcytosine in the DNA sequence, denoted as the base marker M, and adjust the four-base set [A, T, C, G] of DNA to the corresponding five-base set [A, T, C, G, M];

[0125] Among them, the base markers A, T, C, and G are the markers for adenine, thymine, cytosine, and guanine bases respectively;

[0126] Step 12, perform big data collection on the bisulfite sequencing reads of the first species to form a corresponding BS-seq read set;

[0127] Specifically including: Step 121, perform data collection on the bisulfite sequencing reads of the DNA sequence of the first species through a variety of preset data collection channels to obtain a plurality of BS-seq reads;

[0128] Here, the first species in the embodiments of the present invention is defaulted to be human, and it can also be one of other multiple types of species except humans, for example, monkeys, orangutans, etc.; the multiple preset data collections at least include one or more publicly available BS-seq read datasets, and publicly available technical documents with BS-seq read information of the first species;

[0129] Step 122, and set the type range of the corresponding source type based on the sample source information corresponding to all the collected BS-seq reads; and set a corresponding source type code for each source type in the type range of the source type; and use the source type code corresponding to each BS-seq read as the corresponding read segment sample source; and form a corresponding BS-seq read record from each BS-seq read and the corresponding read segment sample source; and perform error read segment deletion and duplicate read segment deduplication processing on all the obtained BS-seq read records;

[0130] Step 123, and form a BS-seq read set from all the remaining BS-seq read records;

[0131] Here, the BS-seq read set in the embodiments of the present invention is composed of multiple BS-seq read records; each BS-seq read record at least includes one BS-seq read and a corresponding read segment sample source; each BS-seq read is a DNA sequence obtained by DNA sequencing after bisulfite sequencing of a DNA sample of the first species, and is sorted by the base markers A, T, C, G; each base marker C on each BS-seq read is associated with a corresponding methylation status marker; the methylation status marker is a binary status marker, including two states: unmethylated and methylated; the read segment sample source is the source type code corresponding to a current BS-seq read, and each source type code corresponds to a specific source type; the type range of the source type is composed of multiple tissue types and / or multiple cell types of the first species;

[0132] Step 13, and perform read segment splicing according to the BS-seq read set to obtain a corresponding first base sequence;

[0133] Specifically, it includes: sequentially splicing all the BS-seq reads in the BS-seq read set to obtain a corresponding first base sequence;

[0134] Here, the first base sequence in the embodiments of the present invention is a marker sequence sorted by the base markers A, T, C, G;

[0135] Step 14, and based on the five-base set [A, T, C, G, M] and the methylation status markers corresponding to the base markers C on the first base sequence, perform five-element sequence conversion on the first base sequence to obtain the corresponding second base sequence;

[0136] Specifically, it includes: performing sequence replication on the first base sequence to obtain the corresponding first replication sequence; modifying the base marker C with the status value of methylated for each current methylation status marker on the first replication sequence to the corresponding base marker M; and using the first replication sequence with the marker modification completed as the corresponding second base sequence;

[0137] Here, the second base sequence in the embodiment of the present invention is a marker sequence sorted by the base markers A, T, C, G, M;

[0138] Step 15, and based on the BPE (Byte Pair Encoding) algorithm and the second base sequence, perform vocabulary learning to obtain the corresponding first vocabulary;

[0139] Here, the first vocabulary in the embodiment of the present invention consists of multiple first sub-word records; the first sub-word record includes the first sub-word encoding, the first sub-word text, and the first sub-word frequency; the first sub-word text is a string sorted by one or more base markers in the five-base set [A, T, C, G, M]; the first sub-word frequency is an integer value; all the first sub-word texts in the first vocabulary are different;

[0140] Specifically, it includes: Step 151, set an empty vocabulary as the corresponding first vocabulary; initialize the first vocabulary; set the record total threshold for the first vocabulary; and use the second base sequence as a corresponding first string;

[0141] Here, the initialized first vocabulary in the embodiment of the present invention consists of six initial first sub-word records; the first sub-word texts of these six initial first sub-word records are the character 'A', the character 'T', the character 'C', the character 'G', the character 'M', and the string 'CG' in sequence, and the first sub-word frequencies of these six initial first sub-word records are all initialized to 1;

[0142] Step 152, use each first sub-word text in the first vocabulary as a corresponding basic sub-word, and form the corresponding current sub-word set from all the obtained basic sub-words; and according to the sub-word sequence splitting method of the BPE algorithm, perform sub-word sequence splitting on the first string according to the current sub-word set to obtain the corresponding current sub-word sequence;

[0143] Here, the current sub-word sequence is sorted by multiple basic sub-words in sequence; the basic sub-word pair is sorted by the front and back two basic sub-words;

[0144] For example, the initialized first vocabulary contains 6 sub-word texts with non-zero word frequencies, namely 'A', 'T', 'C', 'G', 'M', 'CG'; assuming the first string is 'TACCGTAACGCTGC', then the current sub-word sequence obtained by the sub-word sequence splitting method of the dark BPE algorithm is [T, A, C, CG, T, A, A, CG, C, T, G, C];

[0145] Step 153: According to the sub-word pair combination method of the BPE algorithm, every two adjacent basic sub-words in the current sub-word sequence form a corresponding basic sub-word pair; and cluster the same basic sub-word pairs to obtain multiple clustered word pair sets; and take the total number of basic sub-word pairs in each clustered word pair set as the corresponding first word pair count, and take the largest first word pair count as the corresponding high-frequency word pair count; and identify the high-frequency word pair count; if the high-frequency word pair count is 1, go to step 158; if the high-frequency word pair count is greater than 1, take the total number of sets of the clustered word pair set corresponding to the high-frequency word pair count as the corresponding high-frequency word pair total, and identify the high-frequency word pair total. If the high-frequency word pair total is 1, go to step 154; if the high-frequency word pair total is greater than 1, go to step 155;

[0146] Here, each clustered word pair set consists of one or more basic sub-word pairs, and all the basic sub-word pairs in each clustered word pair set are the same;

[0147] For example, if the current sub-word sequence is [T, A, C, CG, T, A, A, CG, C, T, G, C], then the basic sub-word pairs obtained by the sub-word pair combination method of the BPE algorithm are [TA, AC, CCG, CGT, TA, AA, ACG, CGC, CT, TG, GC]; thus, the corresponding relationship of [basic sub-word pair - first word pair count] is [TA - 2], [AC - 1], [CCG - 1], [CGT - 1], [AA - 1], [ACG - 1], [CGC - 1], [CT - 1], [TG - 1], [GC - 1]; where the high-frequency word pair count = 2 corresponds to the basic sub-word pair 'TA', and the high-frequency word pair total is 1;

[0148] It should be noted that two exit conditions will be set when the present invention embodiment performs vocabulary learning: 1) Exit the vocabulary learning process when there are no duplicate word pairs among all the basic sub-word pairs obtained from the current sub-word sequence, that is, when the first word pair count of all basic sub-word pairs is 1. At this time, the high-frequency word pair count = 1, so in the current step 153, it will go to step 158 when the high-frequency word pair count = 1; 2) Exit the vocabulary learning process when the vocabulary of the vocabulary table reaches the upper limit, that is, when the total number of the first sub-word records with non-zero word frequencies in the first vocabulary exceeds the record total threshold, and it can be known from the following that this will be judged at step 157;

[0149] Step 154: Use the basic sub-word pair corresponding to the only clustering word pair set corresponding to the total number of high-frequency word pairs as the corresponding final selected sub-word pair; and go to Step 156;

[0150] For example, if the high-frequency word pair count obtained in Step 153 = 2 corresponds to the basic sub-word pair 'TA', and the total number of high-frequency word pairs is 1, then the final selected sub-word pair should be 'TA';

[0151] Step 155: Denote all the clustering word pair sets corresponding to the total number of high-frequency word pairs as the corresponding candidate sets; and use the basic sub-word pairs corresponding to each candidate set as the corresponding candidate sub-word pairs; and denote the character 'C', the character 'M', and the string 'CG' as methylation-related characters; and denote each candidate sub-word pair containing methylation-related characters as a preferred sub-word pair; and identify the total number of preferred sub-word pairs; if the total number of preferred sub-word pairs is 0, then randomly select one from all the candidate sub-word pairs as the corresponding final selected sub-word pair; if the total number of preferred sub-word pairs is greater than 0, then randomly select one from all the preferred sub-word pairs as the corresponding final selected sub-word pair;

[0152] Here, if it is identified through Step 153 that the total number of high-frequency word pairs > 1, and the first word pair counts of two or more basic sub-word pairs are all equal to the high-frequency word pair count; then, the embodiments of the present invention regard these basic sub-word pairs with the first word pair count equal to the high-frequency word pair count as candidate sub-word pairs, and by default, regard the candidate sub-word pairs containing methylation-related characters (C, M, CG) among them as preferred sub-word pairs and randomly select one from them as the final selected sub-word pair. If there is no basic sub-word pair containing methylation-related characters (C, M, CG) among all the candidate sub-word pairs, then randomly select one from all the candidate sub-word pairs as the final selected sub-word pair;

[0153] Step 156: Add a first sub-word record to the first vocabulary as the corresponding current newly added record, and set the first sub-word text of the current newly added record to the corresponding final selected sub-word pair; and after the record setting is completed, use each first sub-word text in the current first vocabulary as a new basic sub-word, and form the latest current sub-word set from all the newly obtained basic sub-words; and according to the sub-word sequence splitting method of the BPE algorithm, split the first string according to the current sub-word set to obtain the latest current sub-word sequence; and count the total number of each basic sub-word appearing in the current sub-word sequence and use the statistical result as the corresponding first word frequency, and set the first word frequency corresponding to each basic sub-word that does not appear in the current sub-word sequence in the current sub-word set to 0; and update the first word frequency corresponding to each first sub-word text in the current first vocabulary to the corresponding first word frequency; and after the update is completed, delete the first sub-word record with the first word frequency of 0 in the first vocabulary;

[0154] Step 157, count the total number of first sub-word records in the current first vocabulary to obtain the corresponding current record total; and identify whether the current record total is less than the record total threshold; if so, return to Step 153; if not, go to Step 158;

[0155] Step 158, re-sort all the first sub-word records in the latest first vocabulary in descending order of the first sub-word frequency; and based on the preset sub-word encoding rule, set the encoding for the first sub-word codes corresponding to each first sub-word text in the sorted first vocabulary; and output the first vocabulary with the sub-word encoding set as the vocabulary learning result of this time.

[0156] Here, the sub-word encoding rule of the embodiment of the present invention can be customized according to application requirements; conventionally, it can be set with sequentially increasing integers.

[0157] Step 2, construct a methylation feature encoding model with the MosaicBERT model as the core; store the first vocabulary in the word table storage module of the methylation feature encoding model; construct a sequence decoding model for methylation sequence reconstruction; and construct a corresponding binary classification prediction model based on a type of target source type, denoted as the target source prediction model.

[0158] Here, the target source type of the embodiment of the present invention is specifically a type of tissue type or a type of cell type.

[0159] The methylation feature encoding model of the embodiment of the present invention is used to perform feature encoding on the first sequence X input to the model and output the corresponding first encoding vector Y; wherein, the first sequence X is composed of multiple first tokens x i sorted in order; each first token x i is a type of base token in the five-base set [A, T, C, G, M]; 1 ≤ token index i ≤ L, where L is the sequence length of the first sequence X; the first encoding vector Y is composed of multiple first tokenization feature vectors y j ; 1 ≤ tokenization index j ≤ W, where W is the sequence length of the first tokenization sequence S corresponding to the first sequence X.

[0160] As Figure 2 shown, the methylation feature encoding model of the embodiment of the present invention is composed of a sequence tokenizer, a word table storage module, and the MosaicBERT model.

[0161] The connection relationships of the components of the methylation feature encoding model are as follows: the input end of the sequence tokenizer is connected to the model input end of the methylation feature encoding model, and the output end is connected to the input end of the MosaicBERT model; the sequence tokenizer is also connected to the vocabulary storage module; the output end of the MosaicBERT model is connected to the model output end of the methylation feature encoding model.

[0162] The functions of the components of the methylation feature encoding model are as follows:

[0163] 1) The sequence tokenizer is used to split the first sequence X into a first sub-word sequence composed of multiple first sub-word texts according to the sub-word sequence splitting method of the BPE algorithm, based on the first vocabulary stored in the vocabulary storage module; and add a preset classification sub-word text 'CLS' before the first first sub-word text in the first sub-word sequence; and take the sequence length of the first sub-word sequence after the addition as the corresponding sequence length W; and take each sub-word text of the first sub-word sequence as a corresponding first token s j , and from all the first tokens s j to form the corresponding first token sequence S; and based on the first sub-word encoding corresponding to each first sub-word text in the first vocabulary and the classification sub-word encoding corresponding to the classification sub-word text 'CLS', set the corresponding encoding of each first token s j in the first token sequence S to obtain the corresponding first token encoding, and form the corresponding first token encoding sequence from all the obtained first token encodings; and identify the built-in masking rule type; if the masking rule type is no masking, no masking modification is made to the first token encoding sequence; if the masking rule type is a type of random masking, randomly mask the first token encodings in the first token encoding sequence according to the built-in type of random masking ratio and the preset token masking encoding, and ensure that the ratio of the total number of masked first token encodings to the sequence length of the first token encoding sequence matches the type of random masking ratio; if the masking rule type is a second type of random masking, mark the first tokens s j containing methylation-related characters in the first token sequence S as methylation-related tokens, and mark the first token encodings corresponding to each methylation-related token in the first token encoding sequence as methylation-related encodings, and randomly mask the methylation-related encodings in the first token encoding sequence according to the built-in second type of random masking ratio and the token masking encoding, and ensure that the ratio of the total number of masked methylation-related encodings to the total number of methylation-related encodings in the first token encoding sequence matches the second type of random masking ratio; and based on the model input vector embedding encoding rule of the MosaicBERT model, perform embedding encoding processing on the first token encoding sequence to obtain the corresponding first embedding encoding vector E and send it to the MosaicBERT model;

[0164] Here, the first token sequence S of the embodiments of the present invention consists of W first tokens s j The first first token s j=1 is the classification sub-word text 'CLS', and each of the first tokens s 2≤j≤W from the second to the last first token s j corresponds to a first sub-word text in the first vocabulary; the first embedding coding vector E consists of W first token embedding coding vectors e j The first token embedding coding vector e j corresponds one-to-one with the first token s j ;

[0165] The sequence tokenizer has three built-in parameters: the masking rule type, the first type of random masking ratio, and the second type of random masking ratio; the masking rule type includes three types: no masking, the first type of random masking, and the second type of random masking; 1) When the masking rule type is set to no masking, it means that there is no need to perform masking processing on the first token coding sequence; 2) When the masking rule type is set to the first type of random masking, it means that it is necessary to perform random masking processing on the first token coding sequence according to the first type of random masking ratio. For example, if the first token coding sequence contains 100 first token codings and the first type of random masking ratio is set to 15%, it means that 100×15% = 15 first token codings in the first token coding sequence need to be masked. Specifically, 15 first token codings are randomly selected from the first token coding sequence, and the coding values of these 15 first token codings are reset to the token masking coding. Here, the token masking coding is a pre-set coding value; 3) When the masking rule type is set to the second type of random masking, it means that it is necessary to perform random masking processing on the methylation-related codings of the first token coding sequence according to the second type of random masking ratio. For example, if the total number of methylation-related codings in the first token coding sequence is 100 and the second type of random masking ratio is set to 98%, it means that 100×98% = 98 methylation-related codings among these 100 methylation-related codings need to be masked. Specifically, 98 methylation-related codings are randomly selected from these 100 methylation-related codings, and the coding values of these 98 methylation-related codings are reset to the token masking coding.

[0166] 2) The vocabulary storage module is used to store the first vocabulary.

[0167] 3) The MosaicBERT model is used to perform feature coding processing according to the first embedding coding vector E and output the corresponding first coding vector Y.

[0168] It should be noted that the model structure and component functions of the MosaicBERT model have been disclosed in the paper "MosaicBERT: A Bidirectional Encoder Optimized for Fast Pretraining". The MosaicBERT model can be refined according to the content disclosed in this paper. The reason why the embodiments of the present invention select the MosaicBERT model as the core component of the methylation feature encoding model is that MosaicBERT has the following advantages compared with the conventional BERT model: 1) Use the ALiBi positional encoding mechanism for positional encoding. Specifically, it replaces the traditional positional embedding encoding with the Attention with Linear Biases encoding method. This encoding mechanism supports variable-length sequence input and improves the long-sequence modeling ability; 2) Use FlashAttention to optimize the attention calculation process and improve the training efficiency; 3) Use the GeGLU activation function to replace the conventional GLU activation function, enhancing the non-linear representation ability of the model.

[0169] The sequence decoding model of the embodiments of the present invention is used to reconstruct the methylation sequence according to the first encoded vector Y input by the model and output the corresponding first decoded sequence Z; wherein, the first decoded sequence Z is composed of L second tokens z i sequentially sorted; each second token z i is a base token of a five-base set [A, T, C, G, M]; the second token z i corresponds one-to-one with the first token x i one by one.

[0170] As Figure 2 shown, the sequence decoding model of the embodiments of the present invention is composed of a first mapping layer, a first linear layer, and a first Softmax layer; the first mapping layer is implemented based on a fully connected network; the first linear layer is based on another fully connected network.

[0171] The connection relationship of the components of the sequence decoding model is: the input end of the first mapping layer is connected to the model input end of the sequence decoding model, and the output end is connected to the input end of the first linear layer; the output end of the first linear layer is connected to the input end of the first Softmax layer; the output end of the first Softmax layer is connected to the model output end of the sequence decoding model.

[0172] The functions of the components of the sequence decoding model are as follows:

[0173] 1) The first mapping layer is used to map the first tokenization feature vectors y from the 2nd to the Wth in the first encoded vector Y 2≤j≤WExtract and form a first extraction vector A with a length of W - 1; and convert the first extraction vector A into a first mapping vector M with a length of L through a linear transformation and send it to the first linear layer;

[0174] Among them, the first extraction vector A consists of W - 1 first extraction sub - vectors a k which are composed, where 1 ≤ vector index k ≤ (W - 1), and each first extraction sub - vector a k matches the corresponding first tokenization feature vector y j=k+1 ; the first mapping vector M consists of L first sub - mapping vectors m i which are composed;

[0175] 2) The first linear layer is used to perform a fully - connected calculation according to the first mapping vector M to obtain a corresponding first scoring vector B and send it to the first Softmax layer;

[0176] Among them, the first scoring vector B consists of L first scoring values b i which are composed, and each first scoring value b i is a real number;

[0177] 3) The first Softmax layer is used to perform five - classification probability prediction according to the first scoring vector B to obtain a corresponding first prediction vector U; and use the base label corresponding to the largest first prediction probability in each first sub - vector u i in the first prediction vector U as the corresponding second label z i ; and form a corresponding first decoding sequence Z by sorting all the obtained second labels z i in order;

[0178] Among them, the first prediction vector U consists of L first sub - vectors u i which are composed, and each first sub - vector u i is composed of five first prediction probabilities, and each first prediction probability corresponds to a type of base label in the five - element base set [A, T, C, G, M].

[0179] The target source prediction model of the embodiment of the present invention is used to perform binary classification prediction according to the first encoding vector Y input to the model and output a corresponding first prediction result R; among them, the first prediction result R is a binary prediction result, which consists of two prediction results: the target source type and the non - target source type.

[0180] As Figure 2 shown, the target source prediction model of the embodiment of the present invention consists of a first extraction module, a second linear layer, and a second Softmax layer; the second linear layer is based on another fully - connected network.

[0181] The connection relationships of the components of the target source prediction model are as follows: the input end of the first extraction module is connected to the model input end of the target source prediction model, and the output end is connected to the input end of the second linear layer; the output end of the second linear layer is connected to the input end of the second Softmax layer; the output end of the second Softmax layer is connected to the model output end of the target source prediction model.

[0182] The functions of the components of the target source prediction model are as follows:

[0183] 1) The first extraction module is used to extract the first word segmentation feature vector y j=1 in the first encoded vector Y as the corresponding second extraction vector C and send it to the second linear layer;

[0184] 2) The second linear layer is used to perform a fully connected calculation based on the second extraction vector C to obtain a corresponding second scoring vector D and send it to the second Softmax layer;

[0185] wherein, the second scoring vector D consists of two second scores d1 and d2;

[0186] 3) The second Softmax layer is used to perform binary classification probability prediction based on the second scores d1 and d2 of the second scoring vector D to obtain the corresponding second prediction vector V; and identify the two second prediction probabilities v1 and v2 of the second prediction vector V; if the second prediction probability v1 is relatively large, set the corresponding first prediction result R as the target source type; if the second prediction probability v2 is relatively large, set the corresponding first prediction result R as the non-target source type.

[0187] wherein, the second prediction vector V consists of two second prediction probabilities v1 and v2, the second prediction probability v1 is the true value probability, and the second prediction probability v2 is the false value probability.

[0188] Step 3: A pre-training model framework is composed of a methylation feature encoding model and a sequence decoding model; and the pre-training model framework is subjected to first framework training according to the second base sequence to obtain the corresponding pre-training parameters of the encoding model.

[0189] Specifically, it includes: Step 31, a pre-training model framework is composed of a methylation feature encoding model and a sequence decoding model;

[0190] Here, the model connection relationship of the pre-training model framework is as Figure 2 shown;

[0191] Step 32, and the pre-training model framework is subjected to first framework training according to the second base sequence to obtain the corresponding pre-training parameters of the encoding model.

[0192] Specifically, it includes: Step 3201, taking the pre-trained model framework as the corresponding current model framework; taking the methylation feature encoding model and sequence decoding model of the current model framework as the corresponding current encoding model and current decoding model; setting two positive integers as the corresponding first and second stage training times thresholds; setting two proportionality parameters with values between 0 and 1 as the corresponding first and second proportionality parameters; setting the first and second random masking ratios built in the sequence tokenizer of the current encoding model as the corresponding first and second proportionality parameters; setting two counters initialized to 0 as the corresponding first and second counters;

[0193] Here, the first proportionality parameter of the embodiment of the present invention < the second proportionality parameter, 10% < the first proportionality parameter < 20%, 90% < the second proportionality parameter < 100%;

[0194] Step 3202, taking the second base sequence as the corresponding first sequence X; performing label vector conversion on the current first sequence X to obtain the corresponding label vector PG; performing smoothed label vector conversion on the label vector PG to obtain the corresponding smoothed label vector PS;

[0195] Here, the label vector PG of the embodiment of the present invention consists of L sub-vectors pg i ; the sub-vector pg i corresponds one-to-one with the first marker x i of the first sequence X; each sub-vector pg i is a one-hot encoding vector with a length of 5, consisting of 5 base one-hot encodings pgA i , pgT i , pgC i , pgG i , pgM i ; the 5 base one-hot encodings of each sub-vector pg i correspond one-to-one with the five types of base markers in the five-base set [A, T, C, G, M]; only one of the 5 base one-hot encodings of each sub-vector pg i is 1, and the other 4 are 0, and the base marker corresponding to the base one-hot encoding that is 1 matches the first marker x i corresponding to the current sub-vector pg i ;

[0196] The smoothed label vector PS of the embodiment of the present invention consists of L sub-vectors ps i ; the sub-vector ps i corresponds one-to-one with the sub-vector pg i of the label vector PG; the vector length of each sub-vector ps i is 5, consisting of 5 vector encodings psA i , psT i, psC i , psG i , psM i ; Each sub-vector ps i The five vector encodings of are respectively in one-to-one correspondence with the five types of base markers in the five-base set [A, T, C, G, M];

[0197] The sub-vector pg of the embodiment of the present invention i and the sub-vector ps i The conversion relationship is:

[0198]

[0199] where α is a preset smoothing parameter, and its value is between 0 and 1; I is a preset all-1 vector with a length of 5; Condition 1 is: the first marker x i corresponding to the sub-vector pg i is the base marker C or M; Condition 2 is: the first marker x i corresponding to the sub-vector pg i is the base marker A, T or G;

[0200] Step 3203, set the masking rule type built in the sequence tokenizer of the current encoding model to a type of random masking;

[0201] Step 3204, input the current first sequence X into the sequence tokenizer of the current encoding model for processing to obtain the corresponding first embedding encoding vector E;

[0202] Step 3205, input the first embedding encoding vector E into the MosaicBERT model of the current encoding model for processing to obtain the corresponding first encoding vector Y; and input the first encoding vector Y into the current decoding model for processing to obtain the corresponding first decoding sequence Z; and input each first sub-vector u of the first prediction vector U generated during the current processing of the current decoding model i as a corresponding sub-vector pr i , and all the obtained sub-vectors pr i constitute the corresponding prediction vector PR;

[0203] Here, the prediction vector PR of the embodiment of the present invention is composed of L sub-vectors pr i ; The sub-vector pr i is in one-to-one correspondence with the second marker z i of the first decoding sequence Z; The five first prediction probabilities of the sub-vector pr i are denoted as prA i , prT i , prC i , prG i , prM i, corresponding one-to-one with five types of base markers in the five-base set [A, T, C, G, M];

[0204] Step 3206, bring the prediction vector PR and the smoothed label vector PS into the preset first model loss function L M1 for calculation to obtain the corresponding first loss value;

[0205] Among them, the first model loss function L M1 is:

[0206]

[0207] Step 3207, identify whether the first loss value meets the preset first loss value range; if it meets, increment the first counter by 1 and go to Step 3208; if it does not meet, based on the preset first model optimizer, make the first model loss function L M1 reach the minimum value direction to modulate the full model parameters of the current encoding model and the full model parameters of the current decoding model for one round, and return to Step 3205 at the end of this round of modulation;

[0208] Here, the first loss value range of the embodiment of the present invention is a preset loss value range; the first model optimizer includes at least Adam optimizer and SGD optimizer;

[0209] Step 3208, identify whether the first counter exceeds the one-stage training times threshold; if it does not exceed, return to Step 3204 to continue training; if it exceeds, go to Step 3209;

[0210] Step 3209, set the masking rule type built in the sequence tokenizer of the current encoding model to two-category random masking;

[0211] Step 3210, input the current first sequence X into the sequence tokenizer of the current encoding model for processing to obtain the corresponding first embedded encoding vector E;

[0212] Step 3211, input the first embedded encoding vector E into the MosaicBERT model of the current encoding model for processing to obtain the corresponding first encoding vector Y; and input the first encoding vector Y into the current decoding model for processing to obtain the corresponding first decoding sequence Z; and use each first sub-vector u of the first prediction vector U generated during the current processing of the current decoding model i as a corresponding sub-vector pr i and form the corresponding prediction vector PR from all the obtained sub-vectors pr i ;

[0213] Step 3212, bring the prediction vector PR and the smoothed label vector PS into the first model loss function LM1 Perform calculations to obtain the corresponding second loss value;

[0214] Step 3213, identify whether the second loss value meets a preset second loss value range; if it meets, increment the second counter by 1 and go to step 3214; if it does not meet, based on the first model optimizer, modulate the full model parameters of the current encoding model and the full model parameters of the current decoding model in the direction of minimizing the first model loss function L M1 Reach the minimum value direction and perform one round of modulation on the full model parameters of the current encoding model and the full model parameters of the current decoding model, and return to step 3211 at the end of this round of modulation;

[0215] Here, the second loss value range in the embodiments of the present invention is a preset loss value range;

[0216] Step 3214, identify whether the second counter exceeds the two-stage training times threshold; if it does not exceed, return to step 3210 to continue training; if it exceeds, go to step 3215;

[0217] Step 3215, confirm the end of the first framework training, and save the current model parameters of the current encoding model as the corresponding pre-trained parameters of the encoding model; and solidify the framework model parameters of the pre-trained model framework.

[0218] Step 4, compose a first task model framework from a methylation feature encoding model and a target source prediction model; and construct a model data set based on the BS-seq read segment set and the target source type; and perform second framework training on the first task model framework based on the pre-trained parameters of the encoding model and the model data set;

[0219] Specifically include: Step 41, compose a first task model framework from a methylation feature encoding model and a target source prediction model;

[0220] Here, the model connection relationship of the first task model framework is as Figure 2 shown;

[0221] Step 42, and construct a model data set based on the BS-seq read segment set and the target source type;

[0222] Among them, the model data set includes multiple first data records; the first data record includes a first training sequence and a first label classification result; the first label classification result includes a target source type and a non-target source type;

[0223] Specifically include: Step 421, compose a corresponding positive sample set from the BS-seq read segment records in the BS-seq read segment set where the read segment sample source matches the target source type, and compose a corresponding negative sample set from the BS-seq read segment records in the BS-seq read segment set where the read segment sample source does not match the target source type;

[0224] Step 422, and use the BS-seq reads recorded in each BS-seq read segment record of the positive sample set as a corresponding first training sequence, and set a first label classification result specifically set as the target source type; and form a corresponding first data record from the first training sequence and the first label classification result corresponding to each BS-seq read segment record of the positive sample set;

[0225] Step 423, and use the BS-seq reads recorded in each BS-seq read segment record of the negative sample set as a corresponding first training sequence, and set a first label classification result specifically set as the non-target source type; and form a corresponding first data record from the first training sequence and the first label classification result corresponding to each BS-seq read segment record of the negative sample set;

[0226] Step 424, and form a corresponding model data set from all the obtained first data records;

[0227] Step 43, and perform second-frame training on the first task model framework based on the pre-trained parameters of the encoding model and the model data set;

[0228] Specifically, it includes: Step 4301, use the first task model framework as the corresponding current model framework; and use the methylation feature encoding model and the target source prediction model of the current model framework as the corresponding current encoding model and current prediction model; and set the masking rule type built into the sequence tokenizer of the current encoding model to no masking;

[0229] Step 4302, fix the model parameters of the methylation feature encoding model as the corresponding pre-trained parameters of the encoding model; and based on the LoRA fine-tuning mechanism, load a corresponding LoRA adapter on one or more specified modules of the methylation feature encoding model; and form a corresponding adapter parameter set from the adapter parameters of all LoRA adapters;

[0230] Here, the fixed model parameters of each specified module of the methylation feature encoding model in the embodiment of the present invention are denoted as the pre-fine-tuning parameters W0, and the adapter parameters of the LoRA adapter corresponding to each specified module are composed of the matrix parameters of two low-rank matrices A and B; each LoRA adapter is used to fine-tune the pre-fine-tuning parameters W0 of the corresponding specified module to obtain the corresponding fine-tuned parameters W1; the conversion relationship between the fine-tuned parameters W1 and the pre-fine-tuning parameters W0 is: W1X in =W0X in +BAX in ,X in is the input vector of the specified module, and W0X in is the output of the current specified module without adding the corresponding LoRA adapter based on the input vector Xin The output vector obtained by processing, W1X in is the output vector obtained by processing the input vector X based on the current specified module after adding the corresponding LoRA adapter in ;

[0231] It should be noted that, under normal circumstances, some or all of the self-attention modules in the methylation feature encoding model are used as the specified modules;

[0232] Step 4303: Divide the model data set into two sub-data sets according to a preset first splitting ratio, denoted as the corresponding first training set and first evaluation set;

[0233] Here, the first splitting ratio is a preset ratio parameter, such as 8:2; both the first training set and the first evaluation set are composed of multiple first data records; the ratio of the total number of records in the first training set to the total number of records in the first evaluation set satisfies the first splitting ratio; the positive and negative sample ratios in the first training set and the first evaluation set are the same;

[0234] Step 4304: Take the first first data record in the first training set as the corresponding current training record;

[0235] Step 4305: Set a corresponding label vector TG based on the first label classification result of the current training record;

[0236] Here, the label vector TG in the embodiment of the present invention is composed of two one-hot encodings tg1 and tg2; the one-hot encodings tg1 and tg2 correspond to the target source type and the non-target source type respectively; if the first label classification result of the current training record is the target source type, the one-hot encoding tg1 is 1 and the one-hot encoding tg2 is 0; if the first label classification result of the current training record is the non-target source type, the one-hot encoding tg1 is 0 and the one-hot encoding tg2 is 1;

[0237] Step 4306: Take the first training sequence of the current training record as the corresponding first sequence X and input it into the current model framework for processing to obtain the corresponding first prediction result R; and set a corresponding prediction vector TR based on the second prediction vector V generated by the current prediction model during this processing;

[0238] Here, the prediction vector TR in the embodiment of the present invention is composed of two vector encodings tr1 and tr2; the vector encoding tr1 is the second prediction probability v1 of the second prediction vector V, and the vector encoding tr2 is the second prediction probability v2 of the second prediction vector V;

[0239] Step 4307: Input the prediction vector TR and the label vector TG into the preset second model loss function L M2 for calculation to obtain the corresponding third loss value;

[0240] Here, the second model loss function L of the embodiment of the present invention M2 is implemented based on the L1 loss function, the L2 loss function or the cross-entropy loss function;

[0241] Step 4308, identify whether the third loss value meets a preset third loss value range; if the third loss value meets the third loss value range, identify whether the current training record is the last first data record of the first training set. If so, go to step 4309. If not, use the next first data record of the first training set as the new current training record and return to step 4305 to continue training; if the third loss value does not meet the third loss value range, based on a preset second model optimizer, modulate the adapter parameter set and the full model parameters of the current prediction model in the direction of minimizing the second model loss function L M2 for one round, and return to step 4306 at the end of this round of modulation;

[0242] Here, the third loss value range of the embodiment of the present invention is a preset loss value range; the second model optimizer includes at least the Adam optimizer and the SGD optimizer;

[0243] Step 4309, perform one round of traversal on all the first data records of the first evaluation set; and during this round of traversal, use the currently traversed first data record as the corresponding current evaluation record; set a corresponding label vector TG based on the first label classification result of the current evaluation record; use the first training sequence of the current evaluation record as the corresponding first sequence X and input it into the current model framework for processing to obtain the corresponding first prediction result R, and set a corresponding prediction vector TR based on the second prediction vector V generated by the current prediction model during this processing; form a corresponding prediction-label pair from the label vector TG and the prediction vector TR corresponding to the current evaluation record; at the end of this round of traversal, bring all the obtained prediction-label pairs into a preset first model evaluation function for calculation to obtain the corresponding first evaluation value; calculate the accuracy, precision, recall, and F1 score based on all the obtained prediction-label pairs to obtain the corresponding first accuracy, first precision, first recall, and first F1 score;

[0244] Here, the first model evaluation function of the embodiment of the present invention is implemented based on the MSE function or the RMSE function;

[0245] Step 4310: Identify whether the first evaluation value, the first accuracy rate, the first precision rate, the first recall rate, and the first F1 score all meet their corresponding first evaluation value range, first accuracy rate range, first precision rate range, first recall rate range, and first F1 score range. If all meet, proceed to Step 4311; if at least one of the first evaluation value, the first accuracy rate, the first precision rate, the first recall rate, and the first F1 score does not meet the corresponding range, return to Step 4303 to continue training.

[0246] Here, the first evaluation value range, the first accuracy rate range, the first precision rate range, the first recall rate range, and the first F1 score range in the embodiments of the present invention are five pre-set numerical ranges.

[0247] Step 4311: Confirm that the training of the second framework is completed, and save the pre-trained parameters of the encoding model and the latest set of adapter parameters as the corresponding fine-tuned encoding model parameter set; and solidify the framework model parameters of the first task model framework.

[0248] Step 5: After the training of the first and second frameworks is completed, process the methylation sequence reconstruction task of DNA sequences based on the pre-trained model framework; and process the DNA sequence source identification task corresponding to a type of target source type based on the first task model framework.

[0249] Specifically, it includes: Step 51: Process the methylation sequence reconstruction task of DNA sequences based on the pre-trained model framework.

[0250] Specifically, it includes: Receive the DNA sequence input by the user and record it as the corresponding first DNA sequence; and use the first DNA sequence as the corresponding first sequence X to be processed by the pre-trained model framework to obtain the corresponding first decoded sequence Z, and use the current first decoded sequence Z as the corresponding first reconstructed sequence to feedback to the current user.

[0251] Among them, the first DNA sequence is a marker sequence sorted by marker symbols of bases in the four-base set [A, T, C, G]; the first reconstructed sequence is a marker sequence sorted by marker symbols of bases in the five-base set [A, T, C, G, M].

[0252] Here, the user can identify the methylation sites on the current DNA sequence through the base marker M on the first reconstructed sequence.

[0253] Step 52: And process the DNA sequence source identification task corresponding to a type of target source type based on the first task model framework.

[0254] Specifically, it includes: receiving a DNA sequence input by a user and denoting it as a corresponding second DNA sequence; and using the second DNA sequence as a corresponding first sequence X to be processed by a first task model framework to obtain a corresponding first prediction result R, and using the current first prediction result R as a corresponding first recognition result to feedback to the current user;

[0255] Among them, the second DNA sequence is a sequence of marker symbols sorted by marker symbols in a four-base set [A, T, C, G] or a five-base set [A, T, C, G, M].

[0256] As can be seen from the above, the embodiments of the present invention provide a base set upgrade scheme that treats 5-mC in a DNA sequence as a type of extended base marker M and thereby upgrades the conventional four-base set [A, T, C, G] to a five-base set [A, T, C, G, M]; a data acquisition scheme with BS-seq reads as training data; a vocabulary learning scheme that learns a special vocabulary with a large amount of methylation sub-word information based on BS-seq reads, the five-base set [A, T, C, G, M], and the BPE algorithm; an encoding model design scheme for a methylation feature encoding model with the learned special vocabulary as the word segmentation vocabulary and the MosaicBERT model as the core, and a processing scheme for pre-training and fine-tuning the encoding model based on two downstream tasks (denoted as the pre-training and fine-tuning scheme of the encoding model). It should be noted that the implementation principles of the above series of technical solutions (base set upgrade scheme, data acquisition scheme, vocabulary learning scheme, encoding model design scheme, pre-training and fine-tuning scheme of the encoding model) in the embodiments of the present invention are also applicable to other modification forms in the field of epigenetic modification. For example, 6-methyladenine (6-mA) in RNA sequences, 5-hydroxymethylcytosine (5-hmC) in DNA sequences, etc. Specifically, after determining a type of other modification form, for example, assuming the current modification form is 5-hmC, an extended base marker H can be assigned to 5-hmC based on the implementation principle of the base set upgrade scheme of the embodiments of the present invention, and a corresponding five-base set [A, T, C, G, H] or six-base set [A, T, C, G, M, H] can be extended based on the conventional four-base set [A, T, C, G] or the five-base set [A, T, C, G, M] of this case; a sequencing technology type that can detect 5-hmC (such as the oxBS-seq technology) can be selected for the five-base set [A, T, C, G, H] or a sequencing technology type that can simultaneously detect 5-mC and 5-hmC (such as the EM-seq technology) can be selected for the six-base set [A, T, C, G, M, H]; the sequencing reads corresponding to the current sequencing technology can be used as training data for big data acquisition based on the implementation principle of the data acquisition scheme of the embodiments of the present invention; a special vocabulary can be learned according to the currently extended five-base set [A, T, C, G, H] or six-base set [A, T, C, G, M, H] based on the implementation principle of the vocabulary learning scheme of the embodiments of the present invention; a corresponding feature encoding model can be designed for the 5-hmC modification form based on the implementation principle of the encoding model design scheme of the embodiments of the present invention; and the current feature encoding model can be pre-trained and fine-tuned based on the implementation principle of the pre-training and fine-tuning scheme of the encoding model of the embodiments of the present invention.

[0257] Figure 3 The figure is a module structure diagram of a processing device for a methylation feature encoding model provided in the second embodiment of the present invention. The device is a terminal device or a server for implementing the foregoing method embodiment, or may be a device that enables the foregoing terminal device or server to implement the foregoing method embodiment. For example, the device may be a device or a chip system of the foregoing terminal device or server. As Figure 3 shown, the device includes: a preprocessing module 201, a model construction module 202, a first training module 203, a second training module 204, and a model application module 205.

[0258] The preprocessing module 201 is configured to set a corresponding extended identifier for 5-methylcytosine in the DNA sequence, denoted as the base identifier M, and adjust the four-base set [A, T, C, G] of DNA to the corresponding five-base set [A, T, C, G, M]; and perform big data collection on the bisulfite sequencing reads of the first species to form a corresponding BS-seq read set; and perform read splicing according to the BS-seq read set to obtain a corresponding first base sequence; and perform five-element sequence conversion on the first base sequence based on the five-base set [A, T, C, G, M] and the methylation status markers corresponding to each base identifier C on the first base sequence to obtain a corresponding second base sequence; and perform vocabulary learning on the second base sequence based on the BPE algorithm to obtain a corresponding first vocabulary; the first species is defaulted to human, and in addition to humans, it can also be one of other multiple species, such as monkeys, orangutans, and so on.

[0259] The model construction module 202 is configured to construct a methylation feature encoding model with the MosaicBERT model as the core; store the first vocabulary in the vocabulary storage module of the methylation feature encoding model; construct a sequence decoding model for methylation sequence reconstruction; and construct a corresponding binary classification prediction model based on a type of target source type, denoted as the target source prediction model; the target source type is specifically a type of tissue type or a type of cell type.

[0260] The first training module 203 is configured to form a pre-training model framework from the methylation feature encoding model and the sequence decoding model; and perform first framework training on the pre-training model framework according to the second base sequence to obtain corresponding encoding model pre-training parameters.

[0261] The second training module 204 is configured to form a first task model framework from the methylation feature encoding model and the target source prediction model; construct a model data set based on the BS-seq read set and the target source type; and perform second framework training on the first task model framework based on the encoding model pre-training parameters and the model data set.

[0262] The model application module 205 is used to process the methylation sequence reconstruction task of DNA sequences based on the pre-trained model framework after the training of the first and second frameworks; and process the DNA sequence source identification task corresponding to a target source type based on the first task model framework.

[0263] The processing device of a methylation feature encoding model provided by an embodiment of the present invention can execute the method steps in the above method embodiments, and its implementation principle and technical effects are similar, which will not be elaborated here.

[0264] It should be noted that it should be understood that the division of each module of the above device is only a logical function division. In actual implementation, it can be fully or partially integrated into a physical entity, or physically separated. And these modules can all be implemented in the form of software called by a processing element; they can also all be implemented in the form of hardware; or some modules can be implemented in the form of software called by a processing element, and some modules can be implemented in the form of hardware. For example, the preprocessing module can be a separately established processing element, or can be integrated in a certain chip of the above device. In addition, it can also be stored in the memory of the above device in the form of program code, and called and executed by a certain processing element of the above device to perform the functions of the above determined modules. The implementation of other modules is similar. In addition, all or part of these modules can be integrated together or can be independently implemented. The processing element described here can be an integrated circuit with signal processing capabilities. In the implementation process, each step of the above method or each of the above modules can be completed by the integrated logic circuit in the processor element or the instruction in the form of software.

[0265] For example, the above modules can be one or more integrated circuits configured to implement the above method, such as: one or more Application Specific Integrated Circuits (ASICs), or, one or more Digital Signal Processors (DSPs), or, one or more Field Programmable Gate Arrays (FPGAs), etc. Again, when a certain module above is implemented in the form of a program code scheduled by a processing element, the processing element can be a general-purpose processor, such as a Central Processing Unit (CPU) or other processors that can call program code. Again, these modules can be integrated together and implemented in the form of a System-on-a-chip (SOC).

[0266] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the processes or functions described in the foregoing method embodiments are generated in whole or in part. The above computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The above computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the above computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center by wire (such as coaxial cable, optical fiber, Digital Subscriber Line (DSL)) or wireless (such as infrared, wireless, Bluetooth, microwave, etc.). The above computer-readable storage medium can be any available medium that a computer can access or a data storage device such as a server or data center that includes one or more integrated available media. The above available medium can be a magnetic medium (such as a floppy disk, hard disk, magnetic tape), an optical medium (such as a DVD), or a semiconductor medium (such as a solid state disk (SSD)), etc.

[0267] Figure 4 FIG. 3 is a schematic structural diagram of an electronic device provided in Embodiment 3 of the present invention. The electronic device can be a terminal device or a server for implementing the method of the foregoing embodiments, or a terminal device or a server connected to the foregoing terminal device or server for implementing the method of the foregoing embodiments. As Figure 4 shown, the electronic device may include: a processor 301 (such as a CPU), a memory 302, and a transceiver 303; the transceiver 303 is coupled to the processor 301, and the processor 301 controls the transceiver operations of the transceiver 303. Various instructions can be stored in the memory 302 for completing various processing functions and implementing the processing steps described in the foregoing method embodiments. Preferably, the electronic device related to the embodiments of the present invention further includes: a power supply 304, a system bus 305, and a communication port 306. The system bus 305 is used to implement communication connections between components. The above communication port 306 is used for the electronic device to connect and communicate with other peripherals.

[0268] In Figure 4The system bus 305 mentioned above can be a Peripheral Component Interconnect (PCI) bus, an Extended Industry Standard Architecture (EISA) bus, or the like. This system bus can be divided into an address bus, a data bus, a control bus, etc. For the sake of convenience in representation, only a thick line is used in the figure, but it does not mean that there is only one bus or one type of bus. The communication interface is used to implement communication between the database access device and other devices (such as clients, read-write libraries, and read-only libraries). The memory may include Random Access Memory (RAM), and may also include non-volatile memory, such as at least one disk memory.

[0269] The above-mentioned processor can be a general-purpose processor, including a Central Processing Unit (CPU), a Network Processor (NP), a Graphics Processing Unit (GPU), etc.; it can also be a Digital Signal Processor (DSP), an Application Specific Integrated Circuit (ASIC), a Field Programmable Gate Array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components.

[0270] It should be noted that the embodiments of the present invention also provide a computer-readable storage medium. Instructions are stored in this computer-readable storage medium, and when it runs on a computer, it causes the computer to execute the methods and processing procedures provided in the above embodiments.

[0271] An embodiment of the present invention provides a processing method, apparatus, electronic device, and computer-readable storage medium for a methylation feature encoding model. As can be seen from the above, in the embodiment of the present invention, a corresponding base marker M is set for 5-mC in the DNA sequence, and the four-base set [A, T, C, G] of DNA is adjusted to a five-base set [A, T, C, G, M]; big data collection is performed on bisulfite sequencing reads of a certain species, and based on the collected set of BS-seq reads, operations such as read splicing and methylation status marking are performed to obtain a base sequence, and vocabulary learning is performed based on the BPE algorithm and this base sequence; then, an encoder model denoted as a methylation feature encoding model is constructed with the MosaicBERT model as the core, the learned vocabulary is stored in the vocabulary storage module inside the encoding model, a decoder model denoted as a sequence decoding model for reconstructing the methylation sequence is constructed, and a corresponding binary classification prediction model denoted as a target source prediction model is constructed based on a type of target source type; then, a pre-training model framework is composed of the methylation feature encoding model and the sequence decoding model, and a first task model framework is composed of the methylation feature encoding model and the target source prediction model; then, first, the pre-training model framework is trained based on the above base sequence, and then the first task model framework is trained; finally, the methylation sequence reconstruction task of the DNA sequence is processed based on the pre-training model framework, and a type of DNA sequence source identification task is processed based on the first task model framework. The methylation feature encoding model given in the embodiment of the present invention can learn the methylation features of the DNA sequence. The pre-training model framework with this methylation feature encoding model as the core encoder can be used to process the methylation sequence analysis task, and the first task model framework with this methylation feature encoding model as the core encoder can be used to process a type of target source identification task; based on the methylation feature encoding model given in the embodiment of the present invention and the corresponding two task model frameworks to process the methylation sequence analysis task and the target source identification task, not only simplifies the task processing steps, shortens the task processing cycle, improves the real-time performance of the task processing, but also effectively reduces the task processing cost.

[0272] The steps of the methods or algorithms described in connection with the embodiments disclosed herein may be implemented in hardware, software modules executed by a processor, or a combination of both. The software modules may be disposed in a random access memory (RAM), memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, registers, hard disk, removable disk, CD-ROM, or any other form of storage medium known in the art.

[0273] The specific embodiments described above further elaborate on the object, technical solution, and beneficial effects of the present invention. It should be understood that the above description is only for the specific embodiments of the present invention and is not intended to limit the protection scope of the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.

Claims

1. A processing method for a methylation feature encoding model, characterized in that The method includes: Setting a corresponding extended marker for 5-methylcytosine in the DNA sequence, denoted as the base marker M, and adjusting the four-base set [A, T, C, G] of DNA to the corresponding five-base set [A, T, C, G, M]; and collecting big data of bisulfite sequencing reads of the first species to form a corresponding BS-seq read set; and performing read splicing according to the BS-seq read set to obtain a corresponding first base sequence; and performing five-element sequence conversion on the first base sequence based on the five-base set [A, T, C, G, M] and the methylation status marker corresponding to each base marker C on the first base sequence to obtain a corresponding second base sequence; and performing vocabulary learning based on the BPE algorithm and the second base sequence to obtain a corresponding first vocabulary; Constructing a methylation feature encoding model with the Mosa i cBERT model as the core; and storing the first vocabulary in the vocabulary storage module of the methylation feature encoding model; and constructing a sequence decoding model for methylation sequence reconstruction; and constructing a corresponding binary classification prediction model based on a type of target source type, denoted as the target source prediction model; the target source type is specifically a type of tissue type or a type of cell type; Composing a pre-training model framework from the methylation feature encoding model and the sequence decoding model; and performing first framework training on the pre-training model framework according to the second base sequence to obtain corresponding pre-training parameters of the encoding model; Composing a first task model framework from the methylation feature encoding model and the target source prediction model; and constructing a model data set based on the BS-seq read set and the target source type; and performing second framework training on the first task model framework based on the pre-training parameters of the encoding model and the model data set; After the first and second framework trainings are completed, processing the methylation sequence reconstruction task of the DNA sequence based on the pre-training model framework; and processing a DNA sequence source identification task corresponding to a type of the target source type based on the first task model framework.

2. The processing method of the methylation feature encoding model according to claim 1, wherein The base markers A, T, C, G are markers for adenine, thymine, cytosine, and guanine bases respectively; The set of BS-seq reads consists of multiple BS-seq read records; each of the BS-seq read records includes at least one BS-seq read and a corresponding read sample source; each of the BS-seq reads is a DNA sequence obtained by DNA sequencing after bisulfite sequencing treatment of a DNA sample of the first species, and is sorted by the base markers A, T, C, G; each base marker C on each of the BS-seq reads has a corresponding methylation status marker; the methylation status marker is a binary status marker, including two states: unmethylated and methylated; the read sample source is a source type code corresponding to the current BS-seq read, and each source type code corresponds to a specific source type; the type range of the source type consists of multiple tissue types and / or multiple cell types of the first species; The first base sequence is a marker sequence sorted by the base markers A, T, C, G; the second base sequence is a marker sequence sorted by the base markers A, T, C, G, M; The first vocabulary consists of multiple first sub-word records; the first sub-word record includes a first sub-word code, a first sub-word text, and a first sub-word frequency; the first sub-word text is a string sorted by one or more base markers from the five-base set [A, T, C, G, M]; the first sub-word frequency is an integer value; all the first sub-word texts in the first vocabulary are different; The model data set includes multiple first data records; the first data record includes a first training sequence and a first label classification result; the first label classification result includes the target source type and non-target source types.

3. The processing method of the methylation feature coding model according to claim 2, wherein The big data collection of bisulfite sequencing reads of the first species to form a corresponding set of BS-seq reads specifically includes: Data collection of bisulfite sequencing reads of the DNA sequence of the first species is performed through multiple preset data collection channels to obtain multiple BS-seq reads; the multiple preset data collections include at least one or more publicly available BS-seq read data sets, public technical literatures with BS-seq read information of the first species; And based on the sample source information corresponding to all the collected BS-seq reads, the type range of the corresponding source type is set; and a corresponding source type code is set for each source type in the type range of the source type; and the source type code corresponding to each BS-seq read is used as the corresponding read sample source; and each BS-seq read and the corresponding read sample source form a corresponding BS-seq read record; and error read deletion and duplicate read deduplication processing are performed on all the obtained BS-seq read records; And the BS-seq read set is composed of all the remaining BS-seq read records.

4. The processing method of the methylation feature encoding model according to claim 2, characterized in that, Performing read splicing according to the BS-seq read set to obtain a corresponding first base sequence, specifically including: Sequentially splicing all the BS-seq reads in the BS-seq read set to obtain the corresponding first base sequence.

5. The processing method of the methylation feature coding model according to claim 2, characterized in that Based on the five-base set [A, T, C, G, M] and the methylation status markers corresponding to the base marker C on the first base sequence, performing five-element sequence conversion on the first base sequence to obtain a corresponding second base sequence, specifically including: Performing sequence replication on the first base sequence to obtain a corresponding first replication sequence; modifying the base marker C with the status value of methylated for each current methylation status marker on the first replication sequence to the corresponding base marker M; and using the first replication sequence after the marker modification as the corresponding second base sequence.

6. The processing method of the methylation feature coding model according to claim 2, wherein Based on the BPE algorithm and the second base sequence, performing vocabulary learning to obtain a corresponding first vocabulary, specifically including: Step 61, setting an empty vocabulary as the corresponding first vocabulary; initializing the first vocabulary; setting the record total threshold of the first vocabulary; and using the second base sequence as a corresponding first string. Among them, the initialized first vocabulary is composed of six initial first sub-word records; the first sub-word texts of the six initial first sub-word records are the character 'A', the character 'T', the character 'C', the character 'G', the character 'M', and the string 'CG' in sequence; the first sub-word frequencies of the six initial first sub-word records are all initialized to 1. Step 62, using each first sub-word text in the first vocabulary as a corresponding basic sub-word, and forming a corresponding current sub-word set from all the obtained basic sub-words; and according to the sub-word sequence splitting method of the BPE algorithm, splitting the first string according to the current sub-word set to obtain a corresponding current sub-word sequence. Among them, the current sub-word sequence is composed of multiple basic sub-words sorted in order; the basic sub-word pair is composed of two adjacent basic sub-words sorted in front and behind. Step 63: According to the sub-word pair combination method of the BPE algorithm, every two adjacent basic sub-words in the current sub-word sequence form a corresponding basic sub-word pair; cluster the same basic sub-word pairs to obtain multiple clustered word pair sets; take the total number of the basic sub-word pairs in each clustered word pair set as the corresponding first word pair count, and take the largest of the first word pair counts as the corresponding high-frequency word pair count; identify the high-frequency word pair count; if the high-frequency word pair count is 1, go to step 68; if the high-frequency word pair count is greater than 1, take the total number of the clustered word pair sets corresponding to the high-frequency word pair count as the corresponding high-frequency word pair total, and identify the high-frequency word pair total. If the high-frequency word pair total is 1, go to step 64; if the high-frequency word pair total is greater than 1, go to step 65; Among them, each clustered word pair set consists of one or more basic sub-word pairs, and all the basic sub-word pairs in each clustered word pair set are the same; Step 64: Take the basic sub-word pair corresponding to the only clustered word pair set corresponding to the high-frequency word pair total as the corresponding final selected sub-word pair; and go to step 66; Step 65: Denote all the clustered word pair sets corresponding to the high-frequency word pair total as the corresponding candidate sets; take the basic sub-word pairs corresponding to each candidate set as the corresponding candidate sub-word pairs; denote the character 'C', the character 'M', and the string 'CG' as methylation-related characters; denote each candidate sub-word pair containing the methylation-related characters as a preferred sub-word pair; identify the total number of the preferred sub-word pairs; if the total number of the preferred sub-word pairs is 0, randomly select one from all the candidate sub-word pairs as the corresponding final selected sub-word pair; if the total number of the preferred sub-word pairs is greater than 0, randomly select one from all the preferred sub-word pairs as the corresponding final selected sub-word pair; Step 66, add a first sub-word record in the first vocabulary as the corresponding current newly added record, and set the first sub-word text of the current newly added record as the corresponding final selected sub-word pair; after the record setting is completed, use each first sub-word text in the current first vocabulary as a new base sub-word, and form the latest current sub-word set from all the latest obtained base sub-words; and according to the sub-word sequence splitting method of the BPE algorithm, split the first string according to the current sub-word set to obtain the latest current sub-word sequence; and count the total number of each base sub-word appearing in the current sub-word sequence and use the statistical result as the corresponding first word frequency, and set the first word frequency corresponding to each base sub-word that does not appear in the current sub-word sequence in the current sub-word set to 0; and update the first word frequency corresponding to each first sub-word text in the current first vocabulary to the corresponding first word frequency; and after the update is completed, delete the first sub-word record in the first vocabulary with the first word frequency of 0. Step 67, count the total number of the first sub-word records in the current first vocabulary to obtain the corresponding current record total; and identify whether the current record total is less than the record total threshold; if so, return to Step 63; if not, go to Step 68. Step 68, perform a re-sorting on all the first sub-word records in the latest first vocabulary in descending order according to the first word frequency; and based on the preset sub-word encoding rule, perform encoding setting on the first sub-word encoding corresponding to each first sub-word text in the sorted first vocabulary; and output the first vocabulary after the sub-word encoding setting is completed as the vocabulary learning result of this time.

7. The processing method of the methylation feature encoding model according to claim 2, wherein The methylation feature encoding model is used to perform feature encoding on the first sequence X input to the model and output the corresponding first encoding vector Y; the first sequence X is composed of multiple first tokens x i sorted in sequence; each of the first tokens x i is a type of base token in the five-base set [A, T, C, G, M]; 1 ≤ token index i ≤ L, where L is the sequence length of the first sequence X; the first encoding vector Y is composed of multiple first segmented feature vectors y j ; 1 ≤ segmentation index j ≤ W, where W is the sequence length of the first segmented sequence S corresponding to the first sequence X. the methylation feature encoding model is composed of a sequence tokenizer, the word table storage module and the Mosa i cBERT model; the input end of the sequence tokenizer is connected to the model input end of the methylation feature encoding model, and the output end is connected to the input end of the Mosa i cBERT model; the sequence tokenizer is also connected to the word table storage module; the output end of the Mosa i cBERT model is connected to the model output end of the methylation feature encoding model; The sequence tokenizer is used to split the first sequence X into a first sub-word sequence composed of a plurality of the first sub-word texts according to the sub-word sequence splitting method of the BPE algorithm, based on the first vocabulary stored in the vocabulary storage module; and add a preset classification sub-word text 'CLS' before the first sub-word text in the first sub-word sequence; and use the sequence length of the first sub-word sequence after the addition as the corresponding sequence length W; and use each sub-word text of the first sub-word sequence as a corresponding first token s j , and from all the first tokens s j to form the corresponding first token sequence S; and based on the first sub-word encoding corresponding to each of the first sub-word texts in the first vocabulary and the classification sub-word encoding corresponding to the classification sub-word text 'CLS', set the corresponding encoding for each of the first tokens s j in the first token sequence S to obtain the corresponding first token encoding, and form the corresponding first token encoding sequence from all the obtained first token encodings; and identify the built-in masking rule type; if the masking rule type is no masking, then do not perform masking modification on the first token encoding sequence; if the masking rule type is a type of random masking, then randomly mask the first token encodings in the first token encoding sequence according to the built-in type of random masking ratio and the preset token masking encoding, and ensure that the ratio of the total number of the masked first token encodings to the sequence length of the first token encoding sequence matches the type of random masking ratio; if the masking rule type is a second type of random masking, then mark the first tokens s j in the first token sequence S that contain methylation-related characters as methylation-related tokens, and mark the first token encodings corresponding to each of the methylation-related tokens in the first token encoding sequence as methylation-related encodings, and randomly mask the methylation-related encodings in the first token encoding sequence according to the built-in second type of random masking ratio and the token masking encoding, and ensure that the ratio of the total number of the masked methylation-related encodings to the total number of the methylation-related encodings in the first token encoding sequence matches the second type of random masking ratio; and based on the model input vector embedding encoding rule of the MosaicBERT model, perform embedding encoding processing on the first token encoding sequence to obtain the corresponding first embedding encoding vector E and send it to the MosaicBERT model; the first token sequence S is composed of W first tokens s j , the first first token s j=1 is the classification sub-word text 'CLS', and each of the first tokens s 2≤j≤W from the second to the last first token s j corresponding to one of the first sub-word texts in the first vocabulary; the first embedding encoding vector E is composed of W first word-segment embedding encoding vectors e j and the first word-segment embedding encoding vector e j corresponds to the first word segment s j in one-to-one correspondence; the Mosa i cBERT model is used to perform feature encoding processing according to the first embedding encoding vector E and output the corresponding first encoding vector Y.

8. The processing method of the methylation feature encoding model according to claim 7, wherein The sequence decoding model is used to reconstruct the methylation sequence according to the first coding vector Y input to the model and output the corresponding first decoded sequence Z; the first decoded sequence Z is composed of L second tokens z i sequentially sorted; each of the second tokens z i is a base token of the five-base set [A, T, C, G, M]; the second token z i corresponds one-to-one with the first token x i one by one; the sequence decoding model is composed of a first mapping layer, a first linear layer and a first Softmax layer; the first mapping layer is implemented based on a fully connected network; the first linear layer is formed based on another fully connected network; The input end of the first mapping layer is connected to the model input end of the sequence decoding model, and the output end is connected to the input end of the first linear layer; the output end of the first linear layer is connected to the input end of the first Softmax layer; the output end of the first Softmax layer is connected to the model output end of the sequence decoding model; The first mapping layer is used to extract the second to the W-th first word segmentation feature vectors y in the first encoded vector Y to form a first extraction vector A with a length of W - 1; and convert the first extraction vector A into a first mapping vector M with a length of L through a linear transformation and send it to the first linear layer; the first extraction vector A is composed of W - 1 first extraction sub-vectors a 2≤j≤W where 1 ≤ vector index k ≤ (W - 1), and each of the first extraction sub-vectors a k matches the corresponding first word segmentation feature vector y k ; the first mapping vector M is composed of L first sub-mapping vectors m j=k+1 ; i ​ The first linear layer is used to perform a fully connected calculation according to the first mapping vector M to obtain a corresponding first scoring vector B and send it to the first Softmax layer; The first scoring vector B consists of L first scoring values b i and each of the first scoring values b i is a real number; The first Softmax layer is used to perform five-class probability prediction based on the first scoring vector B to obtain a corresponding first prediction vector U; and use the base marker corresponding to the largest first prediction probability among the first sub-vectors u in the first prediction vector U as the corresponding second marker z i ; and use the base marker corresponding to the largest first prediction probability among the first sub-vectors u in the first prediction vector U as the corresponding second marker z i ; and form a corresponding first decoding sequence Z by sequentially sorting all the obtained second markers z i ; and form a corresponding first decoding sequence Z by sequentially sorting all the obtained second markers z The first prediction vector U is composed of L first sub-vectors u i Each of the first sub-vectors u i is composed of five first prediction probabilities, and each of the first prediction probabilities corresponds to a type of base identifier in the five-base set [A, T, C, G, M].

9. The processing method of the methylation feature encoding model according to claim 7, characterized in that, The target source prediction model is used to perform binary classification prediction on the input first encoding vector Y of the model and output a corresponding first prediction result R; the first prediction result R is a binary prediction result, consisting of two prediction results: the target source type and the non-target source type; The target source prediction model consists of a first extraction module, a second linear layer and a second Softmax layer; the second linear layer is based on another fully connected network; The input end of the first extraction module is connected to the model input end of the target source prediction model, and the output end is connected to the input end of the second linear layer; the output end of the second linear layer is connected to the input end of the second Softmax layer; the output end of the second Softmax layer is connected to the model output end of the target source prediction model; The first extraction module is used to extract the first word segmentation feature vector y of the 1st in the first encoding vector Y and send it as the corresponding second extraction vector C to the second linear layer; j=1 ​ The second linear layer is used to perform a fully connected calculation according to the second extraction vector C to obtain a corresponding second scoring vector D and send it to the second Softmax layer; the second scoring vector D consists of two second scores d1 and d2; The second Softmax layer is used to perform binary classification probability prediction on the second scores d1 and d2 of the second scoring vector D to obtain a corresponding second prediction vector V; and identify the two second prediction probabilities v1 and v2 of the second prediction vector V; if the second prediction probability v1 is larger, set the corresponding first prediction result R to the target source type; if the second prediction probability v2 is larger, set the corresponding first prediction result R to the non-target source type; the second prediction vector V consists of two second prediction probabilities v1 and v2, where the second prediction probability v1 is the true value probability and the second prediction probability v2 is the false value probability.

10. The processing method of the methylation feature coding model according to claim 8, characterized in that, The first framework training of the pre-trained model framework according to the second base sequence to obtain corresponding encoding model pre-training parameters specifically includes: Step 1001: Use the pre-trained model framework as the corresponding current model framework; use the methylation feature encoding model and the sequence decoding model of the current model framework as the corresponding current encoding model and current decoding model; set two positive integers as the corresponding first and second stage training times thresholds; set two proportion parameters with values between 0 and 1 as the corresponding first and second proportion parameters; set the first and second random masking ratios built into the sequence tokenizer of the current encoding model as the corresponding first and second proportion parameters; set two counters initialized to 0 as the corresponding first and second counters; Among them, the first proportion parameter < the second proportion parameter, 10% < the first proportion parameter < 20%, 90% < the second proportion parameter < 100%; Step 1002: Use the second base sequence as the corresponding first sequence X; perform label vector conversion on the current first sequence X to obtain the corresponding label vector PG; perform smoothed label vector conversion on the label vector PG to obtain the corresponding smoothed label vector PS; Among them, the label vector PG consists of L sub-vectors pg i ; the sub-vector pg i corresponds one-to-one with the first token x i of the first sequence X; each sub-vector pg i is a one-hot encoding vector of length 5, composed of five-base one-hot encodings pgA i , pgT i , pgC i , pgG i , pgM i ; the five-base one-hot encodings of each sub-vector pg i correspond one-to-one with five types of base tokens in the five-base set [A, T, C, G, M]; only one of the five-base one-hot encodings of each sub-vector pg i is 1, and the other four are 0, and the base token corresponding to the one-hot encoding of 1 matches the first token x i corresponding to the current sub-vector pg i . The smoothed label vector PS consists of L sub-vectors ps i ; The sub-vector ps i corresponds one-to-one with the sub-vector pg i of the label vector PG; The vector length of each sub-vector ps i is 5, and it consists of 5 vector encodings psA i , psT i , psC i , psG i , psM i ; The 5 vector encodings of each sub-vector ps i correspond one-to-one with five types of base markers in the five-base set [A, T, C, G, M]; The sub-vector pg i and the sub-vector ps i have the following conversion relationship: α is a preset smoothing parameter, with a value between 0 and 1; I is a preset all-1 vector with a length of 5; the first condition is that the sub-vector pg i corresponding to the first identifier x i is the base identifier C or M; the second condition is that the sub-vector pg i corresponding to the first identifier x i is the base identifier A, T or G; Step 1003: Set the masking rule type built into the sequence tokenizer of the current encoding model as the first type of random masking; Step 1004: Input the current first sequence X into the sequence tokenizer of the current encoding model for processing to obtain the corresponding first embedded encoding vector E; Step 1005: Input the first embedded encoding vector E into the MosaicBERT model of the current encoding model for processing to obtain the corresponding first encoding vector Y; input the first encoding vector Y into the current decoding model for processing to obtain the corresponding first decoding sequence Z; and use each of the first sub-vectors u of the first prediction vector U generated during the current processing of the current decoding model i as a corresponding sub-vector pr i , and form the corresponding prediction vector PR from all the obtained sub-vectors pr i ; Among them, the prediction vector PR consists of L sub-vectors pr i ; the sub-vector pr i corresponds one-to-one with the second token z of the first decoding sequence Z i ; the five first prediction probabilities of the sub-vector pr i are denoted as prA i , prT i , prC i , prG i , prM i , and respectively correspond one-to-one with five types of base tokens in the five-base set [A, T, C, G, M]; Step 1006, input the prediction vector PR and the smoothed label vector PS into a preset first model loss function L M1 and calculate to obtain the corresponding first loss value; Among them, the first model loss function L M1 is as follows: Step 1007, identify whether the first loss value meets a preset first loss value range; if it meets, increment the first counter by 1 and proceed to Step 1008; if it does not meet, based on a preset first model optimizer, perform one round of modulation on the full model parameters of the current encoding model and the full model parameters of the current decoding model in the direction of minimizing the first model loss function L M1 to reach the minimum value, and return to Step 1005 at the end of this round of modulation; Among them, the first model optimizer includes at least the Adam optimizer and the SGD optimizer; Step 1008: Identify whether the first counter exceeds the first stage training times threshold; if not, return to Step 1004 to continue training; if it exceeds, go to Step 1009; Step 1009: Set the masking rule type built into the sequence tokenizer of the current encoding model as the second type of random masking; Step 1010: Input the current first sequence X into the sequence tokenizer of the current encoding model for processing to obtain the corresponding first embedded encoding vector E; Step 1011, input the first embedded encoding vector E into the MosaicBERT model of the current encoding model for processing to obtain the corresponding first encoding vector Y; input the first encoding vector Y into the current decoding model for processing to obtain the corresponding first decoding sequence Z; and input each of the first sub-vectors u of the first prediction vector U generated during the current processing of the current decoding model i as a corresponding sub-vector pr i , and all the obtained sub-vectors pr i are used to form the corresponding prediction vector PR; Step 1012, input the prediction vector PR and the smoothed label vector PS into the first model loss function L M1 and calculate the corresponding second loss value; Step 1013: Identify whether the second loss value meets a preset second loss value range; if it meets, increment the second counter by 1 and proceed to step 1014; if it does not meet, based on the first model optimizer, perform one round of modulation on the full model parameters of the current encoding model and the full model parameters of the current decoding model in the direction of minimizing the first model loss function L M1 to reach the minimum value, and return to step 1011 at the end of this round of modulation; Step 1014: Identify whether the second counter exceeds the second stage training times threshold; if not, return to Step 1010 to continue training; if it exceeds, go to Step 1015; Step 1015: Confirm that the training of the first framework is completed, use the current model parameters of the current encoding model as the corresponding pre-trained parameters of the encoding model and save them; and solidify the framework model parameters of the pre-trained model framework.

11. The processing method of the methylation feature coding model according to claim 2, characterized in that The construction of the model dataset based on the BS-seq read set and the target source type specifically includes: Compose the corresponding positive sample set from the BS-seq read records in the BS-seq read set whose read sample sources match the target source type, and compose the corresponding negative sample set from the BS-seq read records in the BS-seq read set whose read sample sources do not match the target source type; And use the BS-seq reads recorded in each of the positive sample sets as a corresponding first training sequence, and set a first label classification result specifically set as the target source type; and form a corresponding first data record from the first training sequence and the first label classification result corresponding to each of the BS-seq read records in the positive sample set; And use the BS-seq reads recorded in each of the negative sample sets as a corresponding first training sequence, and set a first label classification result specifically set as the non-target source type; and form a corresponding first data record from the first training sequence and the first label classification result corresponding to each of the BS-seq read records in the negative sample set; And form the corresponding model data set from all the obtained first data records.

12. The processing method of the methylation feature coding model according to claim 9, wherein The second framework training of the first task model framework based on the pre-trained parameters of the encoding model and the model data set specifically includes: Step 1201, use the first task model framework as the corresponding current model framework; and use the methylation feature encoding model and the target source prediction model of the current model framework as the corresponding current encoding model and current prediction model; and set the masking rule type built into the sequence tokenizer of the current encoding model to no masking; Step 1202, solidify the model parameters of the methylation feature encoding model as the corresponding pre-trained parameters of the encoding model; and based on the LoRA fine-tuning mechanism, load a corresponding LoRA adapter on one or more specified modules of the methylation feature encoding model; and form a corresponding adapter parameter set from the adapter parameters of all the LoRA adapters; Among them, the solidified model parameters of each of the specified modules are denoted as the pre-fine-tuning parameters W0, and the adapter parameters of the LoRA adapters corresponding to each of the specified modules are composed of the matrix parameters of two low-rank matrices A and B; each LoRA adapter is used to fine-tune the pre-fine-tuning parameters W0 of the corresponding specified module to obtain the corresponding post-fine-tuning parameters W1; the conversion relationship between the post-fine-tuning parameters W1 and the pre-fine-tuning parameters W0 is: W1X in = W0X in + BAX in , X in is the input vector of the specified module, and W0X in is the output vector obtained by processing the input vector X in by the current specified module without adding the corresponding LoRA adapter, and W1X in is the output vector obtained by processing the input vector X in by the current specified module after adding the corresponding LoRA adapter; Step 1203, divide the model data set into two sub-data sets according to a preset first splitting ratio and record them as the corresponding first training set and first evaluation set; Wherein, both the first training set and the first evaluation set are composed of multiple first data records; the ratio of the total number of records in the first training set to the total number of records in the first evaluation set satisfies the first splitting ratio; the positive and negative sample ratios in the first training set and the first evaluation set are the same; Step 1204, use the first first data record in the first training set as the corresponding current training record; Step 1205, set a corresponding label vector TG based on the first label classification result of the current training record; Among them, the label vector TG is composed of two one-hot encodings tg1 and tg2; the one-hot encodings tg1 and tg2 correspond to the target source type and the non-target source type respectively; if the first label classification result of the current training record is the target source type, then the one-hot encoding tg1 is 1 and the one-hot encoding tg2 is 0; if the first label classification result of the current training record is the non-target source type, then the one-hot encoding tg1 is 0 and the one-hot encoding tg2 is 1; Step 1206, take the first training sequence of the current training record as the corresponding first sequence X and input it into the current model framework for processing to obtain the corresponding first prediction result R; and set a corresponding prediction vector TR based on the second prediction vector V generated by the current prediction model during this processing; Among them, the prediction vector TR is composed of two vector encodings tr1 and tr2; the vector encoding tr1 is the second prediction probability v1 of the second prediction vector V, and the vector encoding tr2 is the second prediction probability v2 of the second prediction vector V; Step 1207, input the prediction vector TR and the label vector TG into a preset second model loss function L M2 and calculate to obtain a corresponding third loss value; Among them, the second model loss function L M2 is implemented based on the L1 loss function, the L2 loss function or the cross-entropy loss function; Step 1208: Identify whether the third loss value meets a preset third loss value range; if the third loss value meets the third loss value range, identify whether the current training record is the last first data record of the first training set. If so, go to Step 1209; if not, use the next first data record of the first training set as the new current training record and return to Step 1205 to continue training; if the third loss value does not meet the third loss value range, based on a preset second model optimizer, perform one round of modulation on the adapter parameter set and the full model parameters of the current prediction model in the direction of minimizing the second model loss function L M2 to reach the minimum value, and return to Step 1206 at the end of this round of modulation; Among them, the second model optimizer includes at least an Adam optimizer and an SGD optimizer; Step 1209, perform a round of traversal on all the first data records in the first evaluation set; and during this round of traversal, take the currently traversed first data record as the corresponding current evaluation record; and set a corresponding label vector TG based on the first label classification result of the current evaluation record; and take the first training sequence of the current evaluation record as the corresponding first sequence X and input it into the current model framework for processing to obtain the corresponding first prediction result R, and set a corresponding prediction vector TR based on the second prediction vector V generated by the current prediction model during this processing; and form a corresponding prediction-label pair from the label vector TG and the prediction vector TR corresponding to the current evaluation record; and at the end of this round of traversal, bring all the obtained prediction-label pairs into a preset first model evaluation function for calculation to obtain the corresponding first evaluation value; and calculate the accuracy, precision, recall, and F1 score based on all the obtained prediction-label pairs to obtain the corresponding first accuracy, first precision, first recall, and first F1 score; Among them, the first model evaluation function is implemented based on the MSE function or the RMSE function; Step 1210, identify whether the first evaluation value, the first accuracy, the first precision, the first recall, and the first F1 score all meet the corresponding first evaluation value range, first accuracy range, first precision range, first recall range, and first F1 score range; if all meet, go to Step 1211; if at least one of the first evaluation value, the first accuracy, the first precision, the first recall, and the first F1 score does not meet the corresponding range, return to Step 1203 to continue training; Step 1211: Confirm the end of the second framework training, save the pre-trained parameters of the encoding model and the latest set of adapter parameters as the corresponding fine-tuned encoding model parameter set; and solidify the framework model parameters of the first task model framework.

13. The processing method of the methylation feature coding model according to claim 8, wherein, The methylation sequence reconstruction task of DNA sequences processed based on the pre-trained model framework specifically includes: Receiving the DNA sequence input by the user and denoting it as the corresponding first DNA sequence; and using the first DNA sequence as the corresponding first sequence X to input into the pre-trained model framework for processing to obtain the corresponding first decoded sequence Z, and feeding back the current first decoded sequence Z as the corresponding first reconstructed sequence to the current user; where the first DNA sequence is a marker sequence sorted by marker symbols of the four-base set [A, T, C, G] of DNA; the first reconstructed sequence is a marker sequence sorted by marker symbols of the five-base set [A, T, C, G, M] of DNA.

14. The processing method of the methylation feature coding model according to claim 9, wherein The DNA sequence source identification task of a type corresponding to the target source type processed based on the first task model framework specifically includes: Receiving the DNA sequence input by the user and denoting it as the corresponding second DNA sequence; and using the second DNA sequence as the corresponding first sequence X to input into the first task model framework for processing to obtain the corresponding first prediction result R, and feeding back the current first prediction result R as the corresponding first identification result to the current user; where the second DNA sequence is a marker sequence sorted by marker symbols of the four-base set [A, T, C, G] or the five-base set [A, T, C, G, M] of DNA.

15. An apparatus for performing the processing method of the methylation feature encoding model according to any one of claims 1-14, characterized in that, The device includes: a preprocessing module, a model construction module, a first training module, a second training module, and a model application module; The preprocessing module is used to set a corresponding extended marker symbol for 5-methylcytosine in the DNA sequence and denote it as the marker symbol M, and adjust the four-base set [A, T, C, G] of DNA to the corresponding five-base set [A, T, C, G, M]; and perform big data collection on the bisulfite sequencing reads of the first species to form the corresponding BS-seq read set; and perform read splicing according to the BS-seq read set to obtain the corresponding first base sequence; and perform five-element sequence conversion on the first base sequence based on the five-base set [A, T, C, G, M] and the methylation status markers corresponding to each base marker C on the first base sequence to obtain the corresponding second base sequence; and perform vocabulary learning based on the BPE algorithm and the second base sequence to obtain the corresponding first vocabulary; The model construction module is used to construct a methylation feature encoding model with the Mosa i cBERT model as the core; store the first vocabulary into the vocabulary storage module of the methylation feature encoding model; construct a sequence decoding model for methylation sequence reconstruction; and construct a corresponding binary classification prediction model based on a type of target source type, denoted as the target source prediction model; the target source type is specifically a type of tissue type or a type of cell type; The first training module is used to form a pre-training model framework by the methylation feature encoding model and the sequence decoding model; and perform first-frame training on the pre-training model framework according to the second base sequence to obtain corresponding encoding model pre-training parameters; The second training module is used to form a first task model framework by the methylation feature encoding model and the target source prediction model; construct a model data set based on the BS-seq read set and the target source type; and perform second-frame training on the first task model framework based on the encoding model pre-training parameters and the model data set; The model application module is used to, after the first and second frame trainings are completed, process the methylation sequence reconstruction task of the DNA sequence based on the pre-training model framework; and process a DNA sequence source identification task corresponding to a type of the target source type based on the first task model framework.

16. An electronic device, characterized in that, Including: A memory, a processor, and a transceiver; The processor is used to be coupled with the memory, read and execute instructions in the memory to implement the method according to any one of claims 1-14; The transceiver is coupled with the processor, and the processor controls the transceiver to perform message sending and receiving.

17. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions, and when the computer instructions are executed by a computer, the computer is caused to execute the method according to any one of claims 1-14.

Citation Information

Patent Citations

  • Deep learning cancer risk prediction method and system based on methylation sequence

    CN117727371A

  • Pre-training method and device for generative large language model

    CN118551750A

  • Method and apparatus for automatically generating inference questions and answers

    WO2021184311A1