Chromatin interaction prediction method and system based on dynamic word segmentation and word embedding

By adopting dynamic word segmentation and word embedding technology in chromatin interaction prediction, combining DNA sequence characteristics and genomic characteristics, and using integrated learning models for prediction, the problem of overfitting the model and insufficient evaluation indicators is solved, and the accuracy and comprehensiveness of the prediction results are improved.

CN120183484APending Publication Date: 2025-06-20SHANDONG UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510243495.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-03
Publication Date
2025-06-20

AI Technical Summary

Technical Problem

The prior art tends to overfit the model when processing highly overlapping chromatin interaction data, and the evaluation indicators are not sufficient to comprehensively evaluate the model performance, resulting in unsatisfactory prediction performance.

Method used

Using a method based on dynamic word segmentation and word embedding, the sequence information utilization rate is improved through bidirectional feature fusion, and DNA sequence features and genomic features are fused to generate joint features, and chromatin interaction prediction is used to use an integrated learning model.

Benefits of technology

It significantly improves the accuracy of chromatin interaction prediction results, avoids the problem of overfitting the model, and provides a more comprehensive performance evaluation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120183484A_ABST
    Figure CN120183484A_ABST
Patent Text Reader

Abstract

The invention belongs to the field of gene data processing, and provides a chromatin interaction prediction method and system based on dynamic word segmentation and word embedding. The method comprises the steps that DNA sequence information of a data sample is obtained, dynamic word segmentation processing is conducted on the DNA sequence information according to the size of a set vocabulary, and then two labeled subsequences are obtained according to the length and the occurrence frequency of words; converting all the labeled subsequences of the data sample into DNA sequence features through an embedded layer; fusing the known genome characteristics of the data sample with the DNA sequence characteristics to generate joint characteristics; and obtaining chromatin interaction prediction results of the sub-models based on a relationship between the joint features and the chromatin interaction prediction results of the sub-models in the integrated learning model, and averaging the chromatin interaction prediction results to obtain a final chromatin interaction prediction result.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of gene data processing, and particularly relates to a method and system for predicting chromatin interactions based on dynamic word segmentation and word embedding. Background Art

[0002] The statements in this section merely provide background technical information related to the present invention and do not necessarily constitute prior art.

[0003] Interactions within chromatin, especially the three-dimensional structure of the genome, have been widely recognized as playing a key role in regulating cell function and pathological states. These interactions are crucial for gene transcription, regulation, and expression. They can regulate the accessibility of regulatory elements to target genes, thereby affecting transcription efficiency. Chromatin interactions have also been found to enhance gene expression, activate transcription by cooperating with the promoters of specific genes, and serve as a bridge for communication between enhancers and genes. Therefore, a deep understanding of the complex network of chromatin interactions is crucial for decoding the gene expression regulation mechanism.

[0004] Advances in related genomic technologies, especially high-throughput chromosome conformation capture (Hi-C) and chromatin interaction analysis (ChIA-PET), have greatly promoted the understanding of chromatin interactions on a genome-wide scale. To achieve low-cost and high-precision analysis, researchers have developed a series of computational methods to predict chromatin interactions, such as TargetFinder, JEME, and RIPPLE, which utilize multiple functional genomic data including open chromatin data of transcription factors, chromatin immunoprecipitation sequencing (ChIP-seq) data, and histone modifications. In addition, there are also some sequence analysis-based methods, such as SPEID, SIMCNN, and EPIVAN.

[0005] However, due to the particularity of chromatin interaction data, traditional sample random splitting methods may lead to overfitting of the model when dealing with highly overlapping data. This problem may cause the performance shown by some methods to be over-optimistically estimated. To solve this problem, Whalen et al. redesigned a new dataset of chromatin interactions and proposed using a chromatin splitting strategy for training and prediction. Nevertheless, the models used are not satisfactory in terms of performance, and the evaluation metrics adopted cannot comprehensively evaluate the performance of the models. Based on the new dataset, IChrom-Deep has made significant progress in related research using a model based on the attention mechanism. However, from the results of sequence prediction only, IChrom-Deep still has room for improvement in sequence feature extraction, and it can only receive sequences of fixed length, with relatively large limitations.

[0006] When exploring the importance of motifs, SPEID uses the average difference value of prediction performance as a measure. That is, it divides the change value of the index corresponding to each motif by the average number of occurrences of that motif in each sample to obtain the unit impact of each motif on the sample prediction result. However, this calculation method ignores an important piece of information, namely the proportion of the number of occurrences of this motif in the total number of occurrences of all motifs. To prevent bias towards longer motifs, after matching the corresponding motif, the mutation range in SPEID is set to a 20bp window centered on the matching center. If the matching center is within 10bp of the end of the input sequence, then the 20bp at the end of the sequence will be mutated. However, this operation actually modifies more information of the original sequence data, especially for short motifs, and cannot reasonably prevent bias towards long motifs, thus affecting the accuracy of chromatin interaction prediction results. Summary of the Invention

[0007] To solve the technical problems existing in the above background art, the present invention provides a chromatin interaction prediction method and system based on dynamic word segmentation and word embedding, which significantly improves the utilization rate of sequence information through bidirectional feature fusion.

[0008] To achieve the above object, the present invention adopts the following technical solutions:

[0009] The first aspect of the present invention provides a chromatin interaction prediction method based on dynamic word segmentation and word embedding.

[0010] A chromatin interaction prediction method based on dynamic word segmentation and word embedding, comprising:

[0011] Obtain the DNA sequence information of the data sample, perform dynamic word segmentation processing on the DNA sequence information according to a preset vocabulary of a set size, and then obtain two tokenized subsequences according to the length and occurrence frequency of the words;

[0012] Extract DNA sequence features from all the tokenized subsequences of the data sample;

[0013] Fuse the known genomic features of the data sample with the DNA sequence features to generate combined features;

[0014] Based on the relationship between the combined features and the chromatin interaction prediction results of each sub-model in the ensemble learning model, obtain the chromatin interaction prediction results of each sub-model and take the average to obtain the final chromatin interaction prediction result.

[0015] As an implementation, the generation process of the preset vocabulary of a set size includes:

[0016] Initialize the vocabulary based on all the unique characters in the DNA sequence information;

[0017] In each iteration, the most frequent character segment is selected as a new word and added to the vocabulary. Then, the corpus is updated by replacing all identical segments with the new word. The iteration continues until the number of words in the vocabulary is reached.

[0018] As an implementation, the process of extracting DNA sequence features from all tokenized subsequences of the data sample is as follows:

[0019] Convert each tokenized subsequence into a corresponding tokenized subsequence feature vector;

[0020] Pass the tokenized subsequence feature vector through the efficient channel attention module to determine the attention weights of each tokenized subsequence feature vector, and then convert it into DNA sequence features through the embedding layer.

[0021] As an implementation, the expression of the efficient channel attention module is:

[0022] Attention(F) = σ(Conv1D k (GAP(F))) ⊙ F

[0023] Where: Attention is the efficient channel attention module function; F is the tokenized subsequence feature vector, GAP is the global average pooling operation, Conv1D k represents a one-dimensional convolution operation with a kernel size of k, σ is the Sigmoid activation function, and ⊙ represents element-wise multiplication.

[0024] As an implementation, the expression of the one-dimensional convolution operation kernel size k is:

[0025]

[0026] Where: C is the number of channels of the input feature, γ and b are hyperparameters; || odd represents rounding the calculation result k down to the nearest odd number.

[0027] As an implementation, the method for predicting chromatin interactions based on dynamic word segmentation and word embedding further includes: calculating the importance score Score of the motif m of the data sample according to the chromatin interaction prediction result m as:

[0028]

[0029] △ m = P - P' m

[0030]

[0031]

[0032]

[0033]

[0034] Among them, P represents the original prediction result of the ensemble learning model for the data sample; P' m represents the prediction result obtained by using the same ensemble learning model after removing the information of the specific motif m from the data sample; △ m represents the prediction result deviation; M represents all motifs; C m represents the average number of times the motif m appears in a sample; C i represents the number of times the motif i appears in a sample; l m represents the length of the motif m, f m represents the proportion of the motif m among all motifs that appear in a sample; NUM s represents the total number of samples, NUM M represents the total number of types of motifs; C M[j] is the average number of times the j-th type of motif appears in a sample; l M[j] is the total length of the j-th type of motif; e is the natural constant, and α is the correction factor.

[0035] The second aspect of the present invention provides a chromatin interaction prediction system based on dynamic word segmentation and word embedding.

[0036] A chromatin interaction prediction system based on dynamic word segmentation and word embedding, comprising:

[0037] A tokenized subsequence conversion module, which is used to obtain the DNA sequence information of the data sample, perform dynamic word segmentation processing on the DNA sequence information according to a set size vocabulary, and then obtain two tokenized subsequences according to the length and occurrence frequency of the words;

[0038] A DNA sequence feature extraction module, which is used to extract DNA sequence features from all the tokenized subsequences of the data sample;

[0039] A joint feature generation module, which is used to fuse the known genomic features of the data sample with the DNA sequence features to generate joint features;

[0040] An interaction prediction module, which is used to obtain the chromatin interaction prediction results of each sub-model based on the relationship between the joint features and the chromatin interaction prediction results of each sub-model in the ensemble learning model, and take the average to obtain the final chromatin interaction prediction result.

[0041] The third aspect of the present invention provides a computer-readable storage medium.

[0042] A computer-readable storage medium stores a computer program thereon, and when the program is executed by a processor, it implements the steps in the chromatin interaction prediction method based on dynamic word segmentation and word embedding as described above.

[0043] The fourth aspect of the present invention provides a computer program product.

[0044] A computer program product includes a computer program / instructions, and when the computer program / instructions are executed by a processor, it implements the steps in the chromatin interaction prediction method based on dynamic word segmentation and word embedding as described above.

[0045] The fifth aspect of the present invention provides an electronic device.

[0046] An electronic device includes a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, it implements the steps in the chromatin interaction prediction method based on dynamic word segmentation and word embedding as described above.

[0047] Compared with the prior art, the beneficial effects of the present invention are as follows:

[0048] The present invention uses dynamic word segmentation and word embedding to initially extract sequence information, proposes a method for extracting the first several words of the positive / negative chain based on the length of sequence word segmentation and word frequency, and significantly improves the utilization rate of sequence information through bidirectional feature fusion; moreover, the combined features generated by fusing the known genomic features and DNA sequence features of the data samples are processed by each sub-model in the integrated learning model to obtain the relationship between the corresponding chromatin interaction prediction results, and then the chromatin interaction prediction results of each sub-model are averaged to obtain the final chromatin interaction prediction result, thereby improving the accuracy of the chromatin interaction prediction result.

[0049] The advantages of the additional aspects of the present invention will be partially given in the following description, partially will become obvious from the following description, or will be understood through the practice of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS

[0050] The specification drawings constituting a part of the present invention are used to provide a further understanding of the present invention. The schematic embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute an improper limitation to the present invention.

[0051] Figure 1 It is the schematic diagram of the chromatin interaction prediction method based on dynamic word segmentation and word embedding in the embodiment of the present invention;

[0052] Figure 2Performance of the chromatin interaction prediction method according to the embodiments of the present invention and two other existing sequence-based methods (SPEID and IChrom-Deep) on three cell lines, namely K562, IMR90, and GM12878;

[0053] Figure 3 The chromatin interaction prediction method according to the embodiments of the present invention compares Inter-Chrom and IChrom-Deep under three different experimental conditions respectively;

[0054] Figure 4 Effect of different input data combinations on the performance of the model in the present invention;

[0055] Figure 5 Effect of different DNA sequence mutation frequencies on the model prediction performance;

[0056] Figure 6 Ranking of important motifs in sequence 1, sequence 2, and both sequences in the samples of the K562 cell line dataset under the condition of using SPEID and the new motif importance calculation method proposed in the present invention respectively;

[0057] Figure 7 Relationship between the proportion of motif importance scores, the proportion of motif lengths, and the correction factor obtained by two calculation methods before and after adding the correction factor;

[0058] Figure 8 Select the top 20% of motifs to analyze their specific performance in different cell lines;

[0059] Figure 9 Flowchart of the training process of the ensemble learning model according to the embodiments of the present invention;

[0060] Figure 10 Actual impact after modifying the motif according to the embodiments of the present invention;

[0061] Figure 11 Flowchart of the chromatin interaction prediction method based on dynamic word segmentation and word embedding according to the embodiments of the present invention;

[0062] Figure 12 Schematic diagram of the structure of the chromatin interaction prediction system based on dynamic word segmentation and word embedding according to the embodiments of the present invention. Detailed implementation manners

[0063] The present invention will be further described below in conjunction with the accompanying drawings and embodiments.

[0064] It should be noted that the following detailed descriptions are all illustrative and are intended to provide further explanations of the present invention. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which the present invention belongs.

[0065] It should be noted that the terms used herein are only for describing specific embodiments and are not intended to limit the exemplary embodiments according to the present invention. As used herein, unless the context clearly indicates otherwise, the singular forms are also intended to include the plural forms. In addition, it should be understood that when the terms "comprising" and / or "including" are used in this specification, they specify the presence of features, steps, operations, devices, components, and / or their combinations.

[0066] Example 1

[0067] The embodiment of the present invention adopts a benchmark dataset derived from TargetFinder2.0. This dataset covers four different cell lines, K562, HeLa-S3, IMR90, and GM12878. All samples contain genomic features, k-mer frequencies, genomic coordinates of chromatin binning, sequence conservation scores, CTCF binding motifs, and distances between chromatin binning. After filtering using chromatin states, the ratio of positive samples (with interactions) to negative samples (without interactions) in each cell line is strictly limited to 1:10. In addition, the HeLa-S3 cell line is discarded because the number of its filtered samples is too small (only 66 positive samples are retained). Finally, specific sequence data is retrieved using BEDTools according to genomic coordinates.

[0068] In the embodiment of the present invention, as Figure 11 shown, a chromatin interaction prediction method based on dynamic word segmentation and word embedding is provided, which includes:

[0069] Step S101: Obtain the DNA sequence information of the data sample, perform dynamic word segmentation processing on the DNA sequence information according to a set-size vocabulary, and then obtain two tokenized subsequences according to the length and occurrence frequency of the words.

[0070] Step S102: Extract DNA sequence features from all the tokenized subsequences of the data sample.

[0071] Step S103: Fuse the known genomic features of the data sample with the DNA sequence features to generate combined features.

[0072] Step S104: Based on the relationship between the combined features and the chromatin interaction prediction results of each sub-model in the ensemble learning model, obtain the chromatin interaction prediction results of each sub-model and take the average to obtain the final chromatin interaction prediction result.

[0073] In this embodiment, in step S101, tokenization of DNA sequences is performed using SentencePiece and Byte Pair Encoding (BPE).

[0074] Among them, SentencePiece, as a language-independent tokenization tool, can regard the input as the original data stream without presupposing any tokenization structure, which is very suitable for processing genomic sequences without clear word and sentence boundaries. BPE, as a compression algorithm, has been widely adopted as a tokenization strategy in the field of natural language processing. It constructs a variable-length token vocabulary of a fixed size by learning the character co-occurrence frequency.

[0075] Table 1 Process of constructing a vocabulary from a specific corpus using BPE

[0076]

[0077] Table 1 outlines the process of constructing a vocabulary from a specific corpus using BPE.

[0078] Specifically, in step S101, the process of generating a vocabulary of a set size includes:

[0079] Initializing the vocabulary based on all unique characters in the DNA sequence information;

[0080] In each iteration, the most frequent character segment is selected as a new word, added to the vocabulary, and then the corpus is updated by replacing all the same segments with this new word. The iteration continues until the number of words set in the vocabulary is reached.

[0081] The present invention selects the word representation vector table trained with a vocabulary size of 4096 (2 12 ) as the initial embedding vector after the sequence is converted into words, and obtains the feature representation most suitable for the chromatin interaction prediction task through retraining.

[0082] In the preliminary processing stage of the sequence, a fixed-size vocabulary is constructed based on the co-occurrence frequency of words. To make more effective use of this information and use it as a convenient input for model training, each sequence is screened according to the word length and occurrence frequency, as Figure 1As shown in a of , the top 500 words in terms of both length and frequency were selected, while maintaining their original order. During the process of constructing the vocabulary, due to the nature of the iterative mechanism, it can be inferred that those longer words obtained through iteration represent, to some extent, the unique elements of the sequence, or the specificity of the sequence. Words with higher frequencies, on the other hand, reflect more general characteristics. Therefore, the subsequences extracted from these two dimensions have different characteristics and advantages in expression. In addition, in order to comprehensively capture the information of the sequence, this processing step was performed on both the forward and reverse directions of the sequence. Overall, each sample originally contained two DNA sequences, but through the forward and reverse processing and obtaining sequences from different perspectives, each sample contains two sequences, and each sequence is obtained in two directions (forward and reverse), and then these sequences are screened from two perspectives of word length and frequency of occurrence to obtain two subsequences. Each sample is finally processed into containing 8 (2 3 ) subsequences composed of different words as input.

[0083] The data preprocessing mainly includes tokenizing the DNA sequences using SentencePiece and BPE to generate a vocabulary of a fixed size. Through the word embedding technology of DNABERT, the DNA sequences are converted into vector representations for subsequent deep learning models.

[0084] In step S102, as Figure 1 shown in b of , the process of extracting DNA sequence features from all tokenized subsequences of the data sample is as follows:

[0085] Step S1021: Convert each tokenized subsequence into a corresponding tokenized subsequence feature vector;

[0086] Step S1022: Pass the tokenized subsequence feature vector through an Efficient Channel Attention module (ECA module) to determine the attention weights of each tokenized subsequence feature vector, and then convert it into DNA sequence features through an embedding layer.

[0087] Among them, the expression of the Efficient Channel Attention module is:

[0088] Attention(F) = σ(Conv1D k (GAP(F))) ⊙ F

[0089] Where: Attention is the function of the Efficient Channel Attention module; F is the tokenized subsequence feature vector, GAP is the global average pooling operation, Conv1D k represents a one-dimensional convolution operation with a kernel size of k, σ is the Sigmoid activation function, and ⊙ represents element-wise multiplication.

[0090] ECA module: Obtain the global average value of each channel through Global Average Pooling (GAP). Use a one-dimensional convolutional kernel size k calculated adaptively to convolve the features to learn the weights of each channel. Apply the Sigmoid activation function to obtain the attention weights of each channel. Multiply the attention weights by the original feature map to enhance important features and suppress unimportant features.

[0091] The expression for the one-dimensional convolutional operation kernel size k is:

[0092]

[0093] where: C is the number of channels of the input features, γ and b are hyperparameters; || odd denotes rounding the calculation result k down to the nearest odd number to ensure that the convolutional kernel size is odd.

[0094] The advantages of the ECA module are its high computational efficiency, preservation of information integrity, adaptive kernel size, and ease of integration. These features enable the ECA module to be easily embedded into any existing CNN architecture, bringing performance improvement to the model without major modifications to the original network architecture.

[0095] In step S103, fuse the DNA sequence features (ECA output) with the genomic features (conservation score, CTCF motif, distance, etc.) to generate a joint feature matrix through a fully connected layer. Among them, the genomic features are provided in the dataset, and each sample contains a DNA sequence and the corresponding genomic features.

[0096] In step S104, during the process of training the ensemble learning model, as Figure 9 shown, to solve the problem that the model is difficult to learn the information of minority classes due to the imbalance of the training set, first calculate the ratio of negative samples to positive samples in the training set, divide the negative samples into multiple subsets according to this ratio, and use each negative sample subset and all positive samples to construct multiple submodels. Finally, take the average of multiple prediction results to obtain the final prediction result.

[0097] The ensemble learning model here consists of several submodels such as Model 1, Model 2,... Model N, and each submodel is trained according to the corresponding sample set. Among them, Model 1, Model 2,... Model N can all select existing neural network models according to the actual situation.

[0098] In this embodiment, known sequence motifs in the HOCOMOCO Human v11 database are used as a reference. First, the position weight matrix (PWM) of the motif is extracted from the file and used to compare with the subsequences in the window during the sliding match on the sequence data. Assuming the length of the motif is Lm, the PWM is constructed as a matrix with L m rows and four columns, and the values of each column correspond to each base (A, C, G, and T) respectively. During the matching process, first, each sequence with a length of L s is segmented into a series of subsequences with a length of L m and a step size of 1. Therefore, for each sequence, L s -L m +1 matches are performed, then the PWM scores of each base are read and summed to obtain the total matching score of the subsequence. This score is then compared with a preset threshold to determine whether the subsequence matches the motif. The matching evaluation criteria for the subsequence are defined by the following equation:

[0099]

[0100] where the variable j is used to represent the four bases in the subsequence fragment, taking values of 0, 1, 2, and 3, corresponding to A, C, G, and T respectively. Q represents the matching score of the subsequence. The P-value threshold for each motif is set. If the matching score Q of a subsequence fragment exceeds this threshold P, it is considered that this fragment matches the corresponding motif.

[0101] The chromatin interaction prediction method based on dynamic word segmentation and word embedding further includes: according to the chromatin interaction prediction result, calculating the importance score Score m of the motif m in the data sample as:

[0102]

[0103] △ m = P - P' m

[0104]

[0105]

[0106]

[0107]

[0108] where P represents the original prediction result of the ensemble learning model for the data sample; P' m represents the prediction result obtained by using the same ensemble learning model after removing the information of the specific motif m from the data sample; △m Indicates the prediction result deviation; M represents all motifs; C m represents the average number of occurrences of motif m in a sample; C i represents the number of occurrences of motif i in a sample; l m represents the length of motif m, f m represents the proportion of motif m among all motifs that appear in a sample; NUM s represents the total number of samples, NUM M represents the total number of motif types; C M[j] is the average number of occurrences of the j-th type of motif in a sample; l M[j] is the total length of the j-th type of motif; e is the natural constant, and α is the correction factor.

[0109] When using computer mutagenesis technology and finding the motif to be modified through matching, there is a challenge: it is impossible to determine whether some parts of the subsequence contain important information after combining with other content. When the occurrence frequency of the motif is low, the increase in the corresponding number of occurrences in the experiment may actually lead to the destruction of more information, as Figure 10 shown. Therefore, a certain degree of suppression is required at this time. However, when the motif frequency is much greater than , the meaning represented by this situation cannot be ignored either, and its proportion in the importance score should be increased.

[0110] The formula here actually borrows a part of the entropy calculation formula in information theory. In the field of information theory, the higher the value of entropy, the greater the uncertainty in the information; the lower the value of entropy, the smaller the uncertainty. The original entropy H(X) calculation formula is defined as follows:

[0111]

[0112] This formula shows that when the probabilities of various events occurring are more uniform, the calculation result on the left side of the formula is higher; on the contrary, when the probability differences are greater, it is lower. When converting the concept of events to elements, it is just the opposite of the calculation logic of expectation. It is considered that when the occurrence frequency of an element x i is more different from that of other elements, it should be given a higher weight. Therefore, the calculation part of p(x i )log p(x i ) is selected. Through the correction factor α, the formula is adjusted so that when the proportion f m is exactly equal to , the minimum value is obtained on the left side, and being larger or smaller can improve the importance score of this motif to a certain extent.

[0113] The method for calculating the importance of the motif can actually be applied in more scenarios, not limited to the 401 motifs in this embodiment. Essentially, it simultaneously considers the proportion of a certain motif in the overall content and its unit impact, and can be directly applied when calculating the importance of different numbers of motifs. Even not only for elements like motifs, as long as the calculation logic is satisfied, it can be modified into a calculation formula applicable to specific scenarios based on this.

[0114] The present invention has developed a chromatin interaction prediction model and a motif importance quantification method based on dynamic word segmentation and DNABERT word embedding, which analyzes and extracts sequence features from multiple perspectives and then predicts whether there is an interaction based on the features.

[0115] First, the framework Inter-Chrom proposed by the present invention was compared with two other existing sequence-based methods (SPEID and IChrom-Deep) in terms of performance. Through a unified training strategy, the performance of these methods on three cell lines, namely K562, IMR90, and GM12878, was re-evaluated, as Figure 2 shown. The results show that the prediction results of SPEID are almost random, which verifies the view of Cao et al. that its performance is overestimated due to overfitting. IChrom-Deep has made obvious progress in considering and attempting to solve this problem, but there is still a significant gap in its ability to extract information from DNA sequences compared with Inter-Chrom. Inter-Chrom shows the best performance among the given evaluation metrics on the three cell lines, demonstrating stronger stability.

[0116] The ability of Inter-Chrom to predict chromatin interactions when crossing cell lines was further verified. Inter-Chrom and IChrom-Deep were compared under three different experimental conditions, and the results are as Figure 3 shown.

[0117] Condition 1: The same number of samples are used for training and testing, and the chromosome splitting strategy is strictly implemented while crossing cell lines. Figure 3 In a of , the values on the diagonal can be seen to be significantly higher than those in other positions, indicating that the ability of IChrom-Deep to predict across cell lines is weak under this condition. The reason for this result is that the accuracy of IChrom-Deep prediction is greatly affected by chromosomes. Especially when training with data from the GM12878 cell line and testing on the other two cell lines, it is affected by many groups of invalid results. Figure 3 b in shows the results of Inter-Chrom under the same experimental conditions. Not only are there no cases of invalid predictions, but the values of valid results are significantly higher than those of IChrom-Deep. Figure 3The diagonal in b is slightly higher than other values, which also reflects the specificity of chromatin interactions between cell lines.

[0118] Condition 2: According to the size of the dataset in different cell lines, different numbers of samples are used for training and testing. The chromosome splitting strategy is also used, as Figure 3 shown in c and d. After comparison with Figure 3 a and b, it is found that the larger the training data volume, the better the prediction performance of the finally obtained model.

[0119] Condition 3: Different numbers of samples are used for training and testing, and general ten-fold cross-validation is performed without considering the grouping of chromosomes, as Figure 3 shown in e and f. The results are significantly higher than the above two groups, which further verifies that the prediction results of chromatin interactions are closely related to chromosomes. This significantly exaggerated result also illustrates the necessity of the chromosome splitting strategy.

[0120] The influence of different input data combinations on the model performance in the present invention is verified, as Figure 4 shown. The results show that under the same sequence processing module, due to the superiority of the feature processing module, single data input can already achieve better prediction effects than other methods. However, adding more feature inputs on the basis of the model in the present invention can still further improve the performance. This also shows that if limited by conditions such as computing resources and time, using a part of the sequence module as input can also obtain near-optimal results. Even if only a part of the forward or reverse sequence is used for prediction, satisfactory results can still be achieved. Finally, when genomic signals are introduced as additional features, the model performance is significantly improved, which confirms the importance of genomic signals in the task of predicting chromatin interactions.

[0121] Analyze the influence of different DNA sequence mutation frequencies on the model prediction performance to verify the robustness of the model, as Figure 5 shown. Three main mutation types including base substitution, deletion mutation, and insertion mutation are considered. Since IChrom-Deep can only accept sequences with a fixed length of 5000bp as input, only base substitution is considered for it, and four groups of experiments with perturbation rates of 1%, 2%, 5%, and 10% are set. For Inter-Chrom, two additional perturbation methods of deletion and insertion are added, and the overall perturbation rate is consistent with the previous two parts. From Figure 5 the two evaluation indicators given, it can be seen that at the four frequencies, the prediction results of the three groups of experiments have similar changing trends, and the two groups of Inter-Chrom are both higher than IChrom-Deep. This shows that when the DNA sequence is perturbed to a certain extent, Inter-Chrom still has advantages in predicting chromatin interactions.

[0122] The impact of motif variations in the input sequence on predicting chromatin interaction results was explored. For a specific designated motif, subsequence fragments were obtained through matching and then replaced with random noise, which is equivalent to erasing the information of this motif in the sequence while collecting the count information of the motif.

[0123] In the specific experimental operation, first, it was trained on the unmodified DNA sequence, and then sequences with specific motif information removed were used for prediction, obtaining 401 different result changes corresponding to the cases of changing 401 motifs respectively. The first two groups of the experiment were operations on the sequences of two chromatin bins in the interaction pair, and the third group was the situation when both sequences were modified according to the designated motif. Taking Figure 5 the K562 cell line shown in as an example, the motif importance rankings and their changes obtained by using the method of calculating importance scores in SPEID and the method proposed in the research of the present invention were observed respectively.

[0124] Figure 6 a, b, and c in show the rankings of only changing sequence 1, only changing sequence 2, and changing sequence 1 and sequence 2 respectively; Figure 6 d, e, and f in show another manifestation corresponding to the three groups of experimental results of the K562 cell line. From the inside out in the circular ring are: the average number of occurrences of a specific motif multiplied by the motif length, and then the proportion P of this value in the total motifs is calculated; the correction factor F is calculated from this proportion; the final importance score S corresponding to this motif. Taking the top ten motifs with the highest importance scores as an example, by comparing the proportion changes of the same element, it can be found that when the motif proportion is too small or too large, the value of the correction factor can be increased. Figure 7 The motif importance rankings obtained by two calculation methods before and after adding the correction factor are respectively shown. Figure 7 In a of, it was found that the motifs with higher importance are concentrated in the part with the smallest proportion, and the motifs with the lowest importance in the last part of the circular ring have the largest proportion, which includes some motifs often discussed in research, such as SP1 at the last position, which plays a role in important aspects such as gene expression regulation and cell physiological function regulation. This contradiction shows that there are limitations in using only the index difference value divided by the occurrence frequency of the motif as the importance score. Figure 7 In b of, after using the correction factor, the importance ranking not only makes the motifs with small proportion and high importance more concentrated, but also improves the ranking of motifs with occurrence frequencies much greater than . Moreover, the conclusion that the motifs with higher importance scores have a very small proportion is very consistent both locally and globally. According to Figure 7For the first part spanned by the dark blue thin ring marked by b in [description], the difference between Ssum and Psum is very obvious. The proportion P of motifs corresponding to 62.4% of the total importance score of 401 motifs is only 6.7%. This result further verifies the necessity of considering the motif proportion factor when designing the formula, and suppresses the situation where the importance score of a motif that appears more frequently may be overestimated because it interferes too much with the roles of other potential motifs.

[0125] After sorting the importance scores of the motifs, the top 20% of the motifs were selected to analyze their specific performances in different cell lines, such as Figure 8 shown. The observation results show that the overall performances of most motifs are similar, and at the same time, there are also some individual motifs that show differences in specific cell lines or chromatin bins. Through the calculation method and sorting of the present invention, important motifs consistent with many previous studies have been identified.

[0126] Example Two

[0127] As Figure 12 shown, the embodiment of the present invention provides a chromatin interaction prediction system based on dynamic word segmentation and word embedding, which includes:

[0128] A tokenized subsequence transformation module 201, which is used to obtain the DNA sequence information of the data sample, perform dynamic word segmentation processing on the DNA sequence information according to a set vocabulary, and then obtain two tokenized subsequences according to the length and occurrence frequency of the words;

[0129] A DNA sequence feature extraction module 202, which is used to extract DNA sequence features from all the tokenized subsequences of the data sample;

[0130] A combined feature generation module 203, which is used to fuse the known genomic features of the data sample with the DNA sequence features to generate combined features;

[0131] An interaction prediction module 204, which is used to obtain the chromatin interaction prediction results of each sub-model based on the relationship between the combined features and the chromatin interaction prediction results of each sub-model in the integrated learning model, and take the average to obtain the final chromatin interaction prediction result.

[0132] It should be noted here that each module in the embodiment of the present invention corresponds to each step in Example One, and the specific implementation process is the same, so it will not be elaborated here.

[0133] Example Three

[0134] This embodiment provides a computer-readable storage medium, on which a computer program is stored. When the program is executed by a processor, it implements the steps in the chromatin interaction prediction method based on dynamic word segmentation and word embedding as described above.

[0135] Embodiment 4

[0136] A computer program product includes a computer program / instructions. When the computer program / instructions are executed by a processor, they implement the steps in the chromatin interaction prediction method based on dynamic word segmentation and word embedding as described above.

[0137] Embodiment 5

[0138] This embodiment provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, it implements the steps in the chromatin interaction prediction method based on dynamic word segmentation and word embedding as described above.

[0139] The present invention is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products of the embodiments of the present invention. It should be understood that each process and / or block in the flowcharts and / or block diagrams, as well as the combination of processes and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate a device for implementing the specified functions in one process Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks.

[0140] The above are only the preferred embodiments of the present invention and are not used to limit the present invention. For those skilled in the art, the present invention can have various changes and modifications. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.

Claims

1. A chromatin interaction prediction method based on dynamic word segmentation and word embedding, characterized in that: include: Acquire DNA sequence information of the data sample, perform dynamic word segmentation on the DNA sequence information according to a set size vocabulary, and then obtain two tokenized subsequences according to the length and frequency of occurrence of the words; Extract DNA sequence features from all tokenized subsequences of the data sample; Merging the known genome features of the data sample with the DNA sequence features to generate a joint feature; Based on the relationship between the joint features and the chromatin interaction prediction results of each sub-model in the integrated learning model, the chromatin interaction prediction results of each sub-model are obtained and averaged to obtain the final chromatin interaction prediction result.

2. The chromatin interaction prediction method based on dynamic word segmentation and word embedding according to claim 1, characterized in that: The process of generating the vocabulary of the specified size includes: Initialize the vocabulary based on all unique characters in the DNA sequence information; In each iteration, the most frequent character segment is selected as a new word and added to the vocabulary. All identical segments are then replaced with the new word to update the corpus. This process continues until the number of words set in the vocabulary is reached.

3. The chromatin interaction prediction method based on dynamic word segmentation and word embedding according to claim 1, characterized in that: The process of extracting DNA sequence features from all tokenized subsequences of a data sample is: Convert each tokenized subsequence into a corresponding tokenized subsequence feature vector; The tokenized subsequence feature vectors are passed through an efficient channel attention module to determine the attention weights of each tokenized subsequence feature vector, and then converted into DNA sequence features through an embedding layer.

4. The chromatin interaction prediction method based on dynamic word segmentation and word embedding according to claim 3, characterized in that: The expression of the efficient channel attention module is: Attention(F)=σ(Conv1D k (GAP(F)))⊙F Among them: Attention is an efficient channel attention module function; F is the tokenized subsequence feature vector, GAP is the global average pooling operation, Conv1D k represents a one-dimensional convolution operation with a kernel size of k, σ is the Sigmoid activation function, and ⊙ represents element-wise multiplication.

5. The chromatin interaction prediction method based on dynamic word segmentation and word embedding according to claim 4, characterized in that: The expression of the kernel size k of a one-dimensional convolution operation is: Where: C is the number of channels of the input feature, γ and b are hyperparameters; || odd Indicates that the calculation result k is rounded down to the nearest odd number.

6. The chromatin interaction prediction method based on dynamic word segmentation and word embedding according to claim 1, characterized in that: Also includes: According to the chromatin interaction prediction results, calculate the importance score of the motif m of the data sample m for: △ m =P-P′ m Among them, P represents the original prediction result of the integrated learning model for the data sample; P′ m Represents the prediction results obtained using the same ensemble learning model after removing the information of a specific motif m from the data sample; △ m Indicates the prediction result deviation; M represents all motifs; C m represents the average number of times motif m appears in a sample; C i Represents the number of times motif i appears in a sample; l m represents the length of motif m, f m Represents the proportion of motif m in all motifs appearing in a sample; NUM s Represents the total number of samples, NUM M Represents the total number of motif types; C M[j] is the average number of occurrences of the j-th motif in a sample; l M[j] is the total length of the jth motif; e is a natural constant, and α is a correction factor.

7. A chromatin interaction prediction system based on dynamic word segmentation and word embedding, characterized in that: include: A tokenized subsequence conversion module is used to obtain DNA sequence information of a data sample, perform dynamic word segmentation on the DNA sequence information according to a set large and small vocabulary, and then obtain two tokenized subsequences according to the length and frequency of occurrence of the words; A DNA sequence feature extraction module, which is used to extract DNA sequence features from all the tokenized subsequences of the data sample; A joint feature generation module, which is used to fuse the known genome features of the data sample with the DNA sequence features to generate a joint feature; The interaction prediction module is used to obtain the chromatin interaction prediction results of each sub-model based on the relationship between the joint features and the chromatin interaction prediction results of each sub-model in the integrated learning model, and take the average to obtain the final chromatin interaction prediction result.

8. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the steps in the chromatin interaction prediction method based on dynamic word segmentation and word embedding as described in any one of claims 1 to 6 are implemented.

9. A computer program product comprising a computer program / instructions, characterized in that When the computer program / instructions are executed by a processor, the steps in the chromatin interaction prediction method based on dynamic word segmentation and word embedding as described in any one of claims 1 to 6 are implemented.

10. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the program, the steps in the chromatin interaction prediction method based on dynamic word segmentation and word embedding as described in any one of claims 1 to 6 are implemented.