Method for predicting proportion of cells in different cycles based on chromatin conformation spectrum
Through the cell cycle proportion prediction model based on deep neural network model, the problem of lack of cell cycle proportion estimation in Bulk Hi-C technology is solved, and fast and accurate cell cycle proportion prediction is achieved to ensure the reliability of the Bulk Hi-C test results.
Patent Information
- Application Number
- CN202510944216.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-09
- Publication Date
- 2025-08-08
- Estimated Expiration
- 2045-07-09
AI Technical Summary
The lack of estimation methods for estimating the proportion of cells in different cell cycles in the sample in existing Bulk Hi-C technology has affected the reliability of the test results.
A cell cycle proportion prediction model based on the deep neural network model is adopted, and the cell cycle proportion prediction results are obtained by training samples and pre-processing of Hi-C data.
The proportion of cells in the Bulk Hi-C data samples with different cell cycles is quickly and accurately estimated, ensuring the reliability of the detection results, and avoiding the average three-dimensional conformational interference and noise of data analysis results caused by the mixing of cells in different cycles.
Smart Images

Figure CN120452554A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of biological detection technology, and in particular to a method for predicting the proportion of cells in different cell cycles based on chromatin conformation profiles. Background Art
[0002] Hi-C is a genomics technique that combines high-throughput sequencing with chromatin conformation capture. It is used to study the spatial relationships of chromatin DNA across the entire genome, obtaining high-resolution three-dimensional chromatin structural information. Bulk Hi-C, on the other hand, is a tissue-sample-level Hi-C technique based on millions of cells. Therefore, it can capture the average three-dimensional conformation of a mixture of cells at different cell cycles.
[0003] In Bulk Hi-C studies, cells are synchronized during the cell cycle, maintaining them in the G0 / G1 phase. This prevents interference from cells in different cycles, particularly those in the G2 / M phase, in the final results. This is because cells in the G2 / M phase have highly condensed chromatin, which is distinct from the loose chromatin of cells in the G0 / G1 phase. Cells in this phase have low transcriptional activity, and allowing a large number of cells in the G2 / M phase to remain in the tissue can introduce noise into the data analysis results and potentially lead to erroneous conclusions.
[0004] Therefore, estimating the proportion of cells at different cell cycles in the sample from which tissue bulk Hi-C data originate is a key step in tissue bulk Hi-C data quality control and is crucial for ensuring the reliability of bulk Hi-C experimental results. However, there is currently a lack of methods for estimating the proportion of cells at different cell cycles based on bulk Hi-C data itself. Summary of the Invention
[0005] This application provides a method for predicting the proportion of cells in different cell cycles based on chromatin conformation profiles, which is used to solve the technical problem in existing Bulk Hi-C technology that lacks a method to estimate the proportion of cells in different cell cycles in samples derived from Bulk Hi-C data, affecting the reliability of detection results.
[0006] According to the first aspect disclosed in the present application, the present application provides a method for predicting the proportion of cells in different cell cycles based on chromatin conformation profiles, comprising: Obtain bulk Hi-C data to be tested; The bulk Hi-C data to be tested is input into a cell cycle ratio prediction model to obtain a cell cycle ratio prediction result output by the cell cycle ratio prediction model; wherein the cell cycle ratio prediction model is obtained based on training a deep neural network model.
[0007] In one feasible implementation, training a deep neural network model includes: Acquire multiple training samples and construct a training set; wherein the training samples are annotated with true values, and the true values include the true proportions of cells in different cell cycles in the training samples; Inputting the training samples contained in the training set into the deep neural network model to obtain a predicted value; wherein the predicted value includes the predicted proportion of cells in different cell cycles in the training samples; Based on a preset loss function, obtaining the loss between the predicted value and the true value in the training set; The network parameters of the deep neural network model are optimized based on the loss amount until the loss function converges or reaches a preset number of training iterations, thereby obtaining a cell cycle ratio prediction model.
[0008] In a feasible implementation, the deep neural network model includes a first DNN model, a second DNN model, and a third DNN model, and inputting the training samples included in the training set into the deep neural network model to obtain a predicted value includes: Inputting the training sample into a first DNN model to obtain a first prediction ratio; Inputting the training sample into a second DNN model to obtain a second prediction ratio; Inputting the training sample into a third DNN model to obtain a third prediction ratio; The predicted proportions of cells in different cell cycles in the first predicted proportion, the second predicted proportion and the third predicted proportion are averaged to obtain a predicted value.
[0009] In a feasible implementation, the first DNN model includes a first input layer, a first first-round linear layer, a first second-round linear layer, a first third-round linear layer, a first fourth-round linear layer, and a first fifth-round linear layer. Inputting the training sample into the first DNN model to obtain a first prediction ratio includes: Extracting features from the training sample through the first input layer to obtain first input extracted features of dimension 1000; The feature dimension of the first input extracted features is reduced to 256 by the first round of linear layers, and then activated by the ReLu function to obtain the first round of extracted features; The feature dimensions of the first round of extracted features are reduced to 128 by the first and second rounds of linear layers, and then activated by the ReLu function to obtain the first and second rounds of extracted features; The feature dimensions of the features extracted from the first and second rounds are reduced to 64 by the first three rounds of linear layers, and then activated by the ReLu function to obtain the first and third rounds of extracted features; The feature dimensions of the features extracted from the first three rounds are reduced to 32 by the first four rounds of linear layers, and then activated by the ReLu function to obtain the first four rounds of extracted features; The feature dimensions of the features extracted in the first four rounds are reduced to 4 through the first five rounds of linear layers, and then the features are normalized by the Softmax function to obtain the first prediction ratio.
[0010] In a feasible implementation, the second DNN model includes a second input layer, a second first-round linear layer, a second second-round linear layer, a second third-round linear layer, a second fourth-round linear layer, and a second fifth-round linear layer. Inputting the training sample into the second DNN model to obtain a second prediction ratio includes: Extract features from the training sample through the second input layer to obtain a second input extracted feature with a dimension of 1000; The feature dimension of the second input extracted features is reduced to 512 by the second round of linear layers, and then activated by the ReLu function to obtain the second round of extracted features; The feature dimension of the second round of extracted features is reduced to 256 by the second round of linear layers, and then activated by the ReLu function. Dropout is applied to randomly discard 30% of the parameters to obtain the second round of extracted features; The feature dimensions of the second-second round extracted features are reduced to 128 through the second-third round linear layer, and then activated by the ReLu function. Dropout is applied to randomly discard 20% of the parameters to obtain the second-third round extracted features; The feature dimensions of the features extracted in the second and third rounds are reduced to 64 by the second four rounds of linear layers, and then activated by the ReLu function. Dropout is applied to randomly discard 10% of the parameters to obtain the second four rounds of extracted features; The feature dimension of the features extracted in the second four rounds is reduced to 4 through the second five rounds of linear layers, and then the features are normalized by the Softmax function to obtain the second prediction ratio.
[0011] In a feasible implementation, the third DNN model includes a third input layer, a third first-round linear layer, a third second-round linear layer, a third third-round linear layer, a third fourth-round linear layer, and a third fifth-round linear layer. Inputting the training sample into the third DNN model to obtain a third prediction ratio includes: Extracting features from the training sample through the third input layer to obtain a third input extracted feature of dimension 1000; The feature dimension of the first input extracted feature is increased to 1024 by the third round of linear layer, and then activated by the ReLu function to obtain the third round of extracted features; The feature dimension of the third round of extracted features is reduced to 512 by the third round of linear layers, and then activated by the ReLu function, and 60% of the parameters are randomly discarded by Dropout to obtain the third round of extracted features; The feature dimension of the third round of extracted features is reduced to 256 by the third round of linear layer, and then activated by ReLu function, and 30% of the parameters are randomly discarded by Dropout to obtain the third round of extracted features; The feature dimension of the third-fourth round of extracted features is reduced to 128 by the third-fourth round of linear layers, and then activated by the ReLu function, and 10% of the parameters are randomly discarded by Dropout to obtain the third-fourth round of extracted features; The feature dimension of the features extracted in the third and fourth rounds is reduced to 4 by the third and fifth rounds of linear layers, and then the features are normalized by the Softmax function to obtain the third prediction ratio.
[0012] In one feasible embodiment, the method includes: Acquire a plurality of single-cell Hi-C data; wherein the single-cell Hi-C data is an interaction matrix; Preprocessing the single-cell Hi-C data to obtain preprocessed Hi-C data; wherein the preprocessing includes depth filtering and depth normalization; Multiple preprocessed Hi-C data are randomly selected for fusion to obtain pseudo Bulk Hi-C data as training samples.
[0013] According to the second aspect disclosed in the present application, the present application provides a device for predicting the proportion of cells in different cell cycles based on chromatin conformation profiles, comprising: Data acquisition module, used to obtain bulk Hi-C data to be tested; A ratio prediction module is used to input the bulk Hi-C data to be tested into a cell cycle ratio prediction model to obtain a cell cycle ratio prediction result output by the cell cycle ratio prediction model; wherein the cell cycle ratio prediction model is obtained based on training a deep neural network model.
[0014] According to a third aspect disclosed in the present application, the present application provides an electronic device, comprising a processor, and a memory communicatively connected to the processor; The memory stores computer-executable instructions; The processor executes the computer-executable instructions stored in the memory to implement the method according to any one of the first aspects.
[0015] According to the fourth aspect disclosed in the present application, the present application provides a computer-readable storage medium, in which computer-executable instructions are stored. When the computer-executable instructions are executed, they are used to implement any one of the methods in the first aspect.
[0016] According to a fifth aspect disclosed in the present application, the present application provides a computer program product, including a computer program, which, when executed, is used to implement any one of the methods in the first aspect.
[0017] Compared with the existing technology, this application has the following beneficial effects: This application provides a method for predicting the proportion of cells in different cell cycles based on chromatin conformation profiles. By training a deep neural network model, the cell cycle proportion prediction results of bulk Hi-C data can be obtained by simply inputting the trained cell cycle proportion prediction model into the bulk Hi-C data. The cell cycle proportion prediction model can quickly and accurately estimate the proportion of cells in different cell cycles in the sample data based on bulk Hi-C data, providing an effective means for organizing the quality control of bulk Hi-C data, helping to ensure the reliability of bulk Hi-C test results, avoiding the average three-dimensional conformation interference caused by the mixing of cells in different cell cycles and the noise in the data analysis results, thereby preventing erroneous conclusions and providing key technical support for ensuring the reliability of bulk Hi-C test results. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the present application and, together with the description, serve to explain the principles of the present application.
[0019] Figure 1 A schematic diagram of a process for predicting the proportion of cells in different cell cycles based on chromatin conformation profiles provided in an embodiment of the present application; Figure 2 A schematic diagram of a process for training a deep neural network model provided in an embodiment of the present application; Figure 3 A schematic diagram of a process for obtaining a prediction value through a deep neural network model provided in an embodiment of the present application; Figure 4 A schematic diagram of the structure of a deep neural network model provided in an embodiment of the present application; Figure 5 Schematic diagram of the interaction matrix for single-cell Hi-C data conversion provided in the embodiments of this application; Figure 6 Schematic diagram of comparative analysis of the predicted and true values of pseudo-bulk Hi-C data provided in the examples of this application; Figure 7 Schematic diagram of data analysis of predicted and true values of Bulk Hi-C data across tissue types provided in the embodiments of this application; Figure 8 A schematic diagram of the structure of a device for predicting the proportion of cells in different cell cycles based on chromatin conformation profiles provided in an embodiment of the present application; Figure 9 A schematic diagram of the structure of an electronic device provided in an embodiment of the present application.
[0020] The above drawings illustrate specific embodiments of the present application, which will be described in more detail below. These drawings and the textual description are not intended to limit the scope of the present application in any way, but rather to illustrate the concepts of the present application to those skilled in the art by reference to specific embodiments. DETAILED DESCRIPTION
[0021] Exemplary embodiments will be described in detail herein, with examples illustrated in the accompanying drawings. In the following description, when referring to the drawings, identical numerals in different figures represent identical or similar elements, unless otherwise indicated. The embodiments described in the following exemplary embodiments are not intended to represent all embodiments consistent with the present application. Rather, they are merely examples of apparatus and methods consistent with certain aspects of the present application, as detailed in the appended claims.
[0022] In mammalian cell nuclei, chromatin is clustered in micron-scale spaces, with each strand of chromatin occupying a specific nuclear compartment. The different nuclear compartments occupied by chromatin are highly correlated with gene density and transcriptional activity. Dynamic changes in chromatin structure are associated with many biological processes, particularly the cell cycle. Chromatin in animal cells undergoes vastly different structural deformations during each cell cycle, typically alternating between a highly condensed state that facilitates chromosome segregation and an interphase structure that accommodates transcription, gene silencing, and DNA replication.
[0023] Hi-C technology is a genomics technique that combines high-throughput sequencing with chromatin conformation capture. It is used to study the spatial relationships of chromatin DNA across the entire genome, obtaining high-resolution three-dimensional chromatin structural information. Hi-C technology can analyze high-level structural features of chromatin, such as A / B compartments, topologically associated domains (TADs), and chromatin loops, providing an important tool for understanding genome function, gene expression regulation, and disease mechanisms.
[0024] Bulk Hi-C is a common application of Hi-C technology, which involves performing Hi-C experiments on large cell populations to obtain population-level chromatin three-dimensional structural information. This technique is based on millions of cells and is performed at the tissue sample level. Therefore, it can capture the average three-dimensional conformation of a mixture of cells at different cell cycles. In some Bulk Hi-C studies, scientists synchronize the cell cycle, keeping all cells in the G0 / G1 phase to prevent cells in different cell cycles, particularly those in the G2 / M phase, from interfering with the final results.
[0025] This is because cells in the G2 / M phase have highly condensed chromatin, which is distinct from the loose chromatin of cells in the G0 / G1 phase. Cells in this phase have low transcriptional activity, and scientists studying transcriptional regulation are generally not interested in this state. Therefore, they try to keep cells in the G0 / G1 phase as much as possible. Allowing a large number of cells in the G2 / M phase in tissues will introduce noise into data analysis results and may lead researchers to draw erroneous conclusions.
[0026] In some normal adult tissue samples, the vast majority of cells are naturally in the G0 / G1 phase, allowing for direct bulk Hi-C experiments without synchronization. However, not all samples share these characteristics. Therefore, estimating the proportion of cells in different cell cycles within the sample from which tissue bulk Hi-C data originate is a key step in tissue bulk Hi-C data quality control and crucial for ensuring the reliability of bulk Hi-C results. However, methods for estimating the proportion of cells in different cell cycles based on bulk Hi-C data are currently lacking.
[0027] To address the above technical issues, this application proposes a method for predicting the proportion of cells in different cell cycles based on chromatin conformation profiles. Through the cell cycle proportion prediction model, the proportion of cells in different cell cycles in the Bulk Hi-C data sample can be quickly and accurately estimated based on Bulk Hi-C data, which helps to ensure the reliability of Bulk Hi-C test results.
[0028] The following detailed description of the technical solution of the method for predicting the proportion of cells in different cell cycles based on chromatin conformation profiles provided by this application is provided below through specific examples. It should be noted that the following examples can exist independently or in combination with each other, and the same or similar content may not be repeated in different examples.
[0029] Figure 1 A flow chart of a method for predicting the proportion of cells in different cell cycles based on chromatin conformation profiles provided in the present application is provided in the embodiment. Figure 1In some embodiments, the method for predicting the proportion of cells in different cell cycles based on chromatin conformation profiles comprises the following steps: S101, obtaining bulk Hi-C data to be tested.
[0030] Among them, the bulk Hi-C data to be tested are obtained for subsequent cell cycle ratio prediction.
[0031] S102, inputting the bulk Hi-C data to be tested into a cell cycle ratio prediction model to obtain a cell cycle ratio prediction result output by the cell cycle ratio prediction model; wherein the cell cycle ratio prediction model is obtained based on training a deep neural network model.
[0032] Among them, the bulk Hi-C data is input into the cell cycle ratio prediction model, and the cell cycle ratio prediction results can be output by the cell cycle ratio prediction model trained based on the deep neural network model.
[0033] Specifically, the cell cycle proportion prediction results include the proportion of cells in the four cycles of G0, G1, G2, and M. Among them, the G0 cycle represents cells that temporarily exit the cell cycle and stop cell division, but can re-enter the cell cycle under certain stimulation; the G1 cycle is the period from the completion of mitosis to DNA replication, which is the main stage of cell growth; the G2 cycle is the period from the completion of DNA replication to the beginning of mitosis; and the M cycle is the period when cells undergo mitosis. Through mitosis, cells evenly distribute the replicated genetic material into their two daughter cells.
[0034] In this example, by training the deep neural network model, the cell cycle ratio prediction results for bulk Hi-C data can be obtained by subsequently inputting the trained cell cycle ratio prediction model into the bulk Hi-C data. The cell cycle ratio prediction model can quickly and accurately estimate the proportion of cells in different cell cycles within the sample data based on bulk Hi-C data. This provides an effective means for quality control of bulk Hi-C data, helps ensure the reliability of bulk Hi-C test results, avoids interference with the average three-dimensional conformation caused by mixing cells at different cell cycles, and avoids noise in data analysis results, thereby preventing erroneous conclusions and providing key technical support for ensuring the reliability of bulk Hi-C test results.
[0035] exist Figure 1 In the method of predicting the proportion of cells in different cycles based on chromatin conformation spectrum shown in the figure, it is necessary to train the deep neural network model. Figure 2, further introduces the content of training the deep neural network model in the technical solution of the above-mentioned method for predicting the proportion of cells in different cycles based on chromatin conformation spectrum.
[0036] Figure 2 A schematic diagram of a process for training a deep neural network model provided in an embodiment of the present application is provided. Figure 2 In some embodiments, the process of training the deep neural network model includes the following steps: S201, obtaining multiple training samples and constructing a training set; wherein the training samples are labeled with true values, and the true values include the true proportions of cells in different cell cycles in the training samples.
[0037] Optionally, obtain multiple training samples and construct a training set, including: Step 1: Acquire multiple single-cell Hi-C data; wherein the single-cell Hi-C data is an interaction matrix.
[0038] Among them, Hi-C raw data (usually paired 150bp long sequences) are reflected as a pair of reads (read pairs) aligned to two different one-dimensional regions on the genome after sequence alignment and filtering. In order to better display and subsequent analysis of the results of these paired reads, Hi-C data are usually converted into an interaction matrix; see Figure 5 Shown is the interaction matrix of the first 10 kb region on chromosome 1 with an accuracy of 1 kb; The number 10 represents the number of paired reads connecting the 1-2kb region and the 6-7kb region of chromosome 1, that is, the unnormalized interaction strength is 10; 8 represents the number of paired reads connecting the 3-4kb region and the 8-9kb region of chromosome 1, that is, the unnormalized interaction strength is 8; because the interaction strength is affected by genomic sequence deviations such as GC content (referring to the proportion of guanine G and cytosine C bases in the genome sequence) and Mappability (the ability of sequencing reads to uniquely map to the genome reference sequence), the ICE (iterative correction and eigenvector decomposition) method is generally used to correct the above interference factors, which is called the in-sample normalization process of the original interaction matrix.
[0039] Step 2: Preprocess the single-cell Hi-C data to obtain preprocessed Hi-C data; wherein the preprocessing includes depth filtering and depth normalization.
[0040] Among them, the single-cell Hi-C data are preprocessed to filter out matrices with a depth greater than 500 k and a depth less than 100 k that is too shallow; depth refers to the total number of paired reads aligned to the genome; generally speaking, samples with a high total depth will show an overall increase in the interaction matrix interaction strength compared to samples with a low total depth, and this effect is considered to be noise caused by the deviation of the sequencing process and needs to be corrected; since samples with comparable depth are required to construct training and test sets, and these samples are depth-normalized so that their final interaction matrix data is not affected by the depth factor, too deep or shallow data will introduce huge deviations that cannot be corrected by conventional depth normalization, so they need to be filtered out and the data depth is normalized to the shallowest data level; normalization has two strategies: downsampling and upsampling. Here, a more robust downsampling method is adopted to randomly sample the effective paired reads of the single-cell Hi-C data, so that the number of effective paired reads, that is, the data depth, is reduced from more than 100,000 <100 k> to 100,000 <100 k>.
[0041] Step 3: Randomly select multiple preprocessed Hi-C data for fusion to obtain pseudo Bulk Hi-C data as training samples.
[0042] Multiple 500 kb matrices of preprocessed Hi-C data were randomly selected and fused together. This means the values of each corresponding grid in these matrices were summed to form a new interaction matrix, which served as the matrix for the pseudo-bulk Hi-C data. Within-sample ICE normalization was then performed. Because each single-cell Hi-C data sample has been normalized to the same depth, the depth of the pseudo-bulk Hi-C data should theoretically be roughly consistent, eliminating the need for depth normalization. Therefore, within-sample ICE normalization was performed to remove interference from genomic sequence biases such as GC content and mappability. Pseudo-bulk Hi-C data were generated by randomly mixing the single-cell Hi-C data for use as training samples.
[0043] Since the cell cycle of single-cell Hi-C data is known, and each pseudo-bulk Hi-C data comes from the mixture of multiple single-cell Hi-C data with known cell cycles, the proportion of cells in different cell cycles in the pseudo-bulk Hi-C data formed by mixing single-cell Hi-C data is also known.
[0044] Optionally, this step is repeated multiple times to eventually generate multiple pseudo BulkHi-C data of cells with different cell cycles, and use them to construct a training set.
[0045] S202, inputting the training samples contained in the training set into the deep neural network model to obtain a predicted value; wherein the predicted value includes the predicted proportion of cells in different cell cycles in the training samples.
[0046] Among them, after the training set is input into the deep neural network model, the deep neural network model will output a predicted value, which includes the predicted proportion of cells in different cell cycles in the training sample.
[0047] S203: Based on a preset loss function, obtain the loss between the predicted value and the true value in the training set.
[0048] Among them, the loss function can be used to obtain the loss between the predicted value and the true value, and then the loss is used to update the network parameters of the model, thereby gradually improving the model's cell composition ratio prediction performance.
[0049] Optionally, the loss function satisfies the following formula:
[0050] Among them, L1 represents the loss function, Indicates the predicted values, Indicates the A true value.
[0051] The absolute error loss, L1, is simple and intuitive. It simply calculates the absolute value of the difference between the predicted and true values and then averages them. This makes it easy to understand and implement, has low computational complexity, is insensitive to outliers, and more robustly reflects the model's predictive performance on mostly normal data, ensuring the stability of the model's overall predictions. Furthermore, its loss value aligns with the actual dimension of the error, making it easy to directly interpret the error magnitude. This provides an intuitive reference for evaluating and improving model performance, helping users quickly understand the model's predictions and make appropriate adjustments.
[0052] S204, optimizing the network parameters of the deep neural network model based on the loss amount until the loss function converges or reaches a preset number of training iterations, thereby obtaining a cell cycle ratio prediction model.
[0053] Among them, the deep neural network model is iteratively trained based on the collected training samples. The network parameters of the deep neural network model are continuously optimized according to the loss amount during the training process until the loss function converges or the preset number of training iterations is reached. The training of the model is terminated. At this time, a cell cycle ratio prediction model for predicting the proportion of cells in different cycles is obtained.
[0054] Specifically, the number of iterations refers to the number of times the model fully traverses the training dataset and updates the model parameters during training. A higher number of iterations can improve model accuracy because the model has more opportunities to learn features from the data. However, too many iterations can also lead to overfitting, meaning the model performs well on the training set but poorly on unseen data. Too few iterations can lead to underfitting, meaning the model does not fully learn the features of the data. Therefore, the number of training iterations should be set based on a combination of factors, such as the size of the training set and the complexity of the model.
[0055] Preferably, the training samples can be divided into a training set, a validation set and a test set.
[0056] Among them, the validation set is used to adjust the network parameters of the model and evaluate the performance of the model during training to avoid overfitting; the test set is used to finally evaluate the generalization ability of the model to ensure that the model can perform well on new data; these three together constitute the complete process of model training, evaluation and optimization.
[0057] In this embodiment, the cell cycle ratio prediction model obtained by training the deep neural network model through the above-mentioned training method can effectively detect the cell cycle ratio, and has the advantages of high detection efficiency and high accuracy.
[0058] exist Figure 2 In the method for training a deep neural network model shown in FIG, it is necessary to train the deep neural network model. Figure 3 , in the technical solution of the above-mentioned method for training a deep neural network model, the content of obtaining the prediction value through the deep neural network model is further introduced.
[0059] Figure 3 A schematic diagram of a process for obtaining a prediction value through a deep neural network model provided in an embodiment of the present application, see Figure 3 In some embodiments, combined with Figure 4 , the deep neural network model includes a first DNN model, a second DNN model and a third DNN model, and the process of obtaining a prediction value through the deep neural network model includes the following steps: S301: Input the training sample into the first DNN model to obtain a first prediction ratio.
[0060] Optionally, the first DNN model includes a first input layer, a first first-round linear layer, a first second-round linear layer, a first third-round linear layer, a first fourth-round linear layer, and a first fifth-round linear layer. Inputting the training sample into the first DNN model to obtain a first prediction ratio includes: Step 1: Extract features from the training sample through the first input layer to obtain the first input extracted features of dimension 1000.
[0061] Among them, the input layer receives training samples and represents the training samples in the form of 1000-dimensional feature vectors as the initial input of the model, providing a basis for subsequent feature extraction and processing.
[0062] Step 2: Reduce the feature dimension of the first input extracted features to 256 through the first round of linear layer, and then activate it through the ReLu function to obtain the first round of extracted features.
[0063] Among them, the first round of linear layer performs weighted summation on the input 1000-dimensional first extracted features to reduce the feature dimension to 256. The ReLu function then performs a nonlinear transformation on the output of the linear layer, introducing nonlinear factors, so that the model can learn more complex feature representations.
[0064] Specifically, the ReLu function sets negative values to 0 and keeps positive values unchanged, which helps alleviate the gradient vanishing problem and accelerates model convergence.
[0065] Step 3: Reduce the feature dimension of the first round of extracted features to 128 through the first and second rounds of linear layers, and then activate them through the ReLu function to obtain the first and second rounds of extracted features.
[0066] The first and second rounds of linear layers perform weighted summation of the 256-dimensional features, reducing the feature dimension to 128. The ReLu function continues to perform nonlinear processing on the features, gradually extracting more representative features.
[0067] Step 4: Reduce the feature dimension of the first and second rounds of extracted features to 64 through the first three rounds of linear layers, and then activate them through the ReLu function to obtain the first and third rounds of extracted features.
[0068] Among them, the first three rounds of linear layers weighted sum the 128-dimensional features and reduce the feature dimension to 64. The ReLu function continues to perform nonlinear processing on the features and gradually extracts more representative features.
[0069] Step 5: Reduce the feature dimension of the first three rounds of extracted features to 32 through the first four rounds of linear layers, and then activate them through the ReLu function to obtain the first four rounds of extracted features.
[0070] The first four linear layers perform weighted summation of the 64-dimensional features, reducing the feature dimension to 32. The ReLu function further compresses the feature dimension and enhances the discriminability of the features through nonlinear transformation, enabling the model to better capture key information in the data.
[0071] In step 6, the feature dimension of the first four rounds of extracted features is reduced to 4 through the first five rounds of linear layers, and then the features are normalized by the Softmax function to obtain the first prediction ratio.
[0072] The first five linear layers perform a weighted summation of the 32-dimensional features, reducing the feature dimension to 4. The softmax function normalizes the output of the linear layer and converts it into a probability distribution, such that the sum of the probabilities of each Bulk Hi-C data sample being in one of the four different cell cycles is 1, thereby obtaining the proportion of samples in each cell cycle.
[0073] Among them, the first DNN model starts with an input layer feature dimension of 1000, and successively reduces the dimension to 256, 128, 64, and 32 through the linear layer, and finally outputs a dimension of 4, thereby gradually extracting key features in the data, reducing the amount of calculation and memory consumption, preventing overfitting, and finally mapping the features to a 4-dimensional space suitable for classification to obtain the proportion of samples in 4 different cell cycles.
[0074] S302: Input the training sample into the second DNN model to obtain a second prediction ratio.
[0075] Optionally, the second DNN model includes a second input layer, a second first-round linear layer, a second second-round linear layer, a second third-round linear layer, a second fourth-round linear layer, and a second fifth-round linear layer. Inputting the training sample into the second DNN model to obtain a second prediction ratio includes: Step 1: Extract features from the training sample through the second input layer to obtain second input extracted features of dimension 1000.
[0076] Among them, the training samples are received as the initial input of the model.
[0077] Step 2: Reduce the feature dimension of the second input extracted features to 512 through the second round of linear layer, and then activate it through the ReLu function to obtain the second round of extracted features.
[0078] The 1000-dimensional features input to the second round of linear layer are weighted and summed to reduce the feature dimension to 512. The ReLu function performs a nonlinear transformation on the output of the linear layer, introducing nonlinear factors to enhance the expressive power of the model.
[0079] Step 3: Reduce the feature dimension of the second round of extracted features to 256 through the second round of linear layer, activate it through the ReLu function, and apply Dropout to randomly discard 30% of the parameters to obtain the second round of extracted features.
[0080] The second linear layer performs a weighted summation of the 512-dimensional features, reducing the feature dimension to 256. The ReLu function further performs nonlinear transformations on the features to extract higher-level features. Dropout randomly discards 30% of the parameters to prevent model overfitting. By randomly discarding some neurons, the model will not be overly dependent on certain neurons during training, thereby improving the model's generalization ability.
[0081] Step 4: Reduce the feature dimension of the second and third rounds of extracted features to 128 through the second and third rounds of linear layers, activate them through the ReLu function, and apply Dropout to randomly discard 20% of the parameters to obtain the second and third rounds of extracted features.
[0082] The second and third linear layers perform a weighted summation of the 256-dimensional features, reducing the feature dimension to 128. The ReLu function further applies nonlinear processing to the features, extracting more representative features. Dropout randomly discards 20% of the parameters to further prevent overfitting and enhance model robustness.
[0083] In step 5, the feature dimension of the second and third rounds of extracted features is reduced to 64 through the second four-round linear layer, and then activated by the ReLu function. Dropout is applied to randomly discard 10% of the parameters to obtain the second four-round extracted features.
[0084] The second four-round linear layer performs a weighted summation of the 128-dimensional features, reducing the feature dimension to 64. The ReLu function further compresses the feature dimension, enhancing feature discrimination. Dropout randomly discards 10% of the parameters. As the model approaches the output layer, the dropout ratio is appropriately reduced to retain more valid information while still preventing overfitting.
[0085] In step 6, the feature dimension of the features extracted in the second four rounds is reduced to 4 through the second five rounds of linear layers, and then the features are normalized by the Softmax function to obtain the second prediction ratio.
[0086] Among them, the second and fifth rounds of linear layers perform weighted summation of the 64-dimensional features to reduce the feature dimension to 4; the Softmax function normalizes the output of the linear layer to obtain the proportion of samples in each cell cycle.
[0087] Among them, the second DNN model input layer feature dimension is 1000, which is first reduced to 512, then to 256, 128, and 64, and finally the output dimension is 4. Important features are extracted through gradual dimensionality reduction, and Dropout technology is combined to prevent overfitting, ensuring that the model maintains good generalization ability while reducing the dimension, and finally the predicted proportion of samples in the four cell cycles is obtained.
[0088] S303: Input the training sample into the third DNN model to obtain a third prediction ratio.
[0089] Optionally, the third DNN model includes a third input layer, a third first-round linear layer, a third second-round linear layer, a third third-round linear layer, a third fourth-round linear layer, and a third fifth-round linear layer. Inputting the training sample into the third DNN model to obtain a third prediction ratio includes: Step 1: Extract features from the training sample through the third input layer to obtain third input extracted features of dimension 1000.
[0090] Among them, the training samples are received as the initial input of the model.
[0091] Step 2: The feature dimension of the first input extracted feature is increased to 1024 through the third round of linear layer, and then activated by the ReLu function to obtain the third round of extracted features.
[0092] The third linear layer performs a weighted summation of the 1000-dimensional input features, increasing the feature dimension to 1024. Applying a nonlinear transformation to the linear layer output and increasing the feature dimension increases the model's capacity, enabling it to learn more complex feature representations. The ReLu activation function introduces nonlinearity, helping the model capture nonlinear relationships in the data.
[0093] Step 3: Reduce the feature dimension of the third round of extracted features to 512 through the third round of linear layers, activate it through the ReLu function, and apply Dropout to randomly discard 60% of the parameters to obtain the third round of extracted features.
[0094] The third and second rounds of linear layers perform a weighted summation of the 1024-dimensional features, reducing the feature dimension to 512. The ReLu function further performs nonlinear transformations on the features to extract higher-level features. Dropout randomly discards 60% of the parameters. Due to the high feature dimensionality, using a larger dropout ratio can more effectively prevent model overfitting and prevent the model from becoming too complex and losing generalization ability.
[0095] Step 4: Reduce the feature dimension of the third round of extracted features to 256 through the third round of linear layer, activate it through the ReLu function, and apply Dropout to randomly discard 30% of the parameters to obtain the third round of extracted features.
[0096] The third linear layer performs a weighted summation of the 512-dimensional features, reducing the feature dimension to 256. The ReLu function further applies nonlinear processing to the features, extracting more representative features. Dropout randomly discards 30% of the parameters. As the feature dimension decreases, the dropout ratio is appropriately reduced to retain more effective information while still preventing overfitting.
[0097] Step 5: Reduce the feature dimension of the third and fourth rounds of extracted features to 128 through the third and fourth rounds of linear layers, activate them through the ReLu function, and apply Dropout to randomly discard 10% of the parameters to obtain the third and fourth rounds of extracted features.
[0098] The third and fourth linear layers perform a weighted summation of the 256-dimensional features, reducing the feature dimension to 128. The ReLu function further compresses the feature dimension, enhancing feature discrimination. Dropout randomly discards 10% of the parameters. As the model approaches the output layer, the dropout ratio is further reduced to ensure that the model can fully utilize the learned features for accurate classification.
[0099] In step 6, the feature dimension of the features extracted in the third and fourth rounds is reduced to 4 through the third and fifth rounds of linear layers, and then the features are normalized by the Softmax function to obtain the third prediction ratio.
[0100] The third and fifth linear layers perform weighted summation of the 128-dimensional features, reducing the feature dimension to 4. The softmax function normalizes the output of the linear layer to obtain the proportion of samples in each cell cycle.
[0101] The third DNN model's input layer feature dimension was 1000, which was first increased to 1024, then sequentially reduced to 512, 256, and 128, with a final output dimension of 4. The purpose of this third DNN model's dimensionality increase was to increase the model's capacity, enabling it to learn more complex feature representations. Subsequently, through gradual dimensionality reduction, noise and redundant information were removed, retaining the most important features. Incorporating the dropout technique to prevent overfitting, predictions for the samples across four cell cycles were ultimately obtained.
[0102] S304: Averaging the predicted proportions of cells in different cell cycles in the first predicted proportion, the second predicted proportion, and the third predicted proportion to obtain a predicted value.
[0103] Among them, by integrating the prediction results of multiple DNN models, the advantages of each DNN model can be combined, the errors and deviations of a single model can be reduced, and the stability and generalization ability of the cell cycle proportion prediction model can be improved, thereby obtaining more accurate and reliable prediction results of the proportion of samples in four different cell cycles.
[0104] In this example, the cell cycle ratio prediction model consists of three deep neural network models with different structures and dropout regularization technology. Each model has four hidden layers of different sizes. The softmax function is used to calculate scores to distinguish cell cycle ratios, and supervised training is performed on simulated data.
[0105] The second and third DNN models employed dropout with a gradually decreasing dropout ratio to accommodate the needs of different training stages. This is because a higher dropout ratio in the early stages of training enhances regularization, forcing the model to learn more robust feature combinations and avoid premature local optima or overfitting. Strong regularization helps suppress oversensitivity to noise, especially in scenarios with less data or high noise levels. Gradually reducing the dropout ratio in the later stages of training reduces artificial interference with neurons, allowing the model to more fully utilize all neurons for parameter fine-tuning, improving the model's convergence stability and finality.
[0106] The model uses ReLu as the activation function, which can speed up training, alleviate the gradient vanishing problem, introduce nonlinear characteristics, and achieve sparse activation of neurons, thereby improving the efficiency and accuracy of the model.
[0107] The model uses linear layers of gradually decreasing size. By reducing the dimensionality layer by layer, it gradually filters out redundant information and extracts high-order features. Low-dimensional layers have fewer parameters, which reduces the risk of overfitting. This structure is equivalent to implicit L2 regularization, complementing explicit regularization methods. Reducing the number of parameters per layer directly reduces memory usage and computational complexity. The decreasing layer dimensionality also alleviates the exploding / vanishing gradient problem.
[0108] By combining the first, second, and third DNN models, the first DNN model uses simple step-by-step dimensionality reduction to extract basic features and classify them. The second DNN model builds on this foundation by introducing dropout to enhance generalization and reduce overfitting. The third DNN model increases model capacity by increasing dimensionality to capture complex features. These three DNN models process the data from different perspectives, enabling a more comprehensive exploration of data features and adapting to data of varying complexity. Finally, ensemble learning is performed by combining the results of the three DNN models. By integrating the strengths of each model, ensemble learning can reduce overfitting, improve model accuracy and stability, and enhance generalization. The three models utilize different hidden layer sizes and varying dropout ratios to capture and extract data features from multiple perspectives. The resulting cell cycle proportion prediction model effectively reduces the error and bias of individual models, improving the stability, accuracy, and robustness of prediction results, and more accurately predicting the proportions of cells in the four different cell cycles in BulkHi-C samples.
[0109] Specifically, the training parameters included the Adam optimizer, a learning rate of 0.0001, and an L1 loss as the optimization objective. Each model was trained independently for 5000 steps, using early stopping to prevent overfitting. Due to the large number of 2D bins in the two-dimensional matrix, only the 2D bins of the pseudo-bulk Hi-C matrix with the highest variance were selected for modeling. The interaction values of these selected 2D bins were then converted into vectors and input into the DNN model.
[0110] Assuming single-cell Hi-C data from mouse embryonic stem cells, a dataset was prepared for training and testing. Using deep neural network (DNN) models and ensemble learning methods from deep learning, single-cell Hi-C data were randomly mixed to generate pseudo-bulk Hi-C data. A model capable of identifying signatures of chromatin conformational changes across cell cycles was constructed and tested. The model was validated using a set of bulk Hi-C data (9 samples, including long-term hematopoietic stem cells, short-term hematopoietic stem cells, multipotent progenitors, common myeloid progenitors, granulocyte-macrophage progenitors, megakaryocyte-erythroid progenitors, common lymphoid progenitors, and megakaryocyte progenitors and granulocytes) containing a mixture of cells at different cell cycles and corresponding known cell cycle proportions.
[0111] See Figure 6 In the test set, the MSE (mean square error) between the model prediction value and the true value is 3.43×10-4, and the Pearson correlation coefficient is 0.99, indicating that the model prediction effect is better when the modeled cell type (mouse ESC stem cell) and the tested cell type (mouse ESC stem cell) are consistent.
[0112] See Figure 7 In the validation set, the Pearson correlation coefficient between the predicted and true values was 0.40, and the MSE was 0.028, indicating reasonable accuracy. However, the prediction performance was not as good as the test set results, indicating that the model's prediction performance was poor when the modeled cell type (mouse ESCs) and the test cell type (mouse adult blood cells) were inconsistent. However, with the increasing popularity of single-cell Hi-C technology, if single-cell Hi-C data of mouse adult blood cells were available, it would be possible to build a model that achieved highly accurate predictions. In other words, when the modeled cell type and the test cell type were consistent, the model's accuracy was very high.
[0113] Figure 8 This is a schematic diagram of the structure of a device for predicting the proportion of cells in different cell cycles based on chromatin conformation profiles provided in an embodiment of the present application. Figure 8The device for predicting the proportion of cells in different cell cycles based on chromatin conformation profiles includes various functional modules for implementing the aforementioned method for predicting the proportion of cells in different cell cycles based on chromatin conformation profiles. Any functional module can be implemented by software and / or hardware.
[0114] In some embodiments, the apparatus 800 for predicting the proportion of cells in different cell cycles based on chromatin conformation profiles includes a data acquisition module 801 and a proportion prediction module 802. The data acquisition module 801 is used to obtain the bulk Hi-C data to be tested; The ratio prediction module 802 is used to input the bulk Hi-C data to be tested into the cell cycle ratio prediction model to obtain the cell cycle ratio prediction result output by the cell cycle ratio prediction model; wherein the cell cycle ratio prediction model is obtained based on the training of the deep neural network model.
[0115] In some embodiments, the apparatus 800 further includes a model training module 803, which is specifically configured to: Acquire multiple training samples and construct a training set; wherein the training samples are labeled with true values, and the true values include the true proportion of cells in different cell cycles in the training samples; Inputting the training samples contained in the training set into the deep neural network model to obtain prediction values; wherein the prediction values include the predicted proportions of cells in different cell cycles in the training samples; Based on the preset loss function, the loss between the predicted value and the true value in the training set is obtained; The network parameters of the deep neural network model are optimized based on the loss amount until the loss function converges or the preset number of training iterations is reached, and a cell cycle ratio prediction model is obtained.
[0116] In some embodiments, the deep neural network model includes a first DNN model, a second DNN model, and a third DNN model. The model training module 803 is further configured to: Input the training sample into the first DNN model to obtain a first prediction ratio; Input the training sample into the second DNN model to obtain a second prediction ratio; Input the training sample into the third DNN model to obtain a third prediction ratio; The predicted proportions of cells in different cell cycles in the first predicted proportion, the second predicted proportion, and the third predicted proportion are averaged to obtain a predicted value.
[0117] In some embodiments, the first DNN model includes a first input layer, a first one-round linear layer, a first two-round linear layer, a first three-round linear layer, a first four-round linear layer, and a first five-round linear layer. The model training module 803 is further configured to: Extract features from the training sample through the first input layer to obtain first input extracted features with a dimension of 1000; The feature dimension of the first input extracted features is reduced to 256 through the first round of linear layers, and then activated by the ReLu function to obtain the first round of extracted features; The feature dimension of the first round of extracted features is reduced to 128 through the first and second rounds of linear layers, and then activated by the ReLu function to obtain the first and second rounds of extracted features; The feature dimensions of the first and second rounds of extracted features are reduced to 64 through the first and third rounds of linear layers, and then activated by the ReLu function to obtain the first and third rounds of extracted features; The feature dimensions of the first three rounds of extracted features are reduced to 32 through the first four rounds of linear layers, and then activated by the ReLu function to obtain the first four rounds of extracted features; The feature dimension of the features extracted in the first four rounds is reduced to 4 through the first five rounds of linear layers, and then the features are normalized by the Softmax function to obtain the first prediction ratio.
[0118] In some embodiments, the second DNN model includes a second input layer, a second first-round linear layer, a second second-round linear layer, a second third-round linear layer, a second fourth-round linear layer, and a second fifth-round linear layer. The training sample is input into the second DNN model. The model training module 803 is further specifically used to: Extract features from the training sample through the second input layer to obtain a second input extracted feature with a dimension of 1000; The feature dimension of the second input extracted features is reduced to 512 through the second round of linear layer, and then activated by the ReLu function to obtain the second round of extracted features; The feature dimension of the second round of extracted features is reduced to 256 through the second round of linear layer, and then activated by the ReLu function. Dropout is applied to randomly discard 30% of the parameters to obtain the second round of extracted features; The feature dimension of the second and third rounds of extracted features is reduced to 128 through the second and third rounds of linear layers, and then activated by the ReLu function. Dropout is applied to randomly discard 20% of the parameters to obtain the second and third rounds of extracted features; The feature dimension of the features extracted in the second and third rounds is reduced to 64 through the second and fourth rounds of linear layers, and then activated by the ReLu function. Dropout is applied to randomly discard 10% of the parameters to obtain the second and fourth rounds of extracted features; The feature dimension of the features extracted in the second four rounds is reduced to 4 through the second five rounds of linear layers, and then the features are normalized by the Softmax function to obtain the second prediction ratio.
[0119] In some embodiments, the third DNN model includes a third input layer, a third first-round linear layer, a third second-round linear layer, a third third-round linear layer, a third fourth-round linear layer, and a third fifth-round linear layer. The training sample is input into the second DNN model. The model training module 803 is further specifically used to: Extract features from the training sample through the third input layer to obtain a third input extracted feature with a dimension of 1000; The feature dimension of the first input extracted features is increased to 1024 through the third round of linear layer, and then activated by the ReLu function to obtain the third round of extracted features; The feature dimension of the features extracted in the third round is reduced to 512 through the third round of linear layers, and then activated by the ReLu function. Dropout is applied to randomly discard 60% of the parameters to obtain the third and second extracted features; The feature dimension of the features extracted in the third round is reduced to 256 through the third round of linear layer, and then activated by the ReLu function. Dropout is applied to randomly discard 30% of the parameters to obtain the third round of extracted features; The feature dimension of the third and fourth rounds of extracted features is reduced to 128 through the third and fourth rounds of linear layers, and then activated by the ReLu function. Dropout is applied to randomly discard 10% of the parameters to obtain the third and fourth rounds of extracted features; The feature dimension of the features extracted in the third and fourth rounds is reduced to 4 through the third and fifth rounds of linear layers, and then the features are normalized by the Softmax function to obtain the third prediction ratio.
[0120] In some embodiments, the model training module 803 is further configured to: Acquire multiple single-cell Hi-C data; the single-cell Hi-C data is an interaction matrix; Preprocessing the single-cell Hi-C data to obtain preprocessed Hi-C data; wherein the preprocessing includes depth filtering and depth normalization; Randomly select multiple preprocessed Hi-C data for fusion to obtain pseudo Bulk Hi-C data as training samples; The device 800 for predicting the proportion of cells in different cycles based on chromatin conformation profiles provided in an embodiment of the present application is used to execute the technical solution provided in the aforementioned method embodiment for predicting the proportion of cells in different cycles based on chromatin conformation profiles. Its implementation principle and technical effects are similar to those in the aforementioned method embodiment and will not be repeated here.
[0121] It should be noted that it should be understood that the division of the various modules of the above device is only a division of logical functions. In actual implementation, they can be fully or partially integrated into one physical entity, or they can be physically separated. And these modules can all be implemented in the form of software called by processing elements, or all in the form of hardware. Some modules can also be implemented in the form of software called by processing elements, and some modules can be implemented in the form of hardware. For example, the model training module 803 can be a separately established processing element, or it can be integrated in a certain chip of the above device. In addition, it can also be stored in the memory of the above device in the form of program code, and called and executed by a certain processing element of the above device. The implementation of other modules is similar. In addition, these modules can be fully or partially integrated together, or they can be implemented independently. The processing element here can be an integrated circuit with signal processing capabilities. In the implementation process, each step of the above method or each module above can be completed by the hardware integrated logic circuit in the processor element or the instruction in the form of software.
[0122] Figure 9 This is a schematic diagram of the structure of an electronic device provided in an embodiment of the present application, see Figure 9 The electronic device 900 includes a processor 901 and a memory 902 communicatively connected to the processor 901; Memory 902 stores computer-executable instructions; The processor 901 executes the computer-executable instructions stored in the memory 902 to implement the technical solution of the aforementioned method for predicting the proportion of cells in different cell cycles based on the chromatin conformation spectrum.
[0123] In the electronic device 900 described above, the memory 902 and processor 901 are directly or indirectly electrically connected to each other to enable data transmission or interaction. For example, these components can be electrically connected via one or more communication buses or signal lines, such as a bus. The bus can be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, or an Extended Industry Standard Architecture (EISA) bus, among others. Buses can be categorized as address buses, data buses, control buses, etc., but this does not mean that there is only one bus or only one type of bus. Memory 902 stores computer-executable instructions for implementing the aforementioned method for predicting the proportion of cells in different cell cycles based on chromatin conformation profiles, including at least one software functional module that can be stored in memory 902 in the form of software or firmware. Processor 901 executes various functional applications and data processing by running the software programs and modules stored in memory 902.
[0124] The memory 902 includes at least one type of readable storage medium, including but not limited to random access memory (RAM), read only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), and electrically erasable programmable read-only memory (EEPROM). The memory 902 is used to store programs, and the processor 901 executes the programs after receiving execution instructions. Furthermore, the software programs and modules in the memory 902 may also include an operating system, which may include various software components and / or drivers for managing system tasks (such as memory management, storage device control, power management, etc.), and may communicate with various hardware or software components to provide an operating environment for other software components.
[0125] The processor 901 can be an integrated circuit chip with signal processing capabilities. The above-mentioned processor 901 can be a general-purpose processor, including a central processing unit (CPU), a network processor (NP), a digital signal processor (DSP), an application-specific integrated circuit (ASIC), etc. It can implement or execute the various methods, steps, and logic block diagrams disclosed in the embodiments of this application. The general-purpose processor can be a microprocessor, or the processor 901 can also be any conventional processor, etc.
[0126] The electronic device 900 is used to execute the technical solution provided by the aforementioned method embodiment for predicting the proportion of cells in different cycles based on chromatin conformation spectrum. Its implementation principle and technical effects are similar to those in the aforementioned method embodiment and will not be repeated here.
[0127] An embodiment of the present application also provides a computer-readable storage medium, which stores computer-executable instructions. When the computer-executable instructions are executed, they are used to implement the technical solution of the method for predicting the proportion of cells in different cycles based on the chromatin conformation spectrum as described above.
[0128] The computer-readable storage medium may be implemented by any type of volatile or non-volatile memory device, or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic storage, flash memory, magnetic disk, or optical disk. The computer-readable storage medium may be any available medium that can be accessed by a general-purpose or special-purpose computer.
[0129] An exemplary readable storage medium is coupled to a processor so that the processor can read information from the readable storage medium and write information to the readable storage medium. Of course, the readable storage medium can also be an integral part of the processor. The processor and the readable storage medium can be located in an application specific integrated circuit (ASIC). Of course, the processor and the readable storage medium can also be present as discrete components in a control device for a device for predicting the proportion of cells in different cell cycles based on a chromatin conformation profile.
[0130] An embodiment of the present application also provides a computer program product, including a computer program, which, when executed, is used to implement the technical solution of the aforementioned method for predicting the proportion of cells in different cell cycles based on chromatin conformation profiles.
[0131] In the above embodiments, those skilled in the art will appreciate that the above method embodiments can be implemented in whole or in part via software, hardware, firmware, or any combination thereof. When implemented using software, they can be implemented in whole or in part in the form of a computer program product. A computer program product comprises one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the processes or functions according to the embodiments of the present application are generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (e.g., coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless network, microwave, etc.) means. The computer-readable storage medium can be any available medium that can be accessed by a computer or a data storage device such as a server or data center that integrates one or more available media. Available media may be magnetic media (eg, floppy disks, hard disks, magnetic tapes), optical media (eg, DVDs), or semiconductor media (eg, solid-state drives (SSDs)).
[0132] In the above embodiments, the description of each embodiment has its own emphasis. For parts not described in detail in a particular embodiment, please refer to the relevant description of other embodiments. The technical features of the above embodiments can be combined in any way. To keep the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0133] Those skilled in the art will readily appreciate other embodiments of the present invention after considering the specification and practicing the invention disclosed herein. This application is intended to cover any variations, uses, or adaptations of the present invention that follow the general principles of the present invention and include common knowledge or customary techniques in the art not disclosed herein. The description and examples are to be considered merely as exemplary, and the true scope and spirit of the present application are indicated by the appended claims.
[0134] It should be understood that the present application is not limited to the exact structure described above and shown in the drawings, and that various modifications and changes may be made without departing from the scope thereof. The scope of the present application is limited only by the appended claims.
Claims
1. A method for predicting the proportion of cells in different cell cycles based on chromatin conformation profiles, characterized in that: include: Obtain bulk Hi-C data to be tested; The bulk Hi-C data to be tested is input into a cell cycle ratio prediction model to obtain a cell cycle ratio prediction result output by the cell cycle ratio prediction model; wherein the cell cycle ratio prediction model is obtained based on training a deep neural network model.
2. The method according to claim 1, characterized in that Training of deep neural network models, including: Acquire multiple training samples and construct a training set; wherein the training samples are annotated with true values, and the true values include the true proportions of cells in different cell cycles in the training samples; Inputting the training samples contained in the training set into the deep neural network model to obtain a predicted value; wherein the predicted value includes the predicted proportion of cells in different cell cycles in the training samples; Based on a preset loss function, obtaining the loss between the predicted value and the true value in the training set; The network parameters of the deep neural network model are optimized based on the loss amount until the loss function converges or reaches a preset number of training iterations, thereby obtaining a cell cycle ratio prediction model.
3. The method according to claim 2, characterized in that The deep neural network model includes a first DNN model, a second DNN model, and a third DNN model. The training samples included in the training set are input into the deep neural network model to obtain a prediction value, including: Inputting the training sample into a first DNN model to obtain a first prediction ratio; Inputting the training sample into a second DNN model to obtain a second prediction ratio; Inputting the training sample into a third DNN model to obtain a third prediction ratio; The predicted proportions of cells in different cell cycles in the first predicted proportion, the second predicted proportion and the third predicted proportion are averaged to obtain a predicted value.
4. The method according to claim 3, characterized in that The first DNN model includes a first input layer, a first first-round linear layer, a first second-round linear layer, a first third-round linear layer, a first fourth-round linear layer, and a first fifth-round linear layer. The training sample is input into the first DNN model to obtain a first prediction ratio, including: Extracting features from the training sample through the first input layer to obtain first input extracted features of dimension 1000; The feature dimension of the first input extracted features is reduced to 256 by the first round of linear layers, and then activated by the ReLu function to obtain the first round of extracted features; The feature dimensions of the first round of extracted features are reduced to 128 by the first and second rounds of linear layers, and then activated by the ReLu function to obtain the first and second rounds of extracted features; The feature dimensions of the features extracted from the first and second rounds are reduced to 64 by the first three rounds of linear layers, and then activated by the ReLu function to obtain the first and third rounds of extracted features; The feature dimensions of the features extracted from the first three rounds are reduced to 32 by the first four rounds of linear layers, and then activated by the ReLu function to obtain the first four rounds of extracted features; The feature dimensions of the features extracted in the first four rounds are reduced to 4 through the first five rounds of linear layers, and then the features are normalized by the Softmax function to obtain the first prediction ratio.
5. The method according to claim 3, characterized in that The second DNN model includes a second input layer, a second first-round linear layer, a second second-round linear layer, a second third-round linear layer, a second fourth-round linear layer, and a second fifth-round linear layer. The training sample is input into the second DNN model to obtain a second prediction ratio, including: Extract features from the training sample through the second input layer to obtain a second input extracted feature with a dimension of 1000; The feature dimension of the second input extracted features is reduced to 512 by the second round of linear layers, and then activated by the ReLu function to obtain the second round of extracted features; The feature dimension of the second round of extracted features is reduced to 256 by the second round of linear layers, and then activated by the ReLu function. Dropout is applied to randomly discard 30% of the parameters to obtain the second round of extracted features; The feature dimensions of the second-second round extracted features are reduced to 128 through the second-third round linear layer, and then activated by the ReLu function. Dropout is applied to randomly discard 20% of the parameters to obtain the second-third round extracted features; The feature dimensions of the features extracted in the second and third rounds are reduced to 64 by the second four rounds of linear layers, and then activated by the ReLu function. Dropout is applied to randomly discard 10% of the parameters to obtain the second four rounds of extracted features; The feature dimension of the features extracted in the second four rounds is reduced to 4 through the second five rounds of linear layers, and then the features are normalized by the Softmax function to obtain the second prediction ratio.
6. The method according to claim 3, characterized in that The third DNN model includes a third input layer, a third first-round linear layer, a third second-round linear layer, a third third-round linear layer, a third fourth-round linear layer, and a third fifth-round linear layer. The training sample is input into the third DNN model to obtain a third prediction ratio, including: Extracting features from the training sample through the third input layer to obtain a third input extracted feature of dimension 1000; The feature dimension of the first input extracted feature is increased to 1024 by the third round of linear layer, and then activated by the ReLu function to obtain the third round of extracted features; The feature dimension of the third round of extracted features is reduced to 512 by the third round of linear layers, and then activated by the ReLu function, and 60% of the parameters are randomly discarded by Dropout to obtain the third round of extracted features; The feature dimension of the third round of extracted features is reduced to 256 by the third round of linear layer, and then activated by ReLu function, and 30% of the parameters are randomly discarded by Dropout to obtain the third round of extracted features; The feature dimension of the third-fourth round of extracted features is reduced to 128 by the third-fourth round of linear layers, and then activated by the ReLu function, and 10% of the parameters are randomly discarded by Dropout to obtain the third-fourth round of extracted features; The feature dimension of the features extracted in the third and fourth rounds is reduced to 4 by the third and fifth rounds of linear layers, and then the features are normalized by the Softmax function to obtain the third prediction ratio.
7. The method according to any one of claims 2 to 6, characterized in that: The method comprises: Acquire a plurality of single-cell Hi-C data; wherein the single-cell Hi-C data is an interaction matrix; Preprocessing the single-cell Hi-C data to obtain preprocessed Hi-C data; wherein the preprocessing includes depth filtering and depth normalization; Multiple preprocessed Hi-C data are randomly selected for fusion to obtain pseudo Bulk Hi-C data as training samples.
8. A device for predicting the proportion of cells in different life cycles based on chromatin conformation profiles, characterized in that: include: Data acquisition module, used to obtain bulk Hi-C data to be tested; A ratio prediction module is used to input the bulk Hi-C data to be tested into a cell cycle ratio prediction model to obtain a cell cycle ratio prediction result output by the cell cycle ratio prediction model; wherein the cell cycle ratio prediction model is obtained based on training a deep neural network model.
9. An electronic device, characterized in that: comprising a processor, and a memory communicatively connected to the processor; The memory stores computer-executable instructions; The processor executes the computer-executable instructions stored in the memory to implement the method according to any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that The computer-readable storage medium stores computer-executable instructions, which, when executed, are used to implement the method according to any one of claims 1 to 7.
Citation Information
Patent Citations
LncRNA IFA and application thereof in porcine ovarian granular cells
CN115992135A
Cell cycle prediction method and system based on single cell Hi-C data
CN116469458A
Cell proportion prediction method, model generation method, equipment and medium
CN116779043A
Method, device and equipment for predicting cell type proportion and storage medium
CN117292750A
Cell type prediction method and system based on multivariate feature set fusion
CN117633630A