Single cell cross-modal data conversion method and device based on diffusion model

Through the single-cell cross-modal data conversion method based on the diffusion model, the diffusion model and conditional generation strategy are used to solve the problem of cross-modal integration in single-cell multimodal analysis, and efficient and accurate cross-modal data conversion is achieved, which improves the robustness and generalization ability of the model.

CN120299517APending Publication Date: 2025-07-11NINGXIA UNIVERSITY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510369121.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-27
Publication Date
2025-07-11

AI Technical Summary

Technical Problem

The existing single-cell multimodal analysis methods face problems such as difficult data acquisition, complex experiments, high cost, low technical throughput and high noise when integrating across modalities, resulting in weak generalization capabilities of models and difficulty in comprehensively capturing the complexity and data diversity between modals.

Method used

A single-cell cross-modal data conversion method based on diffusion model is adopted. By introducing diffusion model and conditional generation strategy, the source modal data is used to guide the probability distribution generation of target modal data, and an encoder and decoder model is built to carry out the disturbance diffusion and denoising process of gradually adding Gaussian noise. The noise prediction parameters are optimized by the Unet network to realize cross-modal data conversion.

Benefits of technology

Effectively capture the cellular heterogeneity between different modes, improve the accuracy and robustness of cross-modal conversion, reduce the influence of batch effects, enhance the robustness of sparse data and noise, and improve the generalization ability of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure BDA0005330835690000037
    Figure BDA0005330835690000037
  • Figure BDA0005330835690000042
    Figure BDA0005330835690000042
  • Figure BDA0005330835690000045
    Figure BDA0005330835690000045
Patent Text Reader

Abstract

The invention provides a single cell cross-modal data conversion method and equipment based on a diffusion model, and belongs to the technical field of bioinformatics. Comprising the following steps: building a single cell cross-modal data conversion model based on a diffusion model, taking single cell RNA and ATAC as examples, and forming an encoder Er, an encoder Ea, a decoder Dr, a decoder Da, an RNA diffusion conversion model and an ATAC diffusion conversion model; carrying out data preprocessing and self-encoder pre-training, optimizing parameters of each encoder and decoder, and freezing the parameters; the method comprises the following steps: training a single cell cross-modal data conversion model by utilizing ATAC and RNA data, including a process of converting ATAC into RNA and a process of converting RNA into ATAC, updating neural network parameters used for predicting RNA noise and neural network parameters used for predicting ATAC noise by utilizing a Unet network, and minimizing the difference between prediction noise and sampling noise through reverse optimization so as to obtain a single cell cross-modal data conversion model. Stopping updating the model parameters after training is finished; and performing single-cell cross-modal data generation by using the single-cell cross-modal data conversion model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of bioinformatics, and particularly to a single-cell cross-modal data conversion method and device based on a diffusion model. Background Art

[0002] With the continuous development of single-cell sequencing technology, researchers can accurately analyze the state of cells from multiple perspectives and reveal different gene regulation mechanisms within cells. However, single-modal sequencing can only capture limited information and is difficult to comprehensively reveal the gene regulation mechanisms at different levels within cells. Single-cell multi-modal analysis technology has emerged. It promotes in-depth research on cell heterogeneity and disease mechanisms by simultaneously analyzing multiple modal data of the same cell. However, single-cell multi-modal analysis faces problems such as difficult data acquisition, complex experiments, high costs, low technical throughput, and large noise. These factors limit the wide application of single-cell multi-modal technology. Although there are already various computational methods for integrating multi-modal data, due to batch effects, data sparsity, and biological background differences between different modalities, cross-modal integration still faces great challenges.

[0003] To solve this problem, it is necessary to explore how to use single-modal data to infer other modal information to make up for the limitations of experimental data. This process is called cross-modal conversion, but it involves difficult problems such as data noise, high-dimensional data modeling, batch effects between modalities, and generalization ability for unseen cell types. Currently, there are already some deep learning algorithms for single-cell cross-modal conversion. BABEL uses two autoencoders to map RNA and ATAC data to a shared latent space and realizes cross-modal conversion through an interoperable encoder-decoder structure. Polarbear first trains variational autoencoders for single-modal RNA and ATAC separately using unsupervised learning, and then uses joint-modal data to supervise the learning of the conversion layer between the two autoencoders. JAMIE generates a shared latent space by aggregating the correlations in the latent space and combining information from different modalities, and uses an adversarial decoder for the representations in the latent space to achieve cross-modal data filling and feature inference. scButterfly uses a double-aligned variational autoencoder to achieve cross-modal conversion and establishes relationships between different modalities by aligning the latent representations of different modalities, thereby realizing data conversion. In addition, there are also some methods that use multi-modal integration models to generate missing data. scMM predicts missing modal data based on a variational autoencoder of a mixture of experts model. CLUE uses a cross-encoder structure to learn joint latent representations of modalities to complete missing modal data. These methods achieve conversion and have achieved certain effects by learning joint latent representations across modalities or aligning the latent space. However, they rely too much on direct mapping and reconstruction of the latent space and the generated data is a single data point, making it difficult to comprehensively capture the complexity and data diversity between modalities. Especially when there is missing data or large data noise, the generalization ability of the model is weak, affecting the cross-modal conversion effect. Summary of the Invention

[0004] In view of this, the present invention provides a single-cell cross-modal data conversion method based on a diffusion model. By introducing a diffusion model and adopting a conditional generation strategy, the probability distribution generation of the target modal data is guided by the source modal data, which can effectively capture the cell heterogeneity between different modalities, realize efficient cross-modal conversion of single-cell data, and has good generalization ability and robustness.

[0005] The technical solution adopted by the embodiments of the present invention to solve its technical problems is:

[0006] A single-cell cross-modal data conversion method based on a diffusion model, comprising:

[0007] Step S1, build a single-cell cross-modal data conversion model based on a diffusion model. Taking the mutual conversion between single-cell RNA and single-cell ATAC as an example, the model consists of an encoder E r and an encoder E a and a decoder D r, decoder D a , composed of an RNA diffusion conversion model and an ATAC diffusion conversion model;

[0008] Step S2, data preprocessing and pre-training of the autoencoder: preprocess single-cell RNA and single-cell ATAC; use the preprocessed RNA and ATAC for pre-training of the autoencoder to optimize E r , E a , D r , D a parameters, and freeze E r , E a , D r , D a parameters;

[0009] Step S3, train the single-cell cross-modal data conversion model:

[0010] The ATAC is input into E a After encoding, a latent representation in the low-dimensional space is obtained The RNA is input into E r After encoding, a latent representation is obtained

[0011] In the process of ATAC converting to RNA, the target modality is input into the RNA diffusion conversion model, first gradually add Gaussian noise for perturbation diffusion, and then the source modality predicts the noise; in the process of RNA converting to ATAC, the target modality is input into the ATAC diffusion conversion model, first gradually add Gaussian noise for perturbation diffusion, and then the source modality predicts the noise; use the Unet network to update the neural network parameters for predicting RNA noise and the neural network parameters for predicting ATAC noise, and minimize the difference between the predicted noise and the sampled noise through reverse optimization;

[0012] After the training of the RNA diffusion conversion model and the ATAC diffusion conversion model is completed, stop updating the parameters;

[0013] Step S4, use the trained single-cell cross-modal data conversion model to generate single-cell cross-modal data. Specifically, input the source modality data, obtain the latent representation of the source modality through encoding by the corresponding encoder, and guide the denoising process of the target modality by the latent representation of the source modality to obtain the latent representation of the target modality. After decoding by the decoder of the target modality, the feature expression of the target modality is obtained.

[0014] Preferably, the preprocessing process of RNA is as follows: First, perform global normalization, add 1 to the normalized data and perform logarithmic transformation, and then select the top 3000 highly variable genes;

[0015] The preprocessing process of ATAC is as follows: First, perform binarization, and discard the data with the proportion of cells with chromatin open peaks lower than 0.5%; use the TF-IDF (term frequency-inverse document frequency) algorithm to perform feature transformation on the binarized and filtered peak matrix to obtain the TF-IDF weighted matrix; finally, perform normalization processing to scale the values to the range of [0, 1].

[0016] Preferably, in the RNA autoencoder composed of E r and D r , both E r and D r adopt a two-layer fully connected network structure. The number of neurons in the two layers of E r is 256 and 128, and D r is symmetrically set with the structure of E r ;

[0017] The RNA autoencoder uses the MSE reconstruction loss function to update the network weights:

[0018]

[0019] In the formula, x r is the input RNA, and D r (E r (x r )) i is the reconstructed data of the i-th feature by the RNA autoencoder, and m is the total number of features of the RNA; is the loss of the autoencoder composed of E r and D r .

[0020] Preferably, in the ATAC autoencoder composed of E a and D a , E a extends the chromosome features to 32 dimensions and maps them to a 128-dimensional latent space through two layers of LeakyReLU activation functions. D a performs feature reconstruction through LeakyReLU, and finally uses the Sigmoid function to constrain the output to the range of [0, 1];

[0021] The ATAC autoencoder uses the BCE reconstruction loss function to update the network weights:

[0022]

[0023] In the formula, xa is the input ATAC, D r (E r (x a )) i is the reconstructed data of the i-th feature by the ATAC autoencoder, and n is the total number of features of ATAC; is the loss of the autoencoder composed of E and D a and D a .

[0024] Preferably, the RNA diffusion conversion model and the ATAC diffusion conversion model adopt the same data processing process. The processing of the input data x0 includes a noise addition process and a denoising process;

[0025] The noise addition process uses the diffusion model DDPM to gradually perturb x0 with Gaussian noise. By gradually adding Gaussian noise to x0 at each time step t, x0 is completely transformed into Gaussian noise. The diffusion process expression is:

[0026]

[0027] In the formula, T is the total number of time steps, and x t is the noise state at time t; α t is the scheduling weight of the noise, and β t is used to control the amount of noise added to the data at each time step t; I in represents the covariance matrix of the standard normal distribution, and the variance of each dimension is 1; is the Gaussian noise obeying the standard normal distribution, represents the cumulative signal retention factor;

[0028] The denoising process uses the deterministic diffusion model DDIM, and the conditional distribution of DDIM is modeled as:

[0029]

[0030] In the formula, σ t is the noise perturbation factor;

[0031] The neural network ∈ θ (x t ) is used to estimate x0, and the estimation result expression is:

[0032]

[0033] In the formula, represents the predicted noise at time step t;

[0034] The generation process is modeled by the conditional probability distribution of DDIM at each time step:

[0035]

[0036] Given the sample x at the current time step t and ∈ θ (x t ), generate the sample x t-1 :

[0037]

[0038] In the formula, the diffusion model DDIM is deterministic sampling, let η = 0, accordingly, σ t = 0, making the sample generation process a deterministic process; ∈ t represents the noise at the current time step;

[0039] During the conversion process, x0 represents when the distribution of the diffusion process is modeled as x0 represents when the distribution of the diffusion process is modeled as Use the Unet network to update the neural network parameters θ for predicting RNA noise in the model r and the neural network parameters θ for predicting ATAC noise a , and minimize the difference between the predicted noise and the sampled noise through backpropagation optimization:

[0040]

[0041] In the formula, L θ (θ r , θ a ) is the loss function, ε r is the standard Gaussian noise added to the RNA modality, ε a is the standard Gaussian noise added to the ATAC modality;

[0042] At each time step t, the reverse inference process is expressed as:

[0043]

[0044] In the formula, is the latent state of the RNA modality at time t - 1, is the latent state of the ATAC modality at time t - 1; represents the neural network for predicting RNA noise, represents the neural network for predicting ATAC noise;

[0045] In the inference stage of cross-modal conversion, define the expressions of the low-dimensional features of the RNA conversion to ATAC process and the ATAC conversion to RNA process:

[0046]

[0047] Wherein, the converted low-dimensional features are obtained by gradually denoising and

[0048] Preferably, each of the two decoders in the single-cell cross-modal data conversion model outputs the reconstruction result of the autoencoder and the low-dimensional feature conversion result respectively:

[0049]

[0050] Preferably, after the step S4, it further includes:

[0051] For the source modal data, a multi-sampling strategy is adopted for cross-modal data conversion, and the mean of the sampling results is used as the final low-dimensional feature prediction result The expression is:

[0052]

[0053] Wherein, N is the number of samplings, represents the low-dimensional feature obtained by prediction using the data obtained from the i-th sampling.

[0054] As can be seen from the above technical solutions, in the single-cell cross-modal data conversion method based on the diffusion model provided by the embodiments of the present invention, taking the mutual conversion of single-cell RNA and single-cell ATAC as an example, first build a single-cell cross-modal data conversion model based on the diffusion model, including the encoder E r and the encoder E a , the decoder D r and the decoder D a, consisting of an RNA diffusion conversion model and an ATAC diffusion conversion model; perform data preprocessing and pre-training of the autoencoder, optimize the parameters of each encoder and decoder, and freeze the parameters; use ATAC and RNA data to train a single-cell cross-modal data conversion model, including the process of converting ATAC to RNA and the process of converting RNA to ATAC, use the Unet network to update the neural network parameters for predicting RNA noise and the neural network parameters for predicting ATAC noise, and minimize the difference between the predicted noise and the sampled noise through reverse optimization. After the training is completed, stop updating the model parameters; finally, use the single-cell cross-modal data conversion model to generate single-cell cross-modal data. The single-cell cross-modal conversion framework proposed by the present invention combines the diffusion model with the low-dimensional latent space, makes full use of the correlation between different modalities, and effectively improves the accuracy and stability of cross-modal conversion through the method of gradually denoising and conditional generation. The model guides data conversion through conditional constraints, can better capture the heterogeneity of cells in different modalities, reduce the influence of batch effects, and improve the robustness to sparse data and noise. BRIEF DESCRIPTION OF THE DRAWINGS

[0055] Figure 1 FIG. is a schematic flow chart of a single-cell cross-modal data conversion model based on a diffusion model. DETAILED DESCRIPTION OF THE INVENTION

[0056] The following further elaborates on the technical solutions and technical effects of the present invention in conjunction with the drawings of the present invention.

[0057] The present invention provides a single-cell cross-modal data conversion method based on a diffusion model, and establishes a single-cell cross-modal data conversion model based on a diffusion model as Figure 1 shown. By introducing the diffusion model and adopting a conditional generation strategy, the probability distribution generation of the target modal data is guided by the source modal data, which can effectively capture the cell heterogeneity between different modalities. Taking single-cell RNA and single-cell ATAC as examples, the mutual conversion of data in two modalities is realized. The specific implementation steps of the method include:

[0058] Step S1, build a single-cell cross-modal data conversion model based on a diffusion model. The model consists of an encoder E r , an encoder E a , a decoder D r , a decoder D a , an RNA diffusion conversion model (the process of converting ATAC to RNA), and an ATAC diffusion conversion model (the process of converting RNA to ATAC);

[0059] Step S2, data preprocessing and pre-training of the autoencoder: preprocess single-cell RNA and single-cell ATAC; use the preprocessed RNA and ATAC for pre-training of the autoencoder, and optimize E r , Ea , D r , D a parameters, and freeze E after the pre-training of the autoencoder ends r , E a , D r , D a parameters;

[0060] Step S3, train the single-cell cross-modal data conversion model:

[0061] The ATAC is input into E a After encoding, a latent representation in the low-dimensional space is obtained The RNA is input into E r After encoding, a latent representation is obtained

[0062] In the process of ATAC converting RNA, the target modality is input into the RNA diffusion conversion model. First, Gaussian noise is gradually added for perturbation diffusion, and then the noise of the source modality is predicted ; In the process of RNA converting ATAC, the target modality is input into the ATAC diffusion conversion model. First, Gaussian noise is gradually added for perturbation diffusion, and then the noise of the source modality is predicted ; Use the Unet network to update the neural network parameters for predicting RNA noise and the neural network parameters for predicting ATAC noise, and minimize the difference between the predicted noise and the sampled noise through reverse optimization;

[0063] After the training of the RNA diffusion conversion model and the ATAC diffusion conversion model ends, stop updating the parameters;

[0064] Step S4, use the trained single-cell cross-modal data conversion model to generate single-cell cross-modal data. Specifically, input the source modality data, obtain the latent representation of the source modality through encoding by the corresponding encoder, and guide the denoising process of the target modality by the latent representation of the source modality to obtain the latent representation of the target modality. After decoding by the decoder of the target modality, the feature expression of the target modality is obtained.

[0065] In Step S2, for different modalities, data preprocessing is performed separately. The preprocessing steps are used to provide a basis for subsequent cross-modal analysis to ensure data quality and analysis consistency, where:

[0066] The preprocessing process of RNA is as follows: First, global normalization is performed. The normalized data is added with 1 and logarithmically transformed, and then the top 3000 highly variable genes are selected to capture the signals that play a key role in cell function;

[0067] The preprocessing process of ATAC is as follows: First, perform binarization and discard the data with the proportion of cells with chromatin open peaks less than 0.5%; use the TF-IDF (Term Frequency-Inverse Document Frequency) algorithm to perform feature transformation on the binarized and filtered peak matrix to obtain the TF-IDF weighted matrix, enhancing the expression differences of the data; finally, perform normalization to scale the values to [0,1]. The TF-IDF algorithm: The term frequency (TF) calculates the open frequency of each peak in the cells; the inverse document frequency (IDF) measures the global specificity of the peak, and TF-IDF strengthens the peak signals that appear frequently in specific cells but are rare globally by calculating the product of the term frequency (TF) and the inverse document frequency (IDF).

[0068] In step S2, train the RNA autoencoder and the ATAC autoencoder respectively:

[0069] In the RNA autoencoder composed of E r and D r , both E r and D r adopt a two-layer fully connected network structure. The number of neurons in the two layers of E r is 256 and 128, and D r is symmetrically set with the structure of E r ;

[0070] The RNA autoencoder updates the network weights using the MSE reconstruction loss function:

[0071]

[0072] In the formula, x r is the input RNA, and D r (E r (x r )) i is the reconstructed data of the RNA autoencoder for the i-th feature . By minimizing the loss, the network can learn an effective low-dimensional representation; m is the total number of features of the RNA; is the loss of the autoencoder composed of E r and D r . The pre-training process uses unsupervised learning to train the data of each modality to ensure that the model effectively captures the key expression information of each single-cell modality, which helps to accelerate the subsequent cross-modal learning speed and improve the performance of the target task.

[0073] In the ATAC autoencoder composed of E a and D a , in order to focus on modeling the gene expression patterns within the same chromosome and reduce the computational burden, E aThe chromosomal features are extended to 32 dimensions and mapped to a 128-dimensional latent space through two layers of LeakyReLU activation functions, D a Feature reconstruction is performed through LeakyReLU, and finally the output is constrained to the interval [0,1] using the Sigmoid function;

[0074] The ATAC autoencoder updates the network weights using the BCE reconstruction loss function:

[0075]

[0076] where x a is the input ATAC, D r (E r (x a )) i is the reconstructed data of the i-th feature by the ATAC autoencoder , and n is the total number of ATAC features; is the loss of the autoencoder composed of E a and D a .

[0077] E a and D a map the data of their respective modalities to a low-dimensional space and respectively, simplify the data distribution and reduce the dimension, thereby improving the feature extraction efficiency. In the pre-training stage, the decoder is only used for self-reconstruction, mapping the low-dimensional features encoded by the encoder back to the original data space to achieve data reconstruction.

[0078] Freeze the autoencoders of each modality in the pre-training stage, E r , E a and D r , D a , and start the training of the diffusion transformation model. The RNA diffusion transformation model and the ATAC diffusion transformation model adopt the same data processing process. The processing of the input data x0 includes a noise addition process and a denoising process:

[0079] The noise addition process uses the diffusion model DDPM. The diffusion process gradually adds noise to the image to turn it into white noise, providing the necessary noise information for the denoising process to restore a clear image. The overall distribution of the diffusion process is expressed as:

[0080]

[0081] where T is the total number of time steps. Gaussian noise is used to gradually perturb x0. By gradually adding Gaussian noise to the input image x0 at each time step t, x0 is finally completely turned into Gaussian noise; x tis the noise state at time t; the diffusion process consists of a Markov chain, and the addition of noise at each time step follows a Gaussian distribution and is represented by Equation (4):

[0082]

[0083] In the formula, represents the cumulative signal retention factor; in which I represents the covariance matrix of the standard normal distribution, and the variance of each dimension is 1; α t is the scheduling weight of the noise, which is scheduled by β t The β t is used to control the amount of noise added to the data at each time step t. β1, …, βt are preset known hyperparameters that control the amount of noise added to the data at the t-th step to ensure stable noise propagation at each time step; the addition of noise at each time step follows Equation (6):

[0084]

[0085] In the formula, T is the total number of time steps, is Gaussian noise that follows the standard normal distribution;

[0086] The denoising process uses the deterministic diffusion model DDIM. By implicitly modeling the generation process at each time step, compared with the DDPM denoising process, DDIM can effectively avoid relying on the Markov chain and allows skip sampling, thus accelerating the generation process. The conditional distribution of DDIM is modeled as:

[0087]

[0088] In the formula, σ t is the noise perturbation factor;

[0089] Use the neural network ∈ θ (x t ) to estimate x0, and the estimation result The expression is:

[0090]

[0091] In the formula, represents the predicted noise at time step t;

[0092] The generation process is modeled at each time step through the conditional probability distribution of DDIM:

[0093]

[0094] According to the sample x t at the current time step and ∈ θ (xt ) Generate the sample \(x\) according to formula (11). t-1 :[[]]

[0095]

[0096]

[0097] In the formula, the diffusion model DDIM is deterministic sampling. Let \(\eta = 0\). Correspondingly, \(\sigma\) t = 0, making the sample generation process a deterministic process; \(\epsilon\) t represents the noise at the current time step;

[0098] During the conversion process, \(x_0\) represents the distribution of the diffusion process at modeled as \(x_0\) represents the distribution of the diffusion process at modeled as Use the Unet network to update the neural network parameters \(\theta\) for predicting RNA noise in the model r and the neural network parameters \(\theta\) for predicting ATAC noise a , and minimize the difference between the predicted noise and the sampled noise through reverse optimization:

[0099]

[0100] In the formula, \(L\) θ (\(\theta\) r , \(\theta\) a ) is the loss function, \(\epsilon\) r is the standard Gaussian noise added to the RNA modality, \(\epsilon\) a is the standard Gaussian noise added to the ATAC modality;

[0101] Since the denoising process of DDPM has a large time overhead, the present invention adopts a more efficient deterministic diffusion model (DDIM) for denoising. By setting the sampling parameter \(\eta\) to 0, the inference process becomes deterministic sampling, thus improving the computational efficiency. At each time step \(t\), the reverse inference process is expressed as:

[0102]

[0103] In the formula, is the latent state of the RNA modality at time \(t - 1\), is the latent state of the ATAC modality at time \(t - 1\); represents the neural network for predicting RNA noise, represents the neural network for predicting ATAC noise;

[0104] In the inference stage of cross-modal conversion, the model no longer updates the parameters, but uses the trained model parameters for inference conversion. Define the expressions of the low-dimensional features for the RNA-to-ATAC conversion process and the ATAC-to-RNA conversion process:

[0105]

[0106] Wherein, the converted low-dimensional features are obtained through step-by-step denoising and

[0107] The present invention can train the model using a training set (including source / target modality paired data) to establish a modality mapping relationship, use a validation set for hyperparameter tuning, and use a test set to verify the denoising conversion process. The target modality data is not used in the test stage, and the target modality data is completely generated under the guidance of the source domain modality:

[0108] In step S4, for the obtained target modality low-dimensional features, each of the two decoders in the single-cell cross-modal data conversion model respectively outputs the autoencoder reconstruction result and the low-dimensional feature conversion result, as shown in formulas (18)-(21):

[0109]

[0110] Since the noise of the diffusion model may cause fluctuations in the results, in order to reduce this influence, a multiple sampling strategy is introduced. By multiple sampling, the influence of noise fluctuations in single sampling can be effectively reduced, thereby improving the smoothness of the conversion process and the robustness of the results. Finally, the mean of these sampling results is taken as the final low-dimensional feature prediction result Described by formula (22):

[0111]

[0112] Wherein, N is the number of samplings, represents the low-dimensional features predicted using the data obtained from the i-th sampling.

[0113] The present invention realizes cross-modal conversion by introducing a diffusion model in the low-dimensional space and Specifically, the latent representation encoded from single-cell RNA is converted into the latent representation of single-cell ATAC, and then the decoder D a of single-cell ATAC is used for decoding to reconstruct the feature expression Vice versa. Since single-cell data of different modalities represent biological information of cells at different levels, conditional constraints are introduced during the conversion process, and the low-dimensional features Z0 of the original data are passed as conditions to the diffusion model. In this way, the model can not only find the conversion path that conforms to the modal topological structure in the latent space to help understand the functional state of cells and gene regulatory mechanisms, but also ensure the accuracy of the results.

[0114] In view of the deficiencies of the prior art, the present invention provides a novel single-cell cross-modal conversion framework that combines a diffusion model with a low-dimensional latent space, aiming to efficiently and accurately achieve data conversion and improve the utilization value of single-modal data. By introducing a diffusion model, the common mode collapse problem in traditional generative models is effectively avoided. Through a step-by-step denoising process, the diffusion model not only enhances the diversity of the generated data, but also improves the richness and quality of cross-modal conversion. Different from other methods in the prior art, the present invention generates the probability distribution of the target-modal data, which can effectively capture the cell heterogeneity between different modalities, reduce the influence brought by technical biases such as batch effects, and enhance the robustness of the model to noise and sparse data. To further improve the expression ability of the model, the present invention adopts a conditional generation strategy, so that the model can still efficiently and accurately complete cross-modal conversion even in the case of complex modal relationships or partial data loss. Finally, the present invention introduces a denoising strategy of multiple sampling, and this method includes the following steps: preprocessing the single-cell sequencing data of different data domains respectively to ensure the quality and usability of the data; for each data domain, constructing an independent low-dimensional space respectively, and using different autoencoders to map the distribution of each data domain to a relatively simple data distribution in the low-dimensional space; introducing a diffusion model in the low-dimensional space for cross-modal conversion, and converting the latent representation Z S encoded from the source domain into the latent representation Z' T of the target domain, and then using the decoder of the target domain for decoding to reconstruct the feature expression X T of the target domain, and vice versa, so as to achieve cross-modal conversion of single-cell data. By constructing an independent low-dimensional space, the present invention enables each data domain to fully retain its unique structural features, and at the same time combines the modal conversion ability of the diffusion model to achieve efficient conversion of single-cell cross-modal data. By performing multiple sampling on the generated data, the noise fluctuations that may be generated by single sampling are effectively reduced, ensuring the smoothness and accuracy of the generated data and further improving the performance of the model in practical applications.

[0115] According to the embodiments disclosed by the present invention, the present invention also provides an electronic device, a readable storage medium, and a computer program product. The electronic device is intended to represent various forms of digital computers. The device includes a computing unit that can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) or a computer program loaded from a storage unit into a random access memory (RAM). In the RAM, various programs and data required for the operation of the device can also be stored. The computing unit, the ROM, and the RAM are connected to each other via a bus. An input / output (I / O) interface is also connected to the bus.

[0116] Multiple components in the device are connected to the I / O interface, including: an input unit, such as a keyboard, a mouse, etc.; an output unit, such as various types of displays, speakers, etc.; a storage unit, such as a disk, an optical disc, etc.; and a communication unit, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit allows the device to exchange information / data with other devices through a computer network such as the Internet and / or various telecommunication networks.

[0117] The computing unit can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any appropriate processor, controller, microcontroller, etc. The computing unit executes the various methods and processes described above, such as the single-cell cross-modal data conversion method based on a diffusion model. For example, in some embodiments, the single-cell cross-modal data conversion method based on a diffusion model can be implemented as a computer software program that is tangibly included in a machine-readable medium, such as a storage unit. In some embodiments, part or all of the computer program can be loaded and / or installed onto the device via the ROM and / or the communication unit. When the computer program is loaded into the RAM and executed by the computing unit, one or more steps of the single-cell cross-modal data conversion method based on a diffusion model described above can be executed. Alternatively, in other embodiments, the computing unit can be configured to execute the single-cell cross-modal data conversion method based on a diffusion model by any other appropriate means (e.g., by means of firmware).

[0118] The program code for implementing the method disclosed in the present invention can be written in any combination of one or more programming languages. These program codes can be provided to a processor or a controller of a general-purpose computer, a special-purpose computer, or other programmable data processing devices, such that when the program codes are executed by the processor or the controller, the functions / operations specified in the flowchart and / or the block diagram are implemented. The program codes can be executed entirely on the machine, partially on the machine, executed partially on the machine and partially on a remote machine as an independent software package, or executed entirely on a remote machine or a server.

[0119] The machine-readable medium can be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. The machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium.

[0120] It should be understood that various forms of the processes shown above can be used, steps can be reordered, added, or deleted. For example, the steps recited in the present disclosure can be executed in parallel, sequentially, or in a different order, as long as the desired results of the technical solution disclosed in the present invention can be achieved, and no limitation is made herein.

[0121] The above-disclosed are only the preferred embodiments of the present invention. Of course, the scope of the rights of the present invention cannot be limited thereby. Those of ordinary skill in the art can understand all or part of the processes of implementing the above embodiments, and the equivalent changes made according to the claims of the present invention still fall within the scope covered by the invention.

Claims

1. A method for single-cell cross-modal data conversion based on a diffusion model, characterized in that Including: Step S1, build a single-cell cross-modal data conversion model based on the diffusion model. Taking single-cell RNA and ATAC as examples, the model consists of encoder E r 、encoder E a 、decoder D r 、decoder D a 、RNA diffusion conversion model, and ATAC diffusion conversion model; Step S2, data preprocessing and pre-training of the autoencoder: Preprocess single-cell RNA and single-cell ATAC; use the preprocessed RNA and ATAC for pre-training of the autoencoder to optimize the parameters of E, E, D, D, and freeze the parameters of E, E, D, D after the pre-training of the autoencoder ends; r E a D r D a and freeze the parameters of E, E, D, D after the pre-training of the autoencoder ends; r E a D r D a ; Step S3, training the single-cell cross-modal data conversion model: ATAC input E a After encoding, a latent representation in the low-dimensional space is obtained RNA input E r After encoding, a latent representation is obtained ATAC conversion RNA process, target modality Input the RNA diffusion conversion model, first gradually add Gaussian noise for perturbation diffusion, and then from the source modality Predict the noise; RNA to ATAC conversion process, target modality Input the ATAC diffusion conversion model, first gradually add Gaussian noise for perturbed diffusion, and then from the source modality Predict the noise; Use the Unet network to update the neural network parameters for predicting RNA noise and the neural network parameters for predicting ATAC noise, and minimize the difference between the predicted noise and the sampled noise through reverse optimization; After the training of the RNA diffusion conversion model and the ATAC diffusion conversion model ends, stop updating the parameters; Step S4, using the trained single-cell cross-modal data conversion model to generate single-cell cross-modal data. Specifically, input the source-modal data, encode it through the corresponding encoder to obtain the latent representation of the source modality, use the latent representation of the source modality to guide the denoising process of the target modality, obtain the latent representation of the target modality, and after decoding by the decoder of the target modality, obtain the feature expression of the target modality.

2. The single-cell cross-modal data conversion method based on the diffusion model according to claim 1, wherein: The preprocessing process of RNA is as follows: First, perform global normalization, add 1 to the normalized data and perform logarithmic transformation, and then select the top 3000 highly variable genes; The preprocessing process of ATAC is as follows: First, perform binarization, and discard the data with the proportion of cells with chromatin open peaks less than 0.5%; use the TF-IDF (term frequency-inverse document frequency) algorithm to perform feature transformation on the binarized and filtered peak matrix to obtain the TF-IDF weighted matrix; finally, perform normalization processing to scale the values to [0, 1].

3. The single-cell cross-modal data conversion method based on a diffusion model according to claim 2, wherein Consisting of E r and D r in the RNA autoencoder, both E r and D r adopt a two-layer fully connected network structure. The number of neurons in the two layers of E r is 256 and 128, and D r is symmetrically configured with E r in structure; The RNA autoencoder updates the network weights using the MSE reconstruction loss function: where x r is the input RNA, and D r (E r (x r )) i is the reconstructed data of the RNA autoencoder for the i-th feature , and m is the total number of features of the RNA; is the loss of the autoencoder composed of E r and D r .

4. The single-cell cross-modal data conversion method based on a diffusion model according to claim 3, wherein, Consisting of E a and D a in the ATAC autoencoder, E a extends the chromosome features to 32 dimensions, maps them to a 128-dimensional latent space through two layers of the LeakyReLU activation function, and D a performs feature reconstruction through LeakyReLU, and finally uses the Sigmoid function to constrain the output to the interval [0, 1]; The ATAC autoencoder updates the network weights using the BCE reconstruction loss function: where x a is the input ATAC, and D r (E r (x a )) i is the reconstructed data of the ATAC auto - encoder for the i - th feature , and n is the total number of features of the ATAC; is the loss of the auto - encoder composed of E a and D a .

5. The single-cell cross-modal data conversion method based on a diffusion model according to claim 4, wherein, The RNA diffusion conversion model and the ATAC diffusion conversion model adopt the same data processing process. The processing of the input data x0 includes a noise addition process and a denoising process; The noise addition process adopts the diffusion model DDPM, gradually perturbs x0 with Gaussian noise, and by gradually adding Gaussian noise to x0 at each time step t, completely turns x0 into Gaussian noise. The diffusion process expression is: where T is the total number of time steps, and x t is the noise state at time t; α t is the scheduling weight of the noise, and β t is used to control the amount of noise added to the data at each time step t; where I represents the covariance matrix of the standard normal distribution, with a variance of 1 for each dimension; is Gaussian noise that follows the standard normal distribution, represents the cumulative signal retention factor; The denoising process adopts the deterministic diffusion model DDIM, and the conditional distribution of DDIM is modeled as: where σ t is the noise disturbance factor; Estimate x0 using the neural network ∈ θ (x t ) and the estimation result is expressed as: wherein, represents the predicted noise at time step t; The generation process is modeled by the conditional probability distribution of DDIM at each time step: According to the sample x at the current time step t and ∈ θ (x t ), generate the sample x t-1 : In the formula, the diffusion model DDIM is deterministic sampling. Let η = 0. Correspondingly, σ t = 0, making the sample generation process a deterministic process; ∈ t represents the noise at the current time step; During the conversion process, x0 represents the distribution of the diffusion process is modeled as x0 represents the distribution of the diffusion process is modeled as Update the neural network parameters θ for predicting RNA noise and the neural network parameters θ for predicting ATAC noise in the model using the Unet network r and minimize the difference between the predicted noise and the sampled noise through backpropagation optimization: a ​ where L θ (θ r , θ a ) is the loss function, ε r is the standard Gaussian noise added to the RNA modality, and ε a is the standard Gaussian noise added to the ATAC modality; At each time step t, the reverse inference process is expressed as: wherein, is the latent state of the RNA modality at time t-1, is the latent state of the ATAC modality at time t-1; represents a neural network for predicting RNA noise, represents a neural network for predicting ATAC noise; In the inference stage of cross-modal conversion, define the expressions of the low-dimensional features of the RNA-to-ATAC process and the ATAC-to-RNA process: In the formula, the converted low-dimensional features are obtained through gradual denoising and 6. The method for single-cell cross-modal data conversion based on a diffusion model according to claim 5, wherein In the single-cell cross-modal data conversion model, each of the two decoders outputs the autoencoder reconstruction result and the low-dimensional feature conversion result respectively:

7. The single-cell cross-modal data conversion method based on a diffusion model according to claim 6, wherein, After the step S4, it further includes: For the source modal data, a multiple sampling strategy is adopted for cross-modal data conversion, and the mean value of the sampling results is used as the final low-dimensional feature prediction result The expression is as follows: where N is the number of sampling times, represents the low-dimensional feature obtained by prediction using the data obtained from the i-th sampling.

8. An electronic device, comprising: At least one processor; And a memory communicatively connected to the at least one processor; wherein, the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the method according to any one of claims 1-7.

9. A non-transitory computer-readable storage medium storing computer instructions, wherein, The computer instructions are used to cause the computer to execute the method according to any one of claims 1-7.

10. A computer program product, including a computer program, which when executed by a processor implements the method according to any one of claims 1-7.