Translation model training method and apparatus, storage medium, and computer program product

By learning the abstract features and regulatory relationships of the transcriptome and proteome through self-supervised pre-training, a translation model is constructed, which solves the problem of low prediction efficiency of intracellular proteome data in single cells and achieves efficient prediction of single-cell proteome data.

CN116994658BActive Publication Date: 2026-05-01TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
TENCENT TECHNOLOGY (SHENZHEN) CO LTD
Filing Date
2022-09-28
Publication Date
2026-05-01

AI Technical Summary

Technical Problem

In existing technologies, labeling single-cell intracellular proteome data requires a lot of resources, resulting in low training efficiency of prediction models and an inability to effectively predict single-cell intracellular proteome data.

Method used

By using self-supervised pre-training based on multi-tissue cell datasets, the general abstract features and regulatory relationships of transcriptomics and proteomics are learned, a translation model is constructed, and fine-tuned on single-cell data that conforms to the pairing relationship, so as to achieve prediction from single-cell transcriptomics data to whole proteomics data.

Benefits of technology

It improves the training efficiency and generalization ability of the model, reduces the dependence on large amounts of labeled data, breaks the limitations of single-cell sequencing technology, and realizes the prediction of single-cell whole proteome data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116994658B_ABST
    Figure CN116994658B_ABST
Patent Text Reader

Abstract

The application discloses a translation model training method and device, a storage medium and a computer program product, and relates to the field of machine learning. The method comprises the following steps: acquiring a multi-tissue cell dataset comprising multi-tissue cell transcriptome data and multi-tissue cell proteome data, and a single-cell dataset comprising single-cell transcriptome data and single-cell proteome data in a first pairing relationship; performing self-supervised training on a transcriptome conversion model and a proteome conversion model based on the multi-tissue cell transcriptome data and the multi-tissue cell proteome data; constructing a translation model; and training the translation model based on the single-cell transcriptome data, the single-cell proteome data and the first pairing relationship, wherein the trained translation model is used for translating single-cell transcriptome data to generate corresponding single-cell whole proteome data, thereby breaking the limitations of single-cell sequencing technology and improving the training efficiency and generalization ability of the model.
Need to check novelty before this filing date? Find Prior Art

Description

Translation model training methods, devices, storage media, and computer program products Technical Field

[0001] This application relates to the field of machine learning, and in particular to a translation model training method, apparatus, storage medium, and computer program product. Background Technology

[0002] Single-cell transcriptome and single-cell proteome data have significant potential applications in clinical medicine. Predicting corresponding single-cell proteome data from single-cell transcriptome data is of great significance for single-cell multi-omics clinical sample analysis, disease mechanism discovery, drug resistance research, and target discovery.

[0003] In related technologies, single-cell transcriptome data is input into a prediction model, which then outputs predicted single-cell surface proteome data. During the training phase, the prediction model is trained using single-cell transcriptome data and single-cell proteome data labeled with cell sequence tags, enabling the prediction of corresponding single-cell protein surface data from single-cell transcriptome data.

[0004] However, the above scheme can only predict single-cell surface proteome data, while the number of single-cell intracellular proteins is huge. Labeling single-cell intracellular proteome data requires a lot of resources, resulting in low sample generation efficiency for training the prediction model and low training efficiency of the prediction model. Summary of the Invention

[0005] This application provides a translation model training method, apparatus, storage medium, and computer program product, capable of predicting corresponding single-cell proteome data from single-cell transcriptome data. The technical solution is as follows.

[0006] On the one hand, a method for training a translation model is provided, the method comprising:

[0007] Obtain a multi-tissue cell dataset, which includes multi-tissue cell transcriptome data and multi-tissue cell proteome data;

[0008] The transcriptome transformation model is trained under self-supervised conditions based on the multi-tissue cell transcriptome data. The transcriptome transformation model includes a transcriptome encoder and a transcriptome decoder.

[0009] The proteome transformation model is trained under self-supervised conditions based on the multi-tissue cell proteome data. The proteome transformation model includes a proteome encoder and a proteome decoder.

[0010] A translation model is constructed that includes the transcriptome encoder and the proteome decoder, and a translation layer is further included between the transcriptome encoder and the proteome decoder;

[0011] Obtain a single-cell dataset, which includes single-cell transcriptome data and single-cell proteome data that have a first pairing relationship;

[0012] Based on the single-cell transcriptome data, the single-cell proteome data, and the first pairing relationship, the translation model is trained. The trained translation model is used to translate the single-cell transcriptome data to generate corresponding single-cell whole proteome data.

[0013] On the other hand, a translation model training device is provided, the device comprising:

[0014] A multi-tissue cell dataset acquisition module is used to acquire a multi-tissue cell dataset, which includes multi-tissue cell transcriptome data and multi-tissue cell proteome data;

[0015] A transcriptome transformation model training module is used to perform self-supervised training of the transcriptome transformation model based on the multi-tissue cell transcriptome data. The transcriptome transformation model includes a transcriptome encoder and a transcriptome decoder.

[0016] A proteome conversion model training module is used to perform self-supervised training of the proteome conversion model based on the multi-tissue cell proteome data. The proteome conversion model includes a proteome encoder and a proteome decoder.

[0017] The translation model construction module is used to construct a translation model that includes the transcriptome encoder and the proteome decoder, and a translation layer is further included between the transcriptome encoder and the proteome decoder;

[0018] A single-tissue cell dataset acquisition module is used to acquire a single-cell dataset, wherein the single-cell dataset includes single-cell transcriptome data and single-cell proteome data that have a first pairing relationship;

[0019] The translation model training module is used to train the translation model based on the single-cell transcriptome data, the single-cell proteome data, and the first pairing relationship. The trained translation model is used to translate the single-cell transcriptome data to generate corresponding single-cell whole proteome data.

[0020] On the other hand, a computer device is provided, the computer device including a processor and a memory, the memory storing at least one instruction, at least one program, code set or instruction set, the at least one instruction, the at least one program, the code set or instruction set being loaded and executed by the processor to implement the translation model training method as described in any of the embodiments of this application above.

[0021] On the other hand, a computer-readable storage medium is provided, wherein at least one instruction, at least one program, code set, or instruction set is stored in the storage medium, wherein the at least one instruction, the at least one program, the code set, or the instruction set is loaded and executed by a processor to implement the translation model training method as described in any of the embodiments of this application above.

[0022] On the other hand, a computer program product or computer program is provided, which includes computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform any of the translation model training methods described in the above embodiments.

[0023] The beneficial effects of the technical solutions provided in this application include at least the following:

[0024] By using self-supervised pre-training based on multi-tissue cell datasets, the model learns general abstract features of the transcriptome, general abstract features of proteins, and the regulatory relationships between the transcriptome and proteome (including membrane surface proteins and intracellular proteins). At the same time, fine-tuning is performed on single-cell transcriptome and proteome data that conform to the pairing relationship, enabling the model to make predictions from single-cell transcriptome data to single-cell proteome data. This breaks through the limitations of existing single-cell sequencing technology. Furthermore, the self-supervised pre-training of the model reduces the dependence on large amounts of labeled data, improving the efficiency of sample generation for training the model, the training efficiency of the model, and the generalization ability of the model. Attached Figure Description

[0025] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0026] Figure 1 is a schematic diagram of the implementation environment provided by an exemplary embodiment of this application;

[0027] Figure 2 is a flowchart of a translation model training method provided in an exemplary embodiment of this application;

[0028] Figure 3 is a schematic diagram of pre-training of a transcriptome transformation model provided in an exemplary embodiment of this application;

[0029] Figure 4 is a schematic diagram of pre-training of a proteome conversion model provided in an exemplary embodiment of this application;

[0030] Figure 5 is a schematic diagram of the translation model structure provided in an exemplary embodiment of this application;

[0031] Figure 6 is a schematic diagram of translation layer pre-training provided in an exemplary embodiment of this application;

[0032] Figure 7 is a flowchart of a pre-trained translation model training method provided in an exemplary embodiment of this application;

[0033] Figure 8 is a schematic diagram of training a pre-trained transcriptome transformation model provided in an exemplary embodiment of this application;

[0034] Figure 9 is a schematic diagram of training a pre-trained proteome conversion model provided in an exemplary embodiment of this application;

[0035] Figure 10 is a schematic diagram of pre-trained translation layer training provided in an exemplary embodiment of this application;

[0036] Figure 11 is a structural diagram of a translation model training device provided in an exemplary embodiment of this application;

[0037] Figure 12 is a structural diagram of a translation model training device provided in an exemplary embodiment of this application;

[0038] Figure 13 is a structural block diagram of a terminal provided in an exemplary embodiment of this application. Detailed Implementation

[0039] To make the objectives, technical solutions, and advantages of this application clearer, the embodiments of this application will be described in further detail below with reference to the accompanying drawings.

[0040] It should be understood that although the terms first, second, etc., may be used in this disclosure to describe various information, such information should not be limited to these terms. These terms are only used to distinguish information of the same type from one another. For example, a first parameter may also be referred to as a second parameter without departing from the scope of this disclosure, and similarly, a second parameter may also be referred to as a first parameter. Depending on the context, the word "if" as used herein may be interpreted as "when," "when," or "in response to determination."

[0041] Proteins are the executors of physiological functions in life activities, directly participating in the body's immune response, immune regulation, and catalytic processes. Differences in protein expression levels more directly reflect the body's physiological state and pathogenesis, making quantitative protein detection crucial for disease diagnosis and prognosis. Traditional tissue protein quantification primarily reflects the average protein expression level within cells. However, due to genetic factors, biochemical noise, and the cellular microenvironment, extensive heterogeneity exists between cells. If we could reflect the direct manifestation of life activities from a single cell, it would not only provide a deeper understanding of the essence of life but also offer strong support for disease diagnosis and treatment. Therefore, single-cell transcriptome and single-cell proteome data have significant application prospects in clinical medicine. Predicting corresponding single-cell proteome data from single-cell transcriptome data is of great significance for single-cell multi-omics clinical sample analysis, disease mechanism discovery, drug resistance research, and target identification. In related technologies, single-cell transcriptome data is input into a prediction model, and the predicted single-cell surface proteome data is output. In this method, the prediction model is trained using single-cell transcriptome and proteome data labeled with cell sequence tags. This allows the model to predict single-cell surface protein data based on the transcriptome data. However, this approach only predicts single-cell surface proteome data. The number of intracellular proteins in single cells is vast, and labeling intracellular proteome data requires significant resources, resulting in low sample generation and training efficiency for the prediction model.

[0042] This application provides a translation model training method. Through self-supervised pre-training based on multi-tissue cell datasets, the model learns general abstract features of the transcriptome, general abstract features of proteins, and the regulatory relationships between the transcriptome and proteome (including membrane surface proteins and intracellular proteins). Simultaneously, fine-tuning is performed on single-cell transcriptome and proteome data that conform to pairing relationships, enabling the model to predict from single-cell transcriptome data to single-cell whole proteome data. This overcomes the limitations of existing single-cell sequencing technologies, making it possible to fill in the entire modality of single-cell proteome data based on single-cell transcriptome data. This is of great significance for single-cell multi-omics clinical sample analysis, disease mechanism discovery, drug resistance research, and target discovery. Furthermore, this application performs self-supervised pre-training on the model, reducing its dependence on large amounts of labeled data, improving the efficiency of sample generation for model training, and enhancing the model's training efficiency and generalization ability.

[0043] First, the implementation environment of this application will be introduced. Please refer to Figure 1, which shows a schematic diagram of the implementation environment provided by an exemplary embodiment of this application. As shown in Figure 1, this implementation environment involves a terminal 110 and a server 120, and the terminal 110 and the server 120 are connected through a communication network 130.

[0044] In some embodiments, terminal 110 is used to send single-cell transcriptome data to server 120. After obtaining a single-cell transcriptome dataset composed of single-cell transcriptome data, server 120 translates the data by inputting it into a translation model to obtain corresponding prediction results, thereby obtaining single-cell proteome data for application in biomedical scenarios, such as: single-cell multi-omics clinical sample analysis, disease mechanism discovery, drug resistance research, and target discovery.

[0045] In some embodiments, the translation model in server 120 is implemented as an encoding layer, a translation layer, and a decoding layer. The encoding layer can be implemented as a software encoder, a hardware encoder, or a wireless encoder, and the decoding layer can be implemented as a software decoder, a hardware decoder, or a wireless decoder.

[0046] In a schematic example, in single-cell multi-omics applications, the encoding layer is implemented as a transcriptome encoder, and the decoding layer is implemented as a proteome decoder. The transcriptome encoder encodes the input single-cell transcriptome data and inputs the output single-cell transcriptome abstract feature representation into the translation layer. The translation layer translates the single-cell transcriptome abstract feature representation into a single-cell proteome abstract feature representation and inputs the single-cell proteome abstract feature representation into the proteome decoder. The proteome decoder decodes the received data and outputs single-cell proteome data as the prediction result.

[0047] In some optional embodiments, terminal 110 can be replaced by a server to perform the same function, and server 120 can be replaced by a terminal to perform the same function; this application does not limit this. The aforementioned terminal is optional and can be a desktop computer, laptop computer, mobile phone, tablet computer, e-book reader, MP3 (Moving Picture Experts Group Audio Layer III) player, MP4 (Moving Picture Experts Group Audio Layer IV) player, smart TV, smart vehicle, and other types of terminal devices; this application does not limit this.

[0048] It is worth noting that the aforementioned servers can be independent physical servers, server clusters or distributed systems composed of multiple physical servers, or cloud servers that provide basic cloud computing services such as cloud services, cloud security, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, content delivery networks (CDN), and big data and artificial intelligence platforms.

[0049] Cloud technology refers to a managed technology that unifies a series of resources such as hardware, software, and networks within a wide area network or local area network to achieve data computing, storage, processing, and sharing.

[0050] In some embodiments, the server described above can also be implemented as a node in a blockchain system.

[0051] It should be noted that the information (including but not limited to user device information, user personal information, etc.), data (including but not limited to data used for analysis, data stored, data displayed, etc.) and signals involved in this application are all authorized by the user or fully authorized by all parties, and the collection, use and processing of related data must comply with the relevant laws, regulations and standards of the relevant countries and regions.

[0052] Please refer to Figure 2 for illustrative purposes. It shows a flowchart of a translation model training method provided in an exemplary embodiment of this application. This method can be applied to a terminal, a server, or both. This embodiment of the application uses the method applied to a terminal as an example for illustration. As shown in Figure 2, the method includes the following steps:

[0053] Step 210: Obtain a multi-tissue cell dataset.

[0054] The multi-tissue cell dataset includes multi-tissue cell transcriptome data and multi-tissue cell proteome data.

[0055] Transcriptome data in the multi-tissue cell dataset is used to represent the total expression level of RNA in multiple cells, while proteome data in the multi-tissue cell dataset is used to represent the total expression level of protein in multiple cells.

[0056] Step 220: Self-supervised training of the transcriptome transformation model based on multi-tissue cell transcriptome data.

[0057] The transcriptome transformation model includes a transcriptome encoder and a transcriptome decoder. The encoder is a module that encodes and converts signals or data into a signal form suitable for communication, transmission, and storage. The decoder is a module that decodes and restores the signal form used for communication, transmission, and storage back into signals or data. In this embodiment, the encoder's input is raw data, and its output is an intermediate feature representation; the decoder's input is the intermediate feature representation, and its output is the reconstructed data corresponding to the raw data. The transcriptome encoder encodes and converts input transcriptome data, such as RNA sequences, and outputs corresponding abstract feature representations, such as atomic sequences of the same length as the RNA sequence that contain the abstract features of RNA. The transcriptome decoder decodes and restores the input abstract feature representations, and outputs corresponding reconstructed transcriptome data, such as RNA sequences of the same length as the input RNA sequence that contain the corresponding abstract features of RNA. Optionally, the transcriptome encoder and transcriptome decoder can be implemented as hardware modules or software modules.

[0058] In some optional embodiments, the transcriptome encoder and transcriptome decoder may include a self-attention layer and a feedforward layer to implement an attention mechanism, automatically learn and calculate the contribution of input data to output data, reduce the loss between transcriptome data and transcriptome reconstructed data, and improve the accuracy of the transformation model.

[0059] In some optional embodiments, the transcriptome transformation model enables the transcriptome encoder to learn to encode multi-tissue cell transcriptome data into a general transcriptome abstract feature representation through self-supervised training based on multi-tissue cell transcriptome data, and the transcriptome decoder to learn to decode the general transcriptome abstract features into multi-tissue cell transcriptome reconstructed data.

[0060] In some optional embodiments, step 220 may be implemented as follows: inputting multi-tissue cell transcriptome data into a transcriptome encoder and outputting an abstract representation of transcriptome features; inputting the abstract representation of transcriptome features into a transcriptome decoder and outputting multi-tissue cell transcriptome reconstructed data; and training a transcriptome transformation model based on the differences between the multi-tissue cell transcriptome data and the multi-tissue cell transcriptome reconstructed data.

[0061] For illustrative purposes, please refer to Figure 3. Figure 3 is a schematic diagram of the pre-training of a transcriptome transformation model provided in an exemplary embodiment of this application. As shown in Figure 3, the transcriptome transformation model 300 includes a transcriptome encoder 310 and a transcriptome decoder 320. Multi-tissue cell transcriptome data 301 is input into the transcriptome encoder 310, and the transcriptome abstract feature representation 302 is output. The transcriptome abstract feature representation 302 is input into the transcriptome decoder 320, and the multi-tissue cell transcriptome reconstructed data 303 is output. Based on the differences between the multi-tissue cell transcriptome data 301 and the multi-tissue cell transcriptome reconstructed data 303, the transcriptome transformation model 300 is trained to optimize the target, i.e., the transcriptome encoder 310 and the transcriptome decoder 320.

[0062] Optionally, the transcriptome transformation model in pre-training can be optimized based on the lowest loss value between multi-tissue cell transcriptome data and multi-tissue cell transcriptome reconstructed data, or based on the convergence of the loss function between multi-tissue cell transcriptome data and multi-tissue cell transcriptome reconstructed data, or based on the number of training iterations reaching a preset threshold.

[0063] Schematic, the reconstruction loss between input and output, i.e., the reconstruction loss between multi-tissue cell transcriptome data and multi-tissue cell transcriptome reconstructed data, is calculated according to the following formula and used as the training target:

[0064]

[0065] in, It is the transcriptome decoder in the transcriptome transformation model. It is the transcriptome encoder in the transcriptome transformation model. It is input multi-tissue cell transcriptome data. It outputs multi-tissue cell transcriptome reconstructed data. This is intermediate data, i.e., an abstract representation of transcriptome features, where R represents RNA. The training objective is to minimize the mean squared error between multi-tissue cell transcriptome data and multi-tissue cell transcriptome reconstruction data.

[0066] Step 230: Perform self-supervised training of the proteome transformation model based on multi-tissue cell proteome data.

[0067] The proteome transformation model includes a proteome encoder and a proteome decoder. The proteome encoder encodes and transforms the input proteome data, such as protein sequences, and outputs a corresponding abstract feature representation, such as an atomic sequence of the same length as the protein sequence that contains the protein's abstract features. The proteome decoder decodes and restores the input abstract feature representation, outputting the corresponding reconstructed proteome data, such as a protein sequence of the same length as the input protein sequence that contains the protein's abstract features. Optionally, the proteome encoder and decoder can be implemented as hardware or software modules.

[0068] In some optional embodiments, the proteome encoder and proteome decoder may include self-attention layers and feedforward layers to implement an attention mechanism, automatically learn and calculate the contribution of input data to output data, reduce the loss between proteome data and proteome reconstructed data, and improve the accuracy of the proteome transformation model.

[0069] In some optional embodiments, the proteome transformation model enables the proteome encoder to learn to encode multi-tissue cell proteome data into a general proteome abstract feature representation through self-supervised training based on multi-tissue cell proteome data, and enables the proteome decoder to learn to decode the general proteome abstract features into multi-tissue cell proteome reconstructed data.

[0070] In some optional embodiments, step 230 may be implemented as follows: inputting multi-tissue cell proteome data into a proteome encoder and outputting an abstract proteome feature representation; inputting the abstract proteome feature representation into a proteome decoder and outputting multi-tissue cell proteome reconstructed data; and training a proteome transformation model based on the differences between the multi-tissue cell proteome data and the multi-tissue cell proteome reconstructed data.

[0071] For illustrative purposes, please refer to Figure 4. Figure 4 is a schematic diagram of the pre-training of a proteome conversion model provided in an exemplary embodiment of this application. As shown in Figure 4, the proteome conversion model 400 includes a proteome encoder 410 and a proteome decoder 420. Multi-tissue cell proteome data 401 is input into the proteome encoder 410, and the output is a proteome abstract feature representation 402. The proteome abstract feature representation 402 is input into the proteome decoder 420, and the output is multi-tissue cell proteome reconstructed data 403. Based on the differences between the multi-tissue cell proteome data 401 and the multi-tissue cell proteome reconstructed data 403, the proteome conversion model 400 is trained according to the optimization target, namely the proteome encoder 410 and the proteome decoder 420.

[0072] Optionally, the proteome transformation model in pre-training can be optimized based on the lowest loss value between multi-tissue cell proteome data and multi-tissue cell proteome reconstructed data, or based on the convergence of the loss function between multi-tissue cell proteome data and multi-tissue cell proteome reconstructed data, or based on the number of training iterations reaching a preset threshold.

[0073] Schematic, the reconstruction loss between input and output, i.e., the reconstruction loss between multi-tissue cell proteome data and multi-tissue cell proteome reconstructed data, is calculated according to the following formula and used as the training target:

[0074]

[0075] in, It is the proteome decoder in the proteome transformation model. It is the proteome encoder in the proteome transformation model. It is the input multi-tissue cell proteomic data. It outputs multi-tissue cell proteome reconstruction data. It is an abstract feature representation of intermediate data proteome, where P stands for Protein. The training objective is to find the proteome encoder and decoder that minimize the mean squared error between multi-tissue cell proteome data and reconstructed multi-tissue cell proteome data.

[0076] Step 240: Construct a translation model that includes a transcriptome encoder and a proteome decoder.

[0077] The transcriptome encoder and proteome decoder are separated by a translation layer.

[0078] The translation layer is used to translate the transcriptome abstract feature representation output by the transcriptome encoder into a proteome abstract feature representation. The translation model is used to predict the corresponding cellular whole proteome data based on the input cellular transcriptome data.

[0079] For illustrative purposes, please refer to Figure 5. Figure 5 is a schematic diagram of the translation model structure provided in an exemplary embodiment of this application. As shown in Figure 5, the translation model 500 includes a transcriptome encoder 510, a translation layer 520, and a proteome decoder 530. Transcriptome data is input into the transcriptome encoder 510, which encodes the data and inputs the output into the translation layer 520. The translation layer 520 translates the output of the transcriptome encoder 510 and inputs the translated data into the proteome decoder 530. The proteome decoder 530 decodes the translated data and outputs proteome data.

[0080] Step 250: Obtain the single-cell dataset.

[0081] The single-cell dataset includes single-cell transcriptome data and single-tissue cell proteome data that have a first pairing relationship.

[0082] Transcriptome data in single-cell datasets are used to represent the total expression level of RNA in a single cell, while proteome data in single-cell datasets are used to represent the total expression level of proteins in a single cell, including intracellular proteins and membrane surface proteins.

[0083] In some alternative embodiments, RNA and proteins can be mapped one-to-one by calculating Pearson similarity, thus obtaining the first pairing relationship between single-cell transcriptome data and single-cell proteome data in a single-cell dataset.

[0084] Step 260: Train the translation model based on single-cell transcriptome data, single-cell proteome data, and the first pairing relationship.

[0085] The trained translation model is used to translate single-cell transcriptome data to generate corresponding single-cell whole proteome data.

[0086] In some optional embodiments, the input single-cell transcriptome data and the generated single-cell proteome data can be used for single-cell multi-omics clinical sample analysis, disease mechanism discovery, drug resistance research, and target discovery.

[0087] It is worth noting that the above application scenarios are merely illustrative examples and are not intended to limit the scope of this application.

[0088] In summary, the method provided in this embodiment clarifies the translation model training process. Through self-supervised pre-training based on multi-tissue cell datasets, the model learns general abstract features of the transcriptome, general abstract features of proteins, and the regulatory relationships between the transcriptome and proteome (including membrane surface proteins and intracellular proteins). Simultaneously, fine-tuning is performed on single-cell transcriptome and proteome data that conform to pairing relationships, enabling the model to make predictions from single-cell transcriptome data to single-cell proteome data. This breaks through the limitations of existing single-cell sequencing technologies. Furthermore, the self-supervised pre-training of the model reduces the dependence on large amounts of labeled data, improving the efficiency of sample generation for model training, the model training efficiency, and the model's generalization ability.

[0089] The method provided in this embodiment offers a pre-training method for transcriptome transformation models, enabling the models to learn to encode multi-tissue cell transcriptome data into abstract feature representations of multi-tissue cell transcriptomes. This reduces the dependence of the transcriptome transformation model on single-cell transcriptome data and improves the sample generation efficiency and generalization ability of the model.

[0090] The method provided in this embodiment offers a pre-training method for proteome transformation models, enabling the proteome transformation models to learn to decode abstract features of multi-tissue cell proteomes into reconstructed multi-tissue cell proteome data. This reduces the dependence of the proteome transformation models on single-cell proteome data and improves the sample generation efficiency and generalization ability of the proteome transformation models.

[0091] In some optional embodiments, the aforementioned multi-tissue cell dataset includes multi-tissue cell transcriptome data and multi-tissue proteome data with a second pairing relationship. Following step 240, a translation layer pre-training step is also included, i.e., training the translation layer in the translation model based on the second pairing relationship between the multi-tissue cell transcriptome data and the multi-tissue cell proteome data.

[0092] The aforementioned translation layer pre-training steps can be implemented as follows: inputting multi-tissue cell transcriptome data into a transcriptome encoder, which outputs an abstract representation of transcriptome features; translating the abstract representation of transcriptome features into an abstract representation of proteome features through a translation layer; inputting the abstract representation of proteome features into a proteome decoder, which outputs reconstructed multi-tissue cell proteome data; and training the translation layer in the translation model based on the differences between the reconstructed multi-tissue cell proteome data and the reference multi-tissue cell proteome data, wherein the reference multi-tissue cell proteome data is data in the multi-tissue cell dataset that conforms to a second pairing relationship with the input multi-tissue cell transcriptome data.

[0093] Schematic illustration, please refer to Figure 6. Figure 6 is a schematic diagram of translation layer pre-training provided in an exemplary embodiment of this application. As shown in Figure 6, the translation model 600 includes a pre-trained transcriptome encoder 610, a translation layer 620, and a pre-trained proteome decoder 630. Multi-tissue cell transcriptome data 601 is input into the transcriptome encoder 610, and the transcriptome abstract feature representation 602 is output. The translation layer 620 translates the transcriptome abstract feature representation 602 into a proteome abstract feature representation 603. The proteome abstract feature representation 603 is input into the proteome decoder 630, and multi-tissue cell proteome reconstructed data 604 is output. Based on the difference between the multi-tissue cell proteome reconstructed data 604 and the reference multi-tissue cell proteome data 605, the translation layer 620 in the translation model 600 is trained, wherein the multi-tissue cell transcriptome data 601 and the reference multi-tissue cell proteome data 605 conform to a second pairing relationship.

[0094] Optionally, the translation layer in pre-training can be optimized based on the lowest loss value between the reference multi-tissue cell proteome data and the reconstructed multi-tissue cell proteome data, or based on the convergence of the loss function between the reference multi-tissue cell proteome data and the reconstructed multi-tissue cell proteome data, or based on the number of training iterations reaching a preset threshold.

[0095] Schematic, the loss between the input and output is calculated according to the following formula, i.e., the loss between the reference multi-tissue cell proteome data and the multi-tissue cell proteome reconstruction data, and is used as the training target:

[0096]

[0097] in, It is the translation layer in the translation model. It is the proteome decoder in the proteome transformation model. It is the transcriptome encoder in the transcriptome transformation model. It references multi-tissue cell proteomic data. It outputs multi-tissue cell proteome reconstruction data. It is input multi-tissue cell transcriptome data. It is an abstract representation of the transcriptome features of intermediate data. It is an abstract feature representation of intermediate data proteome. That is, the training objective is to find the translation layer that minimizes the mean squared error between the reference multi-tissue cell proteome data and the reconstructed multi-tissue cell proteome data.

[0098] In summary, the method provided in this embodiment offers a pre-training method for the translation layer, enabling it to learn the regulatory relationships between multi-tissue cell transcriptomes and multi-tissue cell proteomes. This allows the translation of abstract transcriptome feature representations into abstract proteome feature representations, thereby achieving prediction from transcriptome data to proteome data. This reduces the translation layer's dependence on paired single-cell data and improves the sample generation efficiency and generalization ability of the translation layer.

[0099] Figure 7 is a flowchart of a pre-trained translation model training method provided in an exemplary embodiment of this application. As shown in Figure 7, in some optional embodiments, step 260 can be implemented as steps 261 to 263.

[0100] Step 261: Train the transcriptome transformation model based on single-cell transcriptome data to obtain the transcriptome encoder.

[0101] In some optional embodiments, step 261 may be implemented as follows: inputting single-cell transcriptome data into a transcriptome encoder and outputting an abstract feature representation of the single-cell transcriptome; inputting the abstract feature representation of the single-cell transcriptome into a transcriptome decoder and outputting a reconstructed data representation of the single-cell transcriptome; training the transcriptome transformation model based on the difference between the input single-cell transcriptome data and the reconstructed single-cell transcriptome data to obtain a transcriptome encoder.

[0102] For illustrative purposes, please refer to Figure 8. Figure 8 is a schematic diagram of the training of a pre-trained transcriptome transformation model provided in an exemplary embodiment of this application. As shown in Figure 8, the pre-trained transcriptome transformation model 800 includes a pre-trained transcriptome encoder 810 and a pre-trained transcriptome decoder 820. Single-cell transcriptome data 801 is input into the pre-trained transcriptome encoder 810, which outputs a single-cell transcriptome abstract feature representation 802. The abstract feature representation 802 is input into the pre-trained transcriptome decoder 820, which outputs single-cell transcriptome reconstructed data 803. Based on the difference between the single-cell transcriptome data 801 and the single-cell transcriptome reconstructed data 803, the pre-trained transcriptome transformation model 800 is trained to obtain the trained transcriptome encoder, which is based on the optimization target, i.e., the pre-trained transcriptome encoder 810 and the pre-trained transcriptome decoder 820.

[0103] Optionally, the transcriptome transformation model during training can be optimized based on the lowest loss value between single-cell transcriptome data and single-cell transcriptome reconstructed data, or based on the convergence of the loss function between single-cell transcriptome data and single-cell transcriptome reconstructed data, or based on the number of training iterations reaching a preset threshold.

[0104] Schematic, the reconstruction loss between input and output is calculated according to the following formula, i.e., the reconstruction loss between single-cell transcriptome data and single-cell transcriptome reconstructed data, and is used as the training target:

[0105]

[0106] in, It is the pre-trained transcriptome decoder in the pre-trained transcriptome transformation model. It is the pre-trained transcriptome encoder in the pre-trained transcriptome transformation model. It is the input single-cell transcriptome data. It is the output single-cell transcriptome reconstruction data. This refers to intermediate data, specifically an abstract representation of single-cell transcriptome features, where R represents RNA and scRNA represents single-cell RNA. The training objective is to find the transcriptome encoder and decoder that minimize the mean squared error between the single-cell transcriptome data and the reconstructed single-cell transcriptome data.

[0107] Step 262: Train the proteome transformation model based on single-cell proteome data to obtain the proteome decoder.

[0108] In some optional embodiments, step 262 may be implemented as follows: inputting single-cell proteome data into a proteome encoder and outputting an abstract feature representation of the single-cell proteome; inputting the abstract feature representation of the single-cell proteome into a proteome decoder and outputting a reconstructed data representation of the single-cell proteome; training the proteome transformation model based on the difference between the input single-cell proteome data and the reconstructed single-cell proteome data to obtain a proteome decoder.

[0109] For illustrative purposes, please refer to Figure 9. Figure 9 is a schematic diagram of the training of a pre-trained proteome conversion model provided by an exemplary embodiment of this application. As shown in Figure 9, the pre-trained proteome conversion model 900 includes a pre-trained proteome encoder 910 and a pre-trained proteome decoder 920. Single-cell proteome data 901 is input into the pre-trained proteome encoder 910, which outputs a single-cell proteome abstract feature representation 902. The single-cell proteome abstract feature representation 902 is input into the pre-trained proteome decoder 920, which outputs single-cell proteome reconstructed data 903. Based on the difference between the single-cell proteome data 901 and the single-cell proteome reconstructed data 903, the proteome conversion model 900 is trained to obtain the trained proteome decoder, based on the difference between the pre-trained proteome encoder 910 and the pre-trained proteome decoder 920, targeting the optimization objective.

[0110] Optionally, the proteome transformation model in training can be optimized based on the lowest loss value between single-cell proteome data and single-cell proteome reconstructed data, or based on the convergence of the loss function between single-cell proteome data and single-cell proteome reconstructed data, or based on the number of training iterations reaching a preset threshold.

[0111] Schematic, the reconstruction loss between input and output, i.e., the reconstruction loss between single-cell proteome data and single-cell proteome reconstructed data, is calculated according to the following formula and used as the training objective:

[0112]

[0113] in, It is the pre-trained proteome decoder in the pre-trained proteome transformation model. It is the pre-trained proteome encoder in the pre-trained proteome transformation model. It is the input single-cell proteome data. It outputs single-cell proteome reconstruction data. This refers to intermediate data, specifically abstract feature representations of single-cell proteomes. P represents Protein, and scP represents single-cell protein. The training objective is to find the proteome encoder and decoder that minimize the mean squared error between the single-cell proteome data and the reconstructed single-cell proteome data.

[0114] Step 263: Based on the first pairing relationship, transcriptome encoder, and proteome decoder, train the translation layer of the translation model.

[0115] In some optional embodiments, step 263 may be implemented as follows: inputting single-cell transcriptome data into a transcriptome encoder and outputting an abstract feature representation of the single-cell transcriptome; translating the single-cell transcriptome feature representation into an abstract feature representation of the single-cell proteome through a translation layer; inputting the abstract feature representation of the single-cell proteome into a proteome decoder and outputting reconstructed single-cell proteome data; training the translation layer based on the difference between the reconstructed single-cell proteome data and the reference single-cell proteome data, wherein the reference single-cell proteome data is the data in the single-cell dataset that conforms to a first pairing relationship with the input single-cell transcriptome data.

[0116] For illustrative purposes, please refer to Figure 10. Figure 10 is a schematic diagram of pre-trained translation layer training provided in an exemplary embodiment of this application. As shown in Figure 10, the translation model 1000 includes a pre-trained transcriptome encoder 1010, a pre-trained translator 1020, and a pre-trained proteome decoder 1030. Single-cell transcriptome data 1001 is input into the pre-trained transcriptome encoder 1010, which outputs a single-cell transcriptome abstract feature representation 1002. The pre-trained translation layer 1020 processes the single-cell transcriptome data. The abstract feature representation 1002 is translated into a single-cell proteome abstract feature representation 1003; the single-cell proteome abstract feature representation 1003 is input into the pre-trained proteome decoder 1030, and single-cell proteome data 1004 is output; based on the difference between single-cell proteome data 1004 and reference single-cell proteome data 1005, the translation layer 1020 in the translation model 1000 is trained, wherein the single-cell transcriptome data 1001 and the reference single-cell proteome data 1005 conform to the first pairing relationship.

[0117] Optionally, the translation layer in training can be optimized based on the lowest loss value between the reference single-cell proteome data and the single-cell proteome data, or based on the convergence of the loss function between the reference single-cell proteome data and the single-cell proteome data, or based on the number of training iterations reaching a preset threshold.

[0118] Indicatively, the loss between the input and output is calculated using the following formula, i.e., the loss between the reference single-cell proteome data and the single-cell proteome data, and is used as the training target:

[0119]

[0120] in, It is the pre-trained translation layer in the pre-trained translation model. It is the pre-trained proteome decoder in the pre-trained proteome transformation model. It is the pre-trained transcriptome encoder in the pre-trained transcriptome transformation model. It references single-cell proteomics data. It outputs single-cell proteomic data. It is the input single-cell transcriptome data. It is intermediate data, namely, an abstract representation of the features of a single-cell transcriptome. This is intermediate data, specifically an abstract representation of single-cell proteome features. scRNA represents single-cell RNA, and scP represents single-cell protein. The training objective is to find the translation layer that minimizes the mean squared error between the reference single-cell proteome data and the actual single-cell proteome data.

[0121] In summary, the method provided in this embodiment clarifies the training process of a pre-trained translation model. By training the transcriptome conversion model and the proteome conversion model separately based on a single-cell dataset, the transcriptome encoder learns the abstract features of the single-cell transcriptome, and the proteome decoder learns the abstract features of the single-cell proteome. Based on the trained transcriptome encoder and proteome decoder, the translation layer learns the regulatory relationship between the single-cell transcriptome and the single-cell proteome through paired single-cell data, thereby achieving prediction from the single-cell transcriptome to the single-cell proteome. It can fill in and integrate missing proteomes in all reported single-cell transcriptome clinical samples, effectively making up for the limitations of current single-cell multi-omics sequencing technology and improving the utilization rate of valuable clinical data.

[0122] The method provided in this embodiment offers a pre-trained transcriptome transformation model training method, enabling the transcriptome encoder to learn to encode single-cell transcriptome data into abstract feature representations of single-cell transcriptomes, thereby improving the utilization rate of valuable clinical data and increasing the training efficiency of the transcriptome transformation model.

[0123] The method provided in this embodiment offers a pre-trained proteome transformation model training method, enabling the proteome decoder to learn to decode the abstract feature representation of single-cell proteome into single-cell proteome reconstructed data, thereby improving the utilization rate of valuable clinical data and increasing the training efficiency of the proteome transformation model.

[0124] The method provided in this embodiment offers a pre-training-based translation layer training approach, enabling the translation layer to learn the regulatory relationships between single-cell transcriptomes and single-cell proteomes. This allows the abstract feature representation of the transcriptome to be translated into an abstract feature representation of the proteome, achieving prediction from single-cell transcriptome to single-cell proteome. This effectively compensates for the limitations of current single-cell multi-omics sequencing technologies and improves the training efficiency of the model.

[0125] Figure 11 is a structural block diagram of a translation model training device provided in an exemplary embodiment of this application. As shown in Figure 11, the device includes the following parts:

[0126] The multi-tissue cell dataset acquisition module 1110 is used to acquire a multi-tissue cell dataset, which includes multi-tissue cell transcriptome data and multi-tissue cell proteome data.

[0127] Transcriptome transformation model training module 1120 is used to perform self-supervised training of the transcriptome transformation model based on the multi-tissue cell transcriptome data, wherein the transcriptome transformation model includes a transcriptome encoder and a transcriptome decoder.

[0128] The proteome conversion model training module 1130 is used to perform self-supervised training of the proteome conversion model based on the multi-tissue cell proteome data. The proteome conversion model includes a proteome encoder and a proteome decoder.

[0129] Translation model construction module 1140 is used to construct a translation model that includes the transcriptome encoder and the proteome decoder, and a translation layer is further included between the transcriptome encoder and the proteome decoder;

[0130] The single-tissue cell dataset acquisition module 1150 is used to acquire a single-cell dataset, which includes single-cell transcriptome data and single-cell proteome data that have a first pairing relationship.

[0131] The translation model training module 1160 is used to train the translation model based on the single-cell transcriptome data, the single-cell proteome data, and the first pairing relationship. The trained translation model is used to translate the single-cell transcriptome data to generate corresponding single-cell whole proteome data.

[0132] In some optional embodiments, the transcriptome transformation model training module 1120 is used to input the multi-tissue cell transcriptome data into the transcriptome encoder and output an abstract transcriptome feature representation; input the abstract transcriptome feature representation into the transcriptome decoder and output multi-tissue cell transcriptome reconstructed data; and train the transcriptome transformation model based on the differences between the multi-tissue cell transcriptome data and the multi-tissue cell transcriptome reconstructed data.

[0133] In some optional embodiments, the proteome transformation model training module 1130 is used to input the multi-tissue cell proteome data into the proteome encoder and output an abstract proteome feature representation; input the abstract proteome feature representation into the proteome decoder and output multi-tissue cell proteome reconstruction data; and train the proteome transformation model based on the difference between the multi-tissue cell proteome data and the multi-tissue cell proteome reconstruction data.

[0134] In some optional embodiments, the apparatus further includes a translation layer training module 1170 for training the translation layer in the translation model based on the second pairing relationship between the multi-tissue cell transcriptome data and the multi-tissue cell proteome data.

[0135] In some optional embodiments, the translation layer training module 1170 is used to input the multi-tissue cell transcriptome data into the transcriptome encoder and output an abstract transcriptome feature representation; translate the abstract transcriptome feature representation into an abstract proteome feature representation through the translation layer; input the abstract proteome feature representation into the proteome decoder and output multi-tissue cell proteome reconstructed data; and train the translation layer in the translation model based on the difference between the multi-tissue cell proteome reconstructed data and the reference multi-tissue cell proteome data, wherein the reference multi-tissue cell proteome data is data in the multi-tissue cell dataset that conforms to the second pairing relationship with the input multi-tissue cell transcriptome data.

[0136] In some optional embodiments, the translation model training module 1160 includes:

[0137] Transcriptome encoder training unit 1161 is used to train the transcriptome transformation model based on the single-cell transcriptome data to obtain the transcriptome encoder;

[0138] The proteome decoder training unit 1162 is used to train the proteome transformation model based on the single-cell proteome data to obtain the proteome decoder.

[0139] Translation layer training unit 1163 is used to train the translation layer of the translation model based on the first pairing relationship, the transcriptome encoder and the proteome decoder.

[0140] In some optional embodiments, the transcriptome encoder training unit 1161 is used to input the single-cell transcriptome data into the transcriptome encoder and output a single-cell transcriptome abstract feature representation; input the single-cell transcriptome abstract feature into the transcriptome decoder and output a single-cell transcriptome reconstructed data representation; and train the transcriptome transformation model based on the difference between the input single-cell transcriptome data and the single-cell transcriptome reconstructed data to obtain the transcriptome encoder.

[0141] In some optional embodiments, the proteome decoder training unit 1162 is used to input the single-cell proteome data into the proteome encoder and output a single-cell proteome abstract feature representation; input the single-cell proteome abstract features into the proteome decoder and output a single-cell proteome reconstructed data representation; and train the proteome transformation model based on the difference between the input single-cell proteome data and the single-cell proteome reconstructed data to obtain the proteome decoder.

[0142] In some optional embodiments, the translation layer training unit 1163 is used to input the single-cell transcriptome data into the transcriptome encoder and output a single-cell transcriptome abstract feature representation; through the translation layer, translate the single-cell transcriptome abstract feature representation into a single-cell proteome abstract feature representation; input the single-cell proteome abstract feature representation into the proteome decoder and output single-cell proteome reconstructed data; and train the translation layer based on the difference between the single-cell proteome reconstructed data and the reference single-cell proteome data, wherein the reference single-cell proteome data is data in the single-cell dataset that conforms to the first pairing relationship with the input single-cell transcriptome data.

[0143] In summary, the translation model training device provided in this embodiment enables the model to learn general abstract features of the transcriptome, general abstract features of proteins, and the regulatory relationships between the transcriptome and proteome (including membrane surface proteins and intracellular proteins) through self-supervised pre-training based on multi-tissue cell datasets. Simultaneously, fine-tuning is performed on single-cell transcriptome and proteome data that conform to pairing relationships, allowing the model to make predictions from single-cell transcriptome data to single-cell proteome data. This overcomes the limitations of existing single-cell sequencing technologies. Furthermore, the self-supervised pre-training of the model reduces the dependence on large amounts of labeled data, improving the efficiency of sample generation for model training, the model's training efficiency, and the model's generalization ability.

[0144] It should be noted that the translation model training device provided in the above embodiments is only an example of the division of the above functional modules. In actual applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above.

[0145] Figure 13 shows a structural block diagram of a terminal 1300 provided in an exemplary embodiment of this application. The terminal 1300 may be a smartphone, tablet computer, MP3 player (Moving Picture Experts Group Audio Layer III), MP4 player (Moving Picture Experts Group Audio Layer IV), laptop computer, or desktop computer. The terminal 1300 may also be referred to as a user device, portable terminal, laptop terminal, desktop terminal, or other names.

[0146] Typically, terminal 1300 includes a processor 1301 and a memory 1302.

[0147] Processor 1301 may include one or more processing cores, such as a quad-core processor, an octa-core processor, etc. Processor 1301 may be implemented using at least one hardware form selected from DSP (Digital Signal Processing), FPGA (Field-Programmable Gate Array), and PLA (Programmable Logic Array). Processor 1301 may also include a main processor and a coprocessor. The main processor, also known as a CPU (Central Processing Unit), is used to process data in the wake-up state; the coprocessor is a low-power processor used to process data in the standby state. In some embodiments, processor 1301 may integrate a GPU (Graphics Processing Unit), which is responsible for rendering and drawing the content to be displayed on the screen. In some embodiments, processor 1301 may also include an AI (Artificial Intelligence) processor, which is used to handle computational operations related to machine learning.

[0148] The memory 1302 may include one or more computer-readable storage media, which may be non-transitory. The memory 1302 may also include high-speed random access memory and non-volatile memory, such as one or more disk storage devices or flash memory devices. In some embodiments, the non-transitory computer-readable storage media in the memory 1302 are used to store at least one instruction, which is executed by the processor 1301 to implement the translation model training method provided in the method embodiments of this application.

[0149] In some embodiments, terminal 1300 may also include other components. Those skilled in the art will understand that the structure shown in FIG13 does not constitute a limitation on terminal 1300 and may include more or fewer components than shown, or combine certain components, or adopt different component arrangements.

[0150] Embodiments of this application also provide a computer device, which can be implemented as a terminal or server as shown in FIG1. ​​The computer device includes a processor and a memory, the memory storing at least one instruction, at least one program, code set, or instruction set, the at least one instruction, at least one program, code set, or instruction set being loaded and executed by the processor to implement the translation model training method provided in the above-described method embodiments.

[0151] Embodiments of this application also provide a computer-readable storage medium storing at least one instruction, at least one program, code set, or instruction set, wherein the at least one instruction, at least one program, code set, or instruction set is loaded and executed by a processor to implement the translation model training method provided in the above-described method embodiments.

[0152] Embodiments of this application also provide a computer program product or computer program, which includes computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform any of the translation model training methods described in the above embodiments.

[0153] Optionally, the computer-readable storage medium may include: read-only memory (ROM), random access memory (RAM), solid-state drives (SSDs), or optical discs, etc. The random access memory may include resistive random access memory (ReRAM) and dynamic random access memory (DRAM). The sequence numbers of the embodiments in this application are merely descriptive and do not represent the superiority or inferiority of the embodiments.

[0154] Those skilled in the art will understand that all or part of the steps of the above embodiments can be implemented by hardware or by a program instructing related hardware. The program can be stored in a computer-readable storage medium, such as a read-only memory, a disk, or an optical disk.

[0155] The above description is merely an optional embodiment of this application and is not intended to limit this application. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the protection scope of this application.

Claims

1. A method for training a translation model, characterized in that, The method includes: acquiring a multi-tissue cell dataset, wherein the multi-tissue cell dataset includes multi-tissue cell transcriptome data and multi-tissue cell proteome data; performing self-supervised training on a transcriptome transformation model based on the multi-tissue cell transcriptome data to obtain a pre-trained transcriptome transformation model, wherein the transcriptome transformation model includes a transcriptome encoder and a transcriptome decoder, wherein the transcriptome transformation model is trained based on the difference between the multi-tissue cell transcriptome data and the multi-tissue cell reconstructed transcriptome data, wherein the multi-tissue cell reconstructed transcriptome data is obtained by processing the multi-tissue cell transcriptome data through the transcriptome encoder and the transcriptome decoder, and the training objective includes minimizing the mean squared error between the multi-tissue cell transcriptome data and the multi-tissue cell reconstructed transcriptome data; and performing self-supervised training on a proteome transformation model based on the multi-tissue cell proteome data to obtain a pre-trained proteome transformation model, wherein the proteome transformation model includes a proteome encoder and a proteome decoder, wherein the proteome transformation model is trained based on the difference between the multi-tissue cell proteome data and the multi-tissue cell reconstructed proteome data, wherein the multi-tissue cell proteome reconstructed proteome data is obtained by processing the multi-tissue cell proteome data through the transcriptome encoder and the proteome decoder, and the training objective includes minimizing the mean squared error between the multi-tissue cell transcriptome data and the multi-tissue cell reconstructed proteome data. Based on the outputs of the proteome encoder and the proteome decoder, the training objective includes minimizing the mean squared error between the multi-tissue cell proteome data and the reconstructed multi-tissue cell proteome data; constructing a translation model including a pre-trained transcriptome encoder and a pre-trained proteome decoder, with a translation layer between the transcriptome encoder and the proteome decoder; acquiring a single-cell dataset, which includes single-cell transcriptome data and single-cell proteome data with a first pairing relationship; further training the pre-trained transcriptome conversion model based on the single-cell transcriptome data to obtain a further trained transcriptome encoder; further training the pre-trained proteome conversion model based on the single-cell proteome data to obtain a further trained proteome decoder; and training the translation layer of the translation model based on the first pairing relationship, the further trained transcriptome encoder, and the further trained proteome decoder to obtain a trained translation layer, with the training objective including minimizing the mean squared error between the reconstructed single-cell proteome data and the reference single-cell proteome data. The trained translation model is used to translate the single-cell transcriptome data to generate corresponding single-cell whole proteome data.

2. The method according to claim 1, characterized in that, The step of performing self-supervised training on the transcriptome transformation model based on the multi-tissue cell transcriptome data to obtain a pre-trained transcriptome transformation model includes: inputting the multi-tissue cell transcriptome data into the transcriptome encoder and outputting an abstract representation of transcriptome features; inputting the abstract representation of transcriptome features into the transcriptome decoder and outputting the reconstructed multi-tissue cell transcriptome data; and training the transcriptome transformation model based on the differences between the multi-tissue cell transcriptome data and the reconstructed multi-tissue cell transcriptome data.

3. The method according to claim 1, characterized in that, The step of performing self-supervised training on the proteome transformation model based on the multi-tissue cell proteome data to obtain a pre-trained proteome transformation model includes: inputting the multi-tissue cell proteome data into the proteome encoder and outputting an abstract proteome feature representation; inputting the abstract proteome feature representation into the proteome decoder and outputting the reconstructed multi-tissue cell proteome data; and training the proteome transformation model based on the differences between the multi-tissue cell proteome data and the reconstructed multi-tissue cell proteome data.

4. The method according to claim 1, characterized in that, The multi-tissue cell dataset includes multi-tissue cell transcriptome data and multi-tissue cell proteome data that have a second pairing relationship; After constructing the translation model including the pre-trained transcriptome encoder and the pre-trained proteome decoder, the method further includes: training the translation layer in the translation model based on the second pairing relationship between the multi-tissue cell transcriptome data and the multi-tissue cell proteome data to obtain the pre-trained translation layer. The training objective includes minimizing the mean square error between the multi-tissue cell proteome reconstructed data and the reference multi-tissue cell proteome data.

5. The method according to claim 4, characterized in that, The step of training the translation layer in the translation model based on the second pairing relationship between the multi-tissue cell transcriptome data and the multi-tissue cell proteome data to obtain a pre-trained translation layer includes: inputting the multi-tissue cell transcriptome data into the transcriptome encoder and outputting an abstract transcriptome feature representation; translating the abstract transcriptome feature representation into an abstract proteome feature representation through the translation layer; inputting the abstract proteome feature representation into the proteome decoder and outputting the reconstructed multi-tissue cell proteome data; and training the translation layer in the translation model based on the difference between the reconstructed multi-tissue cell proteome data and the reference multi-tissue cell proteome data to obtain the pre-trained translation layer, wherein the reference multi-tissue cell proteome data is data in the multi-tissue cell dataset that conforms to the second pairing relationship with the input multi-tissue cell transcriptome data.

6. The method according to any one of claims 1 to 5, characterized in that, The step of further training the pre-trained transcriptome transformation model based on the single-cell transcriptome data to obtain a further trained transcriptome encoder includes: inputting the single-cell transcriptome data into the pre-trained transcriptome encoder and outputting an abstract feature representation of the single-cell transcriptome; inputting the abstract feature representation of the single-cell transcriptome into the pre-trained transcriptome decoder and outputting a reconstructed data representation of the single-cell transcriptome; and further training the pre-trained transcriptome transformation model based on the difference between the input single-cell transcriptome data and the reconstructed single-cell transcriptome data to obtain the further trained transcriptome encoder, wherein the training objective includes minimizing the mean squared error between the single-cell transcriptome data and the reconstructed single-cell transcriptome data.

7. The method according to any one of claims 1 to 5, characterized in that, The step of further training the pre-trained proteome transformation model based on the single-cell proteome data to obtain a further trained proteome decoder includes: inputting the single-cell proteome data into the pre-trained proteome encoder and outputting a single-cell proteome abstract feature representation; inputting the single-cell proteome abstract features into the pre-trained proteome decoder and outputting a single-cell proteome reconstructed data representation; and further training the pre-trained proteome transformation model based on the difference between the input single-cell proteome data and the single-cell proteome reconstructed data to obtain the further trained proteome decoder, wherein the training objective includes minimizing the mean square error between the single-cell proteome data and the single-cell proteome reconstructed data.

8. The method according to any one of claims 1 to 5, characterized in that, The step of training the translation layer of the translation model based on the first pairing relationship, the further trained transcriptome encoder, and the further trained proteome decoder to obtain the trained translation layer includes: inputting the single-cell transcriptome data into the further trained transcriptome encoder and outputting a single-cell transcriptome abstract feature representation; translating the single-cell transcriptome abstract feature representation into a single-cell proteome abstract feature representation through the pre-trained translation layer; inputting the single-cell proteome abstract feature representation into the further trained proteome decoder and outputting single-cell proteome reconstructed data; and training the pre-trained translation layer based on the difference between the single-cell proteome reconstructed data and the reference single-cell proteome data to obtain the trained translation layer, wherein the reference single-cell proteome data is data in the single-cell dataset that conforms to the first pairing relationship with the input single-cell transcriptome data.

9. A translation model training apparatus for implementing the translation model training method as described in any one of claims 1 to 8, characterized in that, The device includes: a multi-tissue cell dataset acquisition module for acquiring a multi-tissue cell dataset, the multi-tissue cell dataset including multi-tissue cell transcriptome data and multi-tissue cell proteome data; and a transcriptome transformation model training module for performing self-supervised training of a transcriptome transformation model based on the multi-tissue cell transcriptome data to obtain a pre-trained transcriptome transformation model, the transcriptome transformation model including a transcriptome encoder and a transcriptome decoder, wherein the transcriptome transformation model is trained based on the differences between the multi-tissue cell transcriptome data and the multi-tissue cell reconstructed transcriptome data, the multi-tissue cell reconstructed transcriptome data being the multi-tissue cell transcriptome... The data is obtained through the output of the transcriptome encoder and the transcriptome decoder. The training objective includes minimizing the mean squared error between the multi-tissue cell transcriptome data and the multi-tissue cell transcriptome reconstructed data. A proteome transformation model training module is used to perform self-supervised training of the proteome transformation model based on the multi-tissue cell proteome data to obtain a pre-trained proteome transformation model. The proteome transformation model includes a proteome encoder and a proteome decoder. The proteome transformation model is trained based on the differences between the multi-tissue cell proteome data and the multi-tissue cell proteome reconstructed data. The multi-tissue cell proteome reconstructed data is the multi-tissue cell proteome... The white group data is obtained through the output of the proteome encoder and the proteome decoder. The training objective includes minimizing the mean squared error between the multi-tissue cell proteome data and the reconstructed multi-tissue cell proteome data. A translation model construction module is used to construct a translation model containing a pre-trained transcriptome encoder and a pre-trained proteome decoder, with a translation layer between the transcriptome encoder and the proteome decoder. A single-tissue cell dataset acquisition module is used to acquire a single-cell dataset, which includes single-cell transcriptome data and single-cell proteome data with a first pairing relationship. A translation model training module is used to train the translation model based on the single-cell transcriptome data. The pre-trained transcriptome conversion model is further trained to obtain a further trained transcriptome encoder; the proteome conversion model is further trained based on the pre-trained single-cell proteome data to obtain a further trained proteome decoder; the translation layer of the translation model is trained based on the first pairing relationship, the further trained transcriptome encoder, and the further trained proteome decoder to obtain a trained translation layer. The training objective includes minimizing the mean square error between the reconstructed single-cell proteome data and the reference single-cell proteome data. The trained translation model is used to translate the single-cell transcriptome data to generate corresponding single-cell whole proteome data.

10. A computer device, characterized in that, The computer device includes a processor and a memory, the memory storing at least one program, which is loaded and executed by the processor to implement the translation model training method as described in any one of claims 1 to 8.

11. A computer-readable storage medium, characterized in that, The storage medium stores at least one program segment, which is loaded and executed by a processor to implement the translation model training method as described in any one of claims 1 to 8.

12. A computer program product, characterized in that, It includes a computer program that, when executed by a processor, implements the translation model training method as described in any one of claims 1 to 8.