Methods, devices, equipment, media and products for processing transcription sequencing data

By using gene identifiers to replace gene expression values, a positive and negative sample training set was constructed and a comparative network model was trained, which solved the problem of cross-platform sequencing bias and achieved efficient processing of transcription and sequencing data and preservation of biological feature information.

CN120708704BActive Publication Date: 2025-10-31SHENZHEN INST OF ADVANCED TECH CHINESE ACAD OF SCI
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511203769.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-08-27
Publication Date
2025-10-31
Estimated Expiration
2045-08-27

AI Technical Summary

Technical Problem

Transcriptional sequencing data from different sequencing platforms exhibit significant biases, and traditional correction methods are ineffective, hindering the widespread application of the data.

Method used

Gene identifiers are used instead of gene expression values ​​for model training. Positive and negative sample training sets are constructed by screening highly expressed genes. The comparative network model is iteratively trained to learn the contextual features between gene identifiers.

Benefits of technology

It effectively eliminates cross-platform sequencing bias while preserving the biological characteristics of transcription sequencing data, thus improving the accuracy and consistency of data processing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120708704B_ABST
    Figure CN120708704B_ABST
Patent Text Reader

Abstract

This invention discloses a method, apparatus, device, medium, and product for processing transcription sequencing data. The method includes: for each sample transcription sequencing data in a transcription sequencing dataset, screening for highly expressed genes in the sample transcription sequencing data to obtain a first gene identifier set; constructing a positive and negative sample training set based on at least one first gene identifier set; and iteratively training an untrained contrastive network model based on the positive and negative sample training set to obtain a trained contrastive network model. Each training sample data in the positive and negative sample training set includes a central gene identifier, a set of positive sample identifiers corresponding to the central gene identifier, and a set of negative sample identifiers. The contrastive network model includes an embedding network and an output layer. The embedding network is used to extract gene vector representations corresponding to the central gene identifier. This invention achieves the goal of eliminating sequencing bias across platforms.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the fields of genomics and bioinformatics, and in particular to a method, apparatus, equipment, medium, and product for processing transcription sequencing data. Background Technology

[0002] In recent years, the rapid development of high-throughput gene sequencing technology has provided a variety of technical options for gene expression analysis. However, different sequencing platforms have significant differences in sequencing principles, read lengths, error rates, and GC content preferences. Sequencing data of the same gene on different sequencing platforms show significant deviations, resulting in poor cross-platform data processing performance and limiting the widespread application of sequencing data.

[0003] To eliminate sequencing bias across platforms, current methods mainly rely on assuming similar data distributions across different sequencing platforms or on prior knowledge for data correction. However, due to the complexity and diversity of sequencing data, traditional correction methods are not ideal for eliminating sequencing bias. Summary of the Invention

[0004] This invention provides a method, apparatus, device, medium, and product for processing transcription sequencing data to address the problem that traditional correction methods are ineffective in eliminating sequencing bias across platforms. While eliminating sequencing bias across platforms, the invention preserves the biological characteristic information contained in the transcription sequencing data.

[0005] According to an embodiment of the present invention, a method for processing transcription sequencing data is provided, the method comprising:

[0006] Obtain a transcription sequencing dataset; wherein the transcription sequencing dataset contains sample transcription sequencing data corresponding to at least one sequencing sample, and the sample transcription sequencing data is obtained by sequencing the sequencing sample using a sequencing platform;

[0007] For each sample's transcriptional sequencing data, highly expressed genes are screened from the sample's transcriptional sequencing data to obtain a first set of gene identifiers;

[0008] Construct a positive and negative sample training set based on at least one set of first gene identifiers;

[0009] Based on the positive and negative sample training set, the untrained contrastive network model is iteratively trained to obtain the trained contrastive network model;

[0010] Each training sample data in the positive and negative sample training sets includes a central gene identifier, a set of positive sample identifiers corresponding to the central gene identifier, and a set of negative sample identifiers. The contrast network model includes an embedding network and an output layer. The embedding network is used to extract the gene vector representation corresponding to the central gene identifier.

[0011] According to another embodiment of the present invention, a transcription sequencing data processing apparatus is provided, the apparatus comprising:

[0012] A transcription sequencing dataset acquisition module is used to acquire a transcription sequencing dataset; wherein, the transcription sequencing dataset contains sample transcription sequencing data corresponding to at least one sequencing sample, and the sample transcription sequencing data is obtained by sequencing the sequencing sample through sequencing by a sequencing platform;

[0013] The first gene identifier set determination module is used to screen for highly expressed genes in the transcriptional sequencing data of each sample to obtain the first gene identifier set.

[0014] The positive and negative sample training set construction module is used to construct positive and negative sample training sets based on at least one first gene identifier set.

[0015] The contrast network model training module is used to iteratively train the untrained contrast network model based on the positive and negative sample training set to obtain a trained contrast network model.

[0016] Each training sample data in the positive and negative sample training sets includes a central gene identifier, a set of positive sample identifiers corresponding to the central gene identifier, and a set of negative sample identifiers. The contrast network model includes an embedding network and an output layer. The embedding network is used to extract the gene vector representation corresponding to the central gene identifier.

[0017] According to another embodiment of the present invention, an electronic device is provided, the electronic device comprising:

[0018] At least one processor; and

[0019] A memory communicatively connected to the at least one processor; wherein,

[0020] The memory stores a computer program that can be executed by the at least one processor, which enables the at least one processor to perform the transcription sequencing data processing method according to any embodiment of the present invention.

[0021] According to another embodiment of the present invention, a computer-readable storage medium is provided, the computer-readable storage medium storing computer instructions for causing a processor to execute and implement the transcription sequencing data processing method described in any embodiment of the present invention.

[0022] According to another embodiment of the present invention, a computer program product is provided, including a computer program that, when executed by a processor, implements the transcription sequencing data processing method described in any embodiment of the present invention.

[0023] The technical solution of this invention solves the problem of poor performance of traditional correction methods in eliminating sequencing bias across platforms by using gene identifiers instead of gene expression values ​​in traditional methods for model training. This ensures that the gene information used for training is not affected by the sequencing modes of different sequencing platforms. At the same time, by screening high-expression genes in the transcriptional sequencing data of the sample transcriptional sequencing dataset, a first gene identifier set is obtained. Based on the positive and negative sample training set constructed from the first gene identifier set, the contrastive network model is trained, enabling the contrastive network model to learn the contextual features between gene identifiers. This achieves the goal of preserving the biological feature information contained in the transcriptional sequencing data in the gene vector representation of the embedded network output.

[0024] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of the present invention, nor is it intended to limit the scope of the invention. Other features of the invention will become readily apparent from the following description. Attached Figure Description

[0025] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0026] Figure 1 A flowchart illustrating a method for processing transcription sequencing data according to an embodiment of the present invention;

[0027] Figure 2 A flowchart illustrating another method for processing transcription sequencing data provided in one embodiment of the present invention;

[0028] Figure 3 A model architecture diagram of a comparative network model provided in one embodiment of the present invention;

[0029] Figure 4 A flowchart illustrating a specific example of a method for processing transcription sequencing data provided in an embodiment of the present invention;

[0030] Figure 5 This is a schematic diagram of a transcription sequencing data processing device provided in one embodiment of the present invention;

[0031] Figure 6 This is a schematic diagram of the structure of an electronic device provided in one embodiment of the present invention. Detailed Implementation

[0032] To enable those skilled in the art to better understand the present invention, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of the present invention.

[0033] It should be noted that the terms "first," "second," "initial," and "to be tested," etc., used in the specification, claims, and accompanying drawings of this invention are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that embodiments of the invention described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.

[0034] Figure 1 This is a flowchart illustrating a method for processing transcriptional sequencing data according to an embodiment of the present invention. This embodiment is applicable to situations requiring de-platformization of transcriptional sequencing data. The method can be executed by a transcriptional sequencing data processing device, which can be implemented in hardware and / or software and can be configured in a terminal device. Figure 1 As shown, the method includes:

[0035] S110. Obtain the transcription sequencing dataset.

[0036] In this embodiment, the transcription sequencing dataset contains sample transcription sequencing data corresponding to at least one sequencing sample, which is obtained by sequencing the sequencing sample through sequencing by the sequencing platform.

[0037] Specifically, a sequencing platform refers to a technical system used to determine the base sequence of RNA (ribonucleic acid) molecules. Its core includes sequencers, chemical reagents, signal detection, and data analysis processes. For example, the sequencing technology used in a sequencing platform can be sequencing-while-synthesizing, single-molecule real-time sequencing, or nanopore electrophysiological sequencing, but it is not limited to the examples mentioned above.

[0038] Specifically, sample transcription sequencing data characterizes the quantitative information obtained by the sequencing platform from RNA sequencing of the sequencing sample. The sample transcription sequencing data contains the gene expression value corresponding to at least one sequenced gene, and the gene expression value can be used to measure gene transcription activity.

[0039] In one optional embodiment, the sequencing sample is a single-cell sample, and correspondingly, the sample transcription sequencing data is single-cell transcription sequencing data.

[0040] In another alternative embodiment, the sequencing sample is a tissue sample or a cell population sample, and correspondingly, the sample transcription sequencing data is batch transcription sequencing data. In one specific embodiment, the sequencing sample is a cancer tissue sample.

[0041] The embodiments of the present invention are particularly applicable to batch transcription sequencing data. The technical deviations exhibited by multiple sequencing platforms on batch transcription sequencing data are relatively stable, mainly due to experimental procedures or sequencing techniques. However, due to cell heterogeneity and low coverage, single-cell transcription sequencing data is prone to confusion between sequencing deviations and biological differences at the cell level during the de-platformization process, thereby reducing the effectiveness of de-platformization of single-cell transcription sequencing data.

[0042] For example, sample transcription sequencing data can be represented as: ,in, This indicates the first in the sample transcription sequencing data. Gene expression values ​​corresponding to each sequenced gene. This indicates the number of sequenced genes in the sample transcription sequencing data.

[0043] S120. For each sample transcription sequencing data, high-expression genes are screened in the sample transcription sequencing data to obtain the first gene identifier set.

[0044] In one optional embodiment, screening for highly expressed genes in the sample transcription sequencing data to obtain a first gene identifier set includes: sorting the sequencing genes in the sample transcription sequencing data according to gene expression values, and determining the first gene identifier set according to a first screening condition and the gene sorting results.

[0045] Specifically, the first screening criteria include the number of first screenings or the proportion of first screenings.

[0046] In another optional embodiment, the sample transcription sequencing data is screened for highly expressed genes to obtain a first gene identifier set, including: determining the first gene identifier set based on the sequencing genes in the sample transcription sequencing data whose gene expression values ​​are greater than the expression value threshold.

[0047] In another optional embodiment, the high-expression genes in the sample transcription sequencing data are screened to obtain a first gene identifier set, including: screening the high-expression genes in the sample transcription sequencing data to obtain an initial gene identifier set; extracting differential features based on multiple initial gene identifier sets to obtain differential weight values ​​corresponding to gene identifiers in the initial gene identifier sets; and screening the gene identifiers in the initial gene identifier sets based on the differential weight values ​​to obtain the first gene identifier set.

[0048] The initial set of gene identifiers can be obtained by sorting and filtering or by filtering by expression value thresholds.

[0049] In one optional embodiment, the differential feature extraction algorithm is either a hypothesis testing algorithm or a TF-IDF algorithm (Term Frequency-Inverse Document Frequency). Specifically, the differential weight value characterizes the degree of difference in gene expression values ​​of the sequencing gene corresponding to the gene identifier across multiple sample transcription sequencing data.

[0050] In one optional embodiment, the gene identifiers in the initial gene identifier set are filtered according to the difference weight value to obtain a first gene identifier set, including: deleting gene identifiers in the initial gene identifier set whose difference weight value is less than the weight value threshold to obtain the first gene identifier set; or, sorting the gene identifiers in the initial gene identifier set according to the difference weight value, and filtering the identifier sorting results according to the second screening condition to obtain the first gene identifier set.

[0051] Specifically, the second screening criteria include the number of second screenings or the proportion of second screenings.

[0052] The advantage of this setup is that it highlights the differences between the first set of gene markers, avoids interference from sequencing genes that are highly expressed but do not show sample differences (such as housekeeping genes), and thus improves the learning performance of subsequent comparative network models.

[0053] In this embodiment, the first gene identifier set includes gene identifiers corresponding to multiple high-expression sequencing genes in the sample transcription sequencing data. These gene identifiers are used to uniquely identify the sequencing genes. For example, a gene identifier can be the sequence number of the sequencing gene in a specified sequence, an identifier in a specified database, or a coordinate position on a chromosome, but is not limited to the examples described above.

[0054] For example, the first set of gene identifiers can be represented as: ,in, This indicates the number of sequenced genes that are highly expressed.

[0055] S130. Construct a positive and negative sample training set based on at least one first gene identifier set.

[0056] In this embodiment, each training sample data in the positive and negative sample training sets includes a central gene identifier, a set of positive sample identifiers corresponding to the central gene identifier, and a set of negative sample identifiers.

[0057] In an optional embodiment, constructing a positive and negative sample training set based on at least one first gene identifier set includes: for each gene identifier in each first gene identifier set, using the gene identifier as a central gene identifier, constructing a positive sample identifier set based on gene identifiers in the first gene identifier set that interact with the central gene identifier, constructing a negative sample identifier set based on gene identifiers in the first gene identifier set other than the central gene identifier and the positive sample identifier set, and determining training sample data based on the central gene identifier, the positive sample identifier set, and the negative sample identifier set; and constructing a positive and negative sample training set based on multiple training sample data sets.

[0058] For example, the interaction between gene identifiers in the first gene identifier set and the central gene identifier can be obtained by at least one of the following methods: obtaining from a public database, determining based on the expression correlation between gene identifiers and the central gene identifier, determining by random forest sorting, or predicting by a neural network model, but is not limited to the example scenario.

[0059] Specifically, the number of training sample data corresponding to the first gene identifier set is equal to the number of gene identifiers corresponding to the first gene identifier set.

[0060] In this embodiment, the positive sample identifier set includes gene identifiers that interact with the central gene identifier, and the negative sample identifier set includes gene identifiers that do not interact with the central gene identifier or whose interaction is unknown.

[0061] In another optional embodiment, constructing a positive and negative sample training set based on at least one first gene identifier set includes: determining a second gene identifier set based on a transcription sequencing dataset; for each gene identifier in each first gene identifier set, using the gene identifier as a central gene identifier, and constructing a positive sample identifier set based on gene identifiers in the first gene identifier set other than the central gene identifier; constructing a negative sample identifier set based on at least one gene identifier in the second gene identifier set other than the first gene identifier set; and determining training sample data based on the central gene identifier, the positive sample identifier set, and the negative sample identifier set; and constructing a positive and negative sample training set based on multiple training sample data sets.

[0062] In this embodiment, the second gene identifier set summarizes the gene identifiers of sequenced genes from transcriptional sequencing data of multiple samples.

[0063] In this embodiment, the positive sample identifier set includes gene identifiers that are highly co-expressed with the central gene identifier, indicating a positive interaction with the central gene identifier. The negative sample identifier set includes gene identifiers that are not co-expressed with the central gene identifier, indicating a negative interaction with the central gene identifier. For example, the number of gene identifiers corresponding to the negative sample identifier set is 3 or 4 times the number of gene identifiers corresponding to the positive sample identifier set.

[0064] The advantage of this setup is that by treating the first set of gene identifiers as a "sentence" and the gene identifiers in the first set of gene identifiers as "words," the contrastive network model learns the vector representation of each "word" in the "sentence," thereby further ensuring the preservation of biometric information.

[0065] Based on the above embodiments, optionally, determining a second gene identifier set according to the transcription sequencing dataset includes: determining a first sequencing matrix according to the transcription sequencing dataset; filtering the sequencing genes corresponding to the first sequencing matrix to obtain a second sequencing matrix; and constructing a second gene identifier set according to the matrix encoding corresponding to the sequencing genes in the second sequencing matrix.

[0066] In this embodiment, the matrix parameter of the first sequencing matrix is ​​the expression value of the sequenced gene in the sample transcription sequencing data. For example, the rows of the first sequencing matrix represent the sample transcription sequencing data and the columns represent the sequenced genes, or the rows of the first sequencing matrix represent the sequenced genes and the columns represent the sample transcription sequencing data.

[0067] For example, the first sequencing matrix is ​​represented as follows: ,in, This indicates the number of sequencing samples corresponding to the transcription sequencing dataset. This indicates the number of sequenced genes corresponding to the transcription sequencing dataset.

[0068] For example, the filtering content includes, but is not limited to, filtering out rare sequencing genes, filtering out low-quality sequencing genes, and filtering out contaminated sequencing genes, etc.

[0069] Specifically, the matrix encoding corresponding to the sequenced genes in the second sequencing matrix is ​​either row encoding or column encoding, representing the row or column sequence number of the sequenced gene in the second sequencing matrix.

[0070] Specifically, the second set of gene identifiers is obtained by summarizing and filtering multiple transcriptional sequencing data from the transcriptional sequencing dataset.

[0071] For example, the second set of gene identifiers can be represented as: ,in, This indicates the number of sequencing genes corresponding to the second set of gene identifiers. .

[0072] S140. Based on the positive and negative sample training set, iteratively train the untrained contrastive network model to obtain the trained contrastive network model.

[0073] Specifically, the contrastive network model is used to determine the positive or negative interaction between gene markers and central gene markers.

[0074] In this embodiment, the comparison network model includes an embedding network and an output layer. The embedding network extracts the gene vector representation corresponding to the central gene identifier in the input training sample data. The output layer outputs the interaction labels between the central gene identifier and gene identifiers in the positive sample identifier set, as well as the interaction labels between the central gene identifier and gene identifiers in the negative sample identifier set.

[0075] For example, interaction labels can be binary or multi-classified, such as "1" representing a positive interaction and "0" representing a negative interaction, or interaction labels representing the probability that the two have a positive interaction or a negative interaction.

[0076] Specifically, the embedded network in the trained contrastive network model is used to extract gene vector representations of gene identifiers corresponding to the sequenced genes in the transcriptional sequencing data from any sequencing platform.

[0077] The technical solution of this embodiment solves the problem of poor performance of traditional correction methods in eliminating sequencing bias across platforms by using gene identifiers instead of gene expression values ​​in traditional methods for model training. This ensures that the gene information used for training is not affected by the sequencing modes of different sequencing platforms. At the same time, by screening high-expression genes in the transcriptional sequencing data of the sample transcriptional sequencing dataset, a first gene identifier set is obtained. Based on the positive and negative sample training set constructed from the first gene identifier set, the contrastive network model is trained, enabling the contrastive network model to learn the contextual features between gene identifiers. This achieves the goal of preserving the biological feature information contained in the transcriptional sequencing data in the gene vector representation of the embedded network output.

[0078] Figure 2This is a flowchart illustrating another method for processing transcription sequencing data according to an embodiment of the present invention. This embodiment further refines the "contrast network model" in the above embodiment. In this embodiment, based on positive and negative sample training sets, the untrained contrast network model is iteratively trained to obtain a trained contrast network model. This includes: for each training sample data, inputting the central gene identifiers in the training sample data into the central embedding layer to obtain the output central embedding vector; inputting the positive sample identifier set into the positive sample embedding layer to obtain the output first embedding vector set; inputting the negative sample identifier set into the negative sample embedding layer to obtain the output second embedding vector set; inputting the central embedding vector, the first embedding vector set, and the second embedding vector set into the output layer to obtain the output interaction label set; determining the loss function value based on the interaction label set corresponding to each training sample data; iteratively training the model parameters of the contrast network model based on the loss function value until the loss function value converges to obtain the trained contrast network model. Figure 2 As shown, the method includes:

[0079] S210. Obtain the transcription sequencing dataset.

[0080] S220. For each sample transcription sequencing data, high-expression genes are screened in the sample transcription sequencing data to obtain the first gene identifier set.

[0081] S230. Construct a positive and negative sample training set based on at least one first gene identifier set.

[0082] S210-S230 in this embodiment are the same as those in the above embodiment. Figure 1 The S110-S130 shown are the same or similar, and will not be described again in this embodiment.

[0083] S240. For each training sample data, the central gene identifier in the training sample data is input into the central embedding layer to obtain the output central embedding vector. The positive sample identifier set is input into the positive sample embedding layer to obtain the output first embedding vector set. The negative sample identifier set is input into the negative sample embedding layer to obtain the output second embedding vector set. The central embedding vector, the first embedding vector set and the second embedding vector set are input into the output layer to obtain the output interaction label set.

[0084] Figure 3 This is a model architecture diagram of a contrastive network model provided in one embodiment of the present invention. Specifically, the contrastive network model includes an embedding network and an output layer, wherein the embedding network includes a central embedding layer, a positive sample embedding layer, and a negative sample embedding layer.

[0085] Specifically, the central embedding layer is a learnable weight matrix. For example, the weight matrix... , where 8 represents the dimension of the center embedding vector.

[0086] Specifically, the positive sample embedding layer and the negative sample embedding layer share the model architecture and model parameters. The first embedding vector set contains the embedding vector corresponding to each gene identifier in the positive sample identifier set, and the second embedding vector set contains the embedding vector corresponding to each gene identifier in the negative sample identifier set.

[0087] In an optional embodiment, the output layer includes a fusion module and an activation module; wherein the fusion module is used to fuse the center embedding vector with a first embedding vector set to obtain a first fused vector set, and to fuse the center embedding vector with a second embedding vector set to obtain a second fused vector set; the activation module is used to output an interaction tag set based on the first fused vector set and the second fused vector set.

[0088] For example, the fusion process can be splicing, dot product, addition, or averaging, but it is not limited to the above examples.

[0089] Specifically, the interaction tag set includes the interaction tags corresponding to each gene identifier and the central gene identifier in the positive sample identifier set, as well as the interaction tags corresponding to each gene identifier and the central gene identifier in the negative sample identifier set.

[0090] S250. Determine the loss function value based on the interaction label set corresponding to each training sample data.

[0091] For example, the loss function value Satisfy the following formula:

[0092] ;

[0093] in, Represents the first fusion vector set, Represents the second fusion vector set, This represents the activation function.

[0094] S260. Iteratively train the model parameters of the comparison network model based on the loss function value until the loss function value converges, and obtain the trained comparison network model.

[0095] Based on the above embodiments, the method may optionally further include: using a hyperparameter optimization algorithm to train the comparison network models corresponding to multiple sets of hyperparameters; performing performance tests on each trained comparison network model, and taking the comparison network model with the best test performance as the final comparison network model.

[0096] For example, the hyperparameter optimization algorithm can be a grid search algorithm, and the hyperparameters include, but are not limited to, the number of iterations, the dimension of the embedding vector, and the learning rate.

[0097] The technical solution of this embodiment sets the embedding network in the contrast network model to include a central embedding layer, a positive sample embedding layer, and a negative sample embedding layer. Training sample data is input into the embedding network, and the central embedding vector, the first embedding vector set, and the second embedding vector set output by the embedding network are input into the output layer to obtain the output interaction label set. Based on the interaction label set, the loss function value is determined, and the model parameters of the contrast network model are iteratively trained based on the loss function value until the loss function value converges. This achieves decoupling between the central gene identifier, the gene identifier as a positive sample, and the gene identifier as a negative sample. Therefore, the backpropagation gradients of each embedding layer are independent of each other, avoiding confusion or interference between embedding vectors, improving the stability of the contrast network model, and effectively suppressing the batch effect.

[0098] Based on the above embodiments, the method may optionally further include: determining a third gene identifier set corresponding to the sequencing genes in the transcriptional sequencing data to be tested from any sequencing platform; inputting the third gene identifier set into the embedding network in the trained contrastive network model to obtain the gene vector representation corresponding to each sequencing gene; and determining the sequencing vector representation corresponding to the transcriptional sequencing data to be tested based on the multiple gene vector representations.

[0099] Specifically, the gene identifiers corresponding to each sequenced gene in the transcriptional sequencing data to be tested are obtained from the second gene identifier set to obtain the third gene identifier set.

[0100] Based on the above embodiments, optionally, the method further includes: obtaining sequencing vector representations corresponding to multiple test transcription sequencing data respectively; inputting the multiple sequencing vector representations into a classification and discrimination model to obtain the predicted classification label corresponding to each test transcription sequencing data.

[0101] For example, the predicted classification label can be a cancer category, such as lung cancer or liver cancer. The predicted classification label can also be a tumor subtype, such as histological subtype, molecular subtype, driver gene subtype, MSI status (Microsatellite Instability), etc. The predicted classification label can also be a survival category, such as survival probability, but it is not limited to the above examples.

[0102] Based on the above embodiments, optionally, the method further includes: obtaining the sequencing vector set corresponding to each predicted classification label; for each sequencing vector representation in the sequencing vector set, obtaining the same-cluster distance between the sequencing vector representation and the remaining sequencing vector representations in the sequencing vector set, and obtaining the different-cluster distance between the sequencing vector representation and the sequencing vector representations in other sequencing vector sets; determining the silhouette coefficient corresponding to the sequencing vector representation based on the multiple same-cluster distances and the multiple different-cluster distances; and determining the average silhouette coefficient based on the silhouette coefficients corresponding to the multiple sequencing vector representations.

[0103] The mean silhouette coefficient is an indicator used to measure the quality of clustering. The larger the mean silhouette coefficient, the more biological feature information is retained in the gene vector representation. Conversely, the smaller the mean silhouette coefficient, the less biological feature information is retained in the gene vector representation.

[0104] Based on the above embodiments, the method may optionally further include: obtaining a sequencing vector set corresponding to any predicted classification label, and determining a first transcriptional sequencing dataset corresponding to the sequencing vector set; obtaining a second transcriptional sequencing dataset corresponding to the predicted classification label; and determining neighborhood consistency based on the first transcriptional sequencing dataset and the second transcriptional sequencing dataset.

[0105] In this embodiment, for example, the second transcription sequencing dataset is obtained by clustering multiple transcription sequencing datasets to be tested or by classifying and predicting multiple transcription sequencing datasets to be tested using a neural network model.

[0106] Specifically, neighborhood consistency is used to measure the fidelity of local structures. The greater the neighborhood consistency, the better the effect of removing plateau bias in gene vector representation. Conversely, the smaller the neighborhood consistency, the worse the effect of removing plateau bias in gene vector representation.

[0107] Based on the above embodiments, optionally, the method further includes: determining a fourth gene identifier set corresponding to the test transcription sequencing data from a public database, inputting the fourth gene identifier set into the embedding network in the trained contrastive network model to obtain multiple gene vector representations, and determining the sequencing vector representation corresponding to the test transcription sequencing data based on the multiple gene vector representations.

[0108] Specifically, test transcription sequencing data from public databases were used as an independent test set to validate the generalization performance of the comparative network model.

[0109] Figure 4This is a flowchart illustrating a specific example of a method for processing transcriptional sequencing data according to an embodiment of the present invention. Specifically, in the data preparation stage, transcriptional sequencing data of sequencing samples are collected from multiple sequencing platforms, and data alignment, format standardization, and gene filtering are performed on the multiple sample transcriptional sequencing data to obtain a transcriptional sequencing dataset. In the model training stage, highly expressed genes are screened for in each sample transcriptional test data in the transcriptional sequencing dataset, and gene identifiers are assigned to the screened sequencing genes to obtain a first gene identifier set. Based on multiple first gene identifier sets, positive and negative sample training sets are constructed. Using the positive and negative sample training sets, the contrastive network model is iteratively trained to obtain the trained contrastive network model.

[0110] In the model application phase, based on the second gene identifier set, a third gene identifier set is determined for the transcribed sequencing data to be tested on any sequencing platform. This third gene identifier set is then input into the embedding network of the trained contrastive network model to obtain the gene vector representation output by the intermediate embedding layer. Multiple gene vector representations constitute the sequencing vector representation corresponding to the transcribed sequencing data to be tested. Downstream tasks, such as cancer classification, tumor subtype discrimination, and survival classification, are then performed on the sequencing vector representations corresponding to the multiple transcribed sequencing data to be tested. In the method validation phase, generalization validation is performed using test transcribed sequencing data from a public database. Based on the task data corresponding to the aforementioned downstream tasks, metrics such as average contour width and neighborhood consistency are calculated, and the results are visualized.

[0111] The following are embodiments of the transcription sequencing data processing device provided in this invention. This device and the transcription sequencing data processing method in the above embodiments belong to the same inventive concept. For details not described in detail in the embodiments of the transcription sequencing data processing device, please refer to the content of the transcription sequencing data processing method in the above embodiments.

[0112] Figure 5 This is a schematic diagram of a transcription sequencing data processing device provided in one embodiment of the present invention. Figure 5 As shown, the device includes: a transcription sequencing dataset acquisition module 310, a first gene identifier set determination module 320, a positive and negative sample training set construction module 330, and a contrastive network model training module 340.

[0113] The system includes a transcription sequencing dataset acquisition module 310, which acquires a transcription sequencing dataset containing transcription sequencing data corresponding to at least one sequencing sample. The transcription sequencing data is obtained by sequencing the samples using a sequencing platform. A first gene identifier set determination module 320 screens for highly expressed genes in each sample's transcription sequencing data to obtain a first gene identifier set. A positive and negative sample training set construction module 330 constructs a positive and negative sample training set based on at least one first gene identifier set. A contrastive network model training module 340 iteratively trains an untrained contrastive network model using the positive and negative sample training sets to obtain a trained contrastive network model.

[0114] Each training sample in the positive and negative sample training sets includes a central gene identifier, a set of positive sample identifiers corresponding to the central gene identifier, and a set of negative sample identifiers. The contrast network model includes an embedding network and an output layer. The embedding network is used to extract the gene vector representation corresponding to the central gene identifier.

[0115] The technical solution of this embodiment solves the problem of poor performance of traditional correction methods in eliminating sequencing bias across platforms by using gene identifiers instead of gene expression values ​​in traditional methods for model training. This ensures that the gene information used for training is not affected by the sequencing modes of different sequencing platforms. At the same time, by screening high-expression genes in the transcriptional sequencing data of the sample transcriptional sequencing dataset, a first gene identifier set is obtained. Based on the positive and negative sample training set constructed from the first gene identifier set, the contrastive network model is trained, enabling the contrastive network model to learn the contextual features between gene identifiers. This achieves the goal of preserving the biological feature information contained in the transcriptional sequencing data in the gene vector representation of the embedded network output.

[0116] In an optional embodiment, the positive and negative sample training set construction module 330 includes:

[0117] The second gene identifier set determination unit is used to determine the second gene identifier set based on the transcription sequencing dataset; wherein the second gene identifier set consists of gene identifiers from sequencing genes corresponding to transcription sequencing data from multiple samples;

[0118] The training sample data determination unit is used to, for each gene identifier in each first gene identifier set, take the gene identifier as the central gene identifier, construct a positive sample identifier set based on the gene identifiers in the first gene identifier set other than the central gene identifier, construct a negative sample identifier set based on at least one gene identifier in the second gene identifier set other than the first gene identifier set, and determine the training sample data based on the central gene identifier, the positive sample identifier set, and the negative sample identifier set.

[0119] The positive and negative sample training set construction unit is used to construct a positive and negative sample training set based on multiple training sample data.

[0120] In one optional embodiment, the second gene identifier set determination unit is specifically used for:

[0121] Based on the transcription sequencing dataset, a first sequencing matrix is ​​determined; wherein, the matrix parameter of the first sequencing matrix is ​​the expression value of the sequenced gene in the sample transcription sequencing data;

[0122] The sequencing genes corresponding to the first sequencing matrix are filtered to obtain the second sequencing matrix, and a second gene identifier set is constructed based on the matrix codes corresponding to the sequencing genes in the second sequencing matrix.

[0123] In an optional embodiment, the embedding network includes a central embedding layer, a positive sample embedding layer, and a negative sample embedding layer. Correspondingly, the contrastive network model training module 340 is specifically used for:

[0124] For each training sample data, the central gene identifier in the training sample data is input into the central embedding layer to obtain the output central embedding vector. The positive sample identifier set is input into the positive sample embedding layer to obtain the output first embedding vector set. The negative sample identifier set is input into the negative sample embedding layer to obtain the output second embedding vector set. The central embedding vector, the first embedding vector set and the second embedding vector set are input into the output layer to obtain the output interaction label set.

[0125] The loss function value is determined based on the interaction label set corresponding to each training sample data.

[0126] The model parameters of the comparison network model are iteratively trained based on the loss function value until the loss function value converges, resulting in a trained comparison network model.

[0127] In one optional embodiment, the output layer includes a fusion module and an activation module;

[0128] The fusion module is used to fuse the center embedding vector with the first embedding vector set to obtain the first fused vector set, and to fuse the center embedding vector with the second embedding vector set to obtain the second fused vector set.

[0129] The activation module is used to output an interaction tag set based on the first fusion vector set and the second fusion vector set.

[0130] In an optional embodiment, the first gene identifier set determination module 320 is specifically used for:

[0131] An initial set of gene identifiers was obtained by screening highly expressed genes from the sample transcription and sequencing data;

[0132] Differential feature extraction is performed based on multiple initial gene identifier sets to obtain the differential weight values ​​corresponding to the gene identifiers in the initial gene identifier sets;

[0133] Based on the difference weight values, the gene identifiers in the initial gene identifier set are filtered to obtain the first gene identifier set.

[0134] In an optional embodiment, the device further includes:

[0135] The sequencing vector representation determination module is used to determine the set of third gene identifiers corresponding to the sequencing genes in the transcriptome sequencing data from any sequencing platform.

[0136] The third set of gene identifiers is input into the embedding network of the trained contrastive network model to obtain the gene vector representation corresponding to each sequenced gene.

[0137] Based on multiple gene vector representations, the sequencing vector representation corresponding to the transcriptional sequencing data to be tested is determined.

[0138] The transcription sequencing data processing apparatus provided in the embodiments of the present invention can execute the transcription sequencing data processing method provided in any embodiment of the present invention, and has the corresponding functional modules and beneficial effects of executing the method.

[0139] Figure 6 This is a schematic diagram of an electronic device provided according to one embodiment of the present invention. The electronic device 10 is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workbenches, servers, blade servers, mainframe computers, and other suitable computers. The electronic device may also represent various forms of mobile devices, such as personal digital assistants, cellular phones, smartphones, wearable devices (such as helmets, glasses, watches, etc.), and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the invention described and / or claimed herein.

[0140] like Figure 6As shown, the electronic device 10 includes at least one processor 11 and a memory, such as a read-only memory (ROM) 12 or a random access memory (RAM) 13, communicatively connected to the at least one processor 11. The memory stores computer programs executable by the at least one processor 11. The processor 11 can perform various appropriate actions and processes based on the computer program stored in the ROM 12 or loaded from storage unit 18 into the RAM 13. The RAM 13 can also store various programs and data required for the operation of the electronic device 10. The processor 11, ROM 12, and RAM 13 are interconnected via a bus 14. An input / output (I / O) interface 15 is also connected to the bus 14.

[0141] Multiple components in electronic device 10 are connected to I / O interface 15, including: input unit 16, such as keyboard, mouse, etc.; output unit 17, such as various types of displays, speakers, etc.; storage unit 18, such as disk, optical disk, etc.; and communication unit 19, such as network card, modem, wireless transceiver, etc. Communication unit 19 allows electronic device 10 to exchange information or data with other devices through computer networks such as the Internet and / or various telecommunications networks.

[0142] Processor 11 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of processor 11 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, digital signal processors (DSPs), and any suitable processor, controller, microcontroller, etc. Processor 11 performs the various methods and processes described above, such as the transcription sequencing data processing method provided in the above embodiments.

[0143] In some embodiments, the transcription sequencing data processing method provided in the above embodiments can be implemented as a computer program tangibly contained in a computer-readable storage medium, such as storage unit 18. In some embodiments, part or all of the computer program can be loaded and / or installed on electronic device 10 via ROM 12 and / or communication unit 19. When the computer program is loaded into RAM 13 and executed by processor 11, one or more steps of the transcription sequencing data processing method described above can be performed. Alternatively, in other embodiments, processor 11 can be configured to perform the transcription sequencing data processing method by any other suitable means (e.g., by means of firmware).

[0144] In particular, according to embodiments of the present invention, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments of the present invention include a computer program product comprising a computer program carried on a non-transitory computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via communication unit 19, or installed from storage unit 18, or installed from ROM 12. When the computer program is executed by processor 11, it performs the functions defined in the methods of the embodiments of the present invention.

[0145] Various embodiments of the systems and techniques described above herein can be implemented in the following systems or combinations thereof: digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard parts (ASSPs), system-on-chips (SOCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.

[0146] Computer programs for implementing the transcription sequencing data processing methods of the present invention can be written in any combination of one or more programming languages. These computer programs can be provided to the processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when executed by the processor, the computer programs cause the functions / operations specified in the flowcharts and / or block diagrams to be implemented. The computer programs can be executed entirely on the machine, partially on the machine, as a standalone software package partially on the machine and partially on a remote machine, or entirely on a remote machine or server.

[0147] In the context of this application, a computer-readable storage medium can be a tangible medium that may contain or store a computer program for use by or in conjunction with an instruction execution system, apparatus, or device. A computer-readable storage medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. Alternatively, a computer-readable storage medium can be a machine-readable storage medium. Examples of machine-readable storage media include, based on an electrical connection of at least one wire, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM), flash memory, optical fiber, compact disc read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.

[0148] To provide interaction with a user, the systems and techniques described herein can be implemented on a terminal device having: a display device for displaying information to the user (e.g., a cathode-ray tube (CRT) or liquid crystal display (LCD) monitor); and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the terminal device. Other types of devices can also provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).

[0149] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as data servers), or middleware components (e.g., application servers), or frontend components (e.g., user computers with graphical user interfaces or web browsers through which users can interact with implementations of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., communication networks). Examples of communication networks include local area networks (LANs), wide area networks (WANs), blockchain networks, and the Internet.

[0150] A computing system can include clients and servers. Clients and servers are generally located far apart and typically interact through communication networks. The client-server relationship is created by computer programs running on the respective computers and having a client-server relationship with each other. The server can be a cloud server, also known as a cloud computing server or cloud host, which is a hosting product within the cloud computing service system. It addresses the shortcomings of traditional physical hosts and Virtual Private Server (VPS) services, such as high management difficulty and weak business scalability.

[0151] It should be understood that the various forms of processes shown above can be used, with steps reordered, added, or deleted. For example, the steps described in this invention can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution of this invention can be achieved, and this is not limited herein.

[0152] The specific embodiments described above do not constitute a limitation on the scope of protection of this invention. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this invention should be included within the scope of protection of this invention.

Claims

1. A method for processing transcription sequencing data, characterized in that, include: Obtain a transcription sequencing dataset; wherein the transcription sequencing dataset contains sample transcription sequencing data corresponding to at least one sequencing sample, and the sample transcription sequencing data is obtained by sequencing the sequencing sample using a sequencing platform; For each sample's transcriptional sequencing data, highly expressed genes are screened in the sample's transcriptional sequencing data to obtain a first set of gene identifiers; Construct a positive and negative sample training set based on at least one set of first gene identifiers; Based on the positive and negative sample training set, the untrained contrastive network model is iteratively trained to obtain the trained contrastive network model. Each training sample data in the positive and negative sample training sets includes a central gene identifier, a set of positive sample identifiers corresponding to the central gene identifier, and a set of negative sample identifiers. The contrast network model includes an embedding network and an output layer. The embedding network is used to extract the gene vector representation corresponding to the central gene identifier.

2. The method for processing transcription sequencing data according to claim 1, characterized in that, The construction of a positive and negative sample training set based on at least one first gene identifier set includes: Based on the transcription sequencing dataset, a second gene identifier set is determined; wherein the second gene identifier set summarizes the gene identifiers of sequencing genes from transcription sequencing data of multiple samples; For each gene identifier in each first gene identifier set, the gene identifier is used as the central gene identifier, and a positive sample identifier set is constructed based on the gene identifiers in the first gene identifier set other than the central gene identifier. A negative sample identifier set is constructed based on at least one gene identifier in the second gene identifier set other than the first gene identifier set. Training sample data is determined based on the central gene identifier, the positive sample identifier set, and the negative sample identifier set. A training set of positive and negative samples is constructed based on multiple training sample data.

3. The method for processing transcription sequencing data according to claim 2, characterized in that, The step of determining the second gene identifier set based on the transcription sequencing dataset includes: Based on the transcription sequencing dataset, a first sequencing matrix is ​​determined; wherein the matrix parameter of the first sequencing matrix is ​​the expression value of the sequenced gene in the sample transcription sequencing data; The sequencing genes corresponding to the first sequencing matrix are filtered to obtain the second sequencing matrix, and a second gene identifier set is constructed based on the matrix encoding of the sequencing genes in the second sequencing matrix.

4. The method for processing transcription sequencing data according to claim 1, characterized in that, The embedding network includes a central embedding layer, a positive sample embedding layer, and a negative sample embedding layer. Correspondingly, the step of iteratively training the untrained contrastive network model to obtain a trained contrastive network model based on the positive and negative sample training set includes: For each training sample data, the central gene identifier in the training sample data is input into the central embedding layer to obtain the output central embedding vector, and the positive sample identifier set is input into the positive sample embedding layer to obtain the output first embedding vector set, and the negative sample identifier set is input into the negative sample embedding layer to obtain the output second embedding vector set. The central embedding vector, the first embedding vector set and the second embedding vector set are input into the output layer to obtain the output interaction tag set. The loss function value is determined based on the interaction label set corresponding to each training sample data. The model parameters of the comparison network model are iteratively trained based on the loss function value until the loss function value converges, thus obtaining the trained comparison network model.

5. The method for processing transcription sequencing data according to claim 4, characterized in that, The output layer includes a fusion module and an activation module; The fusion module is used to fuse the center embedding vector with the first embedding vector set to obtain a first fused vector set, and to fuse the center embedding vector with the second embedding vector set to obtain a second fused vector set. The activation module is used to output an interaction tag set based on the first fusion vector set and the second fusion vector set.

6. The method for processing transcription sequencing data according to claim 1, characterized in that, The process of screening highly expressed genes from the transcriptional sequencing data of the sample yields a first set of gene identifiers, including: The transcriptional sequencing data of the samples were used to screen for highly expressed genes to obtain an initial set of gene identifiers; Differential feature extraction is performed based on multiple initial gene identifier sets to obtain the differential weight values ​​corresponding to the gene identifiers in the initial gene identifier sets; Based on the difference weight values, the gene identifiers in the initial gene identifier set are filtered to obtain the first gene identifier set.

7. The method for processing transcription sequencing data according to any one of claims 1-6, characterized in that, Also includes: Determine the set of third gene identifiers corresponding to the sequenced genes in the transcriptome sequencing data from any sequencing platform; The third set of gene identifiers is input into the embedding network of the trained contrastive network model to obtain the gene vector representation corresponding to each sequenced gene. Based on multiple gene vector representations, the sequencing vector representation corresponding to the transcriptional sequencing data to be tested is determined.

8. A device for processing transcription sequencing data, characterized in that, include: A transcription sequencing dataset acquisition module is used to acquire a transcription sequencing dataset; wherein, the transcription sequencing dataset contains sample transcription sequencing data corresponding to at least one sequencing sample, and the sample transcription sequencing data is obtained by sequencing the sequencing sample through sequencing by a sequencing platform; The first gene identifier set determination module is used to screen for highly expressed genes in the transcriptional sequencing data of each sample to obtain the first gene identifier set. The positive and negative sample training set construction module is used to construct positive and negative sample training sets based on at least one first gene identifier set. The contrast network model training module is used to iteratively train the untrained contrast network model based on the positive and negative sample training set to obtain a trained contrast network model. Each training sample data in the positive and negative sample training sets includes a central gene identifier, a set of positive sample identifiers corresponding to the central gene identifier, and a set of negative sample identifiers. The contrast network model includes an embedding network and an output layer. The embedding network is used to extract the gene vector representation corresponding to the central gene identifier.

9. An electronic device, characterized in that, The electronic device includes: At least one processor; and A memory communicatively connected to the at least one processor; wherein, The memory stores a computer program that can be executed by the at least one processor to enable the at least one processor to perform the method for processing transcription sequencing data according to any one of claims 1-7.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions that cause a processor to execute the method for processing transcription sequencing data according to any one of claims 1-7.

11. A computer program product, characterized in that, It includes a computer program that, when executed by a processor, implements a method for processing transcriptional sequencing data according to any one of claims 1-7.

Citation Information

Patent Citations

  • Computer-implemented method and computer system for rank normalization for differential expression analysis of transcriptome sequencing data

    CN103377317A

  • Ribosome imprinting sequencing data analysis method and system

    CN111243665A