Gene expression deletion filling method and system based on deep learning
By dividing spatial transcriptome data into tile and spot, combined with KNN algorithm and comparison learning module, the problems of insufficient spatial accuracy of the filling results of gene expression missing value and insufficient local region interaction capture in the prior art are solved, and a more efficient and accurate filling effect is achieved.
Patent Information
- Application Number
- CN202510672128.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-23
- Publication Date
- 2025-06-20
- Estimated Expiration
- 2045-05-23
AI Technical Summary
In the prior art, when processing spatial transcriptome data, there are problems such as insufficient spatial accuracy of gene expression missing value filling results, insufficient capture of local regional interactions, and imbalance in model complexity and computational efficiency.
By dividing the detection points into multiple tiles and spots, counting genes whose expression frequency in the spot has higher than the set threshold, using the KNN algorithm for preliminary interpolation, and calculating the population feature embedding of positive and negative samples through the tile feature embedding module and the spot feature embedding module. The model is trained using the comparative learning module, and finally selecting tiles with similar embedding representations for reinterpolation.
It improves the local consistency and spatial continuity of the filling results, balances the model complexity and computing efficiency, and is suitable for filling large-scale spatial transcriptome data.
Smart Images

Figure CN120183513A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of electronic digital data processing, and particularly relates to a method and system for filling gene expression deletions based on deep learning. Background Art
[0002] Spatial transcriptomics technology, with its unique advantage of measuring gene expression levels while preserving tissue spatial location information, has become a key tool for exploring tissue microenvironment and cellular heterogeneity in biomedical research. However, due to technical limitations and experimental conditions, gene expression deletions are prevalent in spatial transcriptomics data, which not only destroys data integrity but also directly affects the accuracy of subsequent analysis.
[0003] Currently, methods for filling missing values in spatial transcriptomics gene expression mainly fall into three categories: interpolation methods based on statistical models, prediction methods based on machine learning, and generative models based on deep learning. Interpolation methods based on statistical models, such as mean interpolation and K-nearest neighbor interpolation, are easy to implement but have poor filling result accuracy due to ignoring the complex relationships of gene expression and spatial location information; prediction methods based on machine learning, using models such as random forest and support vector machine, improve filling accuracy to a certain extent but fail to fully exploit the spatial information and local region interactions of spatial transcriptomics data; generative models based on deep learning can capture non-linear relationships of gene expression but have insufficient consideration of local region spatial dependence when dealing with high-resolution data, resulting in poor spatial continuity of filling results.
[0004] Existing methods have three significant defects: First, the utilization of spatial information is insufficient, leading to insufficient spatial accuracy of filling results; second, the local region interactions cannot be effectively captured, making the local filling results lack consistency; third, although some deep learning methods improve filling accuracy, the models are complex and the computational efficiency is low, making it difficult to adapt to large-scale datasets. Summary of the Invention
[0005] To solve the problems in the background art, one aspect of the present invention provides a method for filling gene expression deletions based on deep learning, including:
[0006] S1: Preprocess the original data of detection points on tissue sections to obtain high-resolution spatial transcriptomics data, where the high-resolution spatial transcriptomics data includes gene expression vectors of detection points and position information of detection points;
[0007] S2: Divide the detection points into multiple tiles and spots, where each spot contains multiple tiles and each tile contains multiple detection points;
[0008] S3: Statistically analyze genes with expression frequencies higher than a set threshold within the spot as expressed genes to construct an expressed gene set; perform preliminary imputation on tiles within the spot that lack the expression of the expressed genes through the KNN algorithm;
[0009] S4: Select tiles from the spot to construct positive and negative samples, and calculate the population feature embeddings of the positive and negative samples through the tile feature embedding module and the spot feature embedding module; wherein, the tile feature embedding module is used to obtain the embedding representation of the tile; the spot feature embedding module is used to obtain the population feature embedding representation of the positive and negative samples according to the embedding representation of the tile;
[0010] S5: Train the tile feature embedding module and the spot feature embedding module through the contrastive learning module according to the population feature embedding representations of the positive and negative samples; extract the embedding representation of the tile through the trained tile feature embedding module;
[0011] S6: Use the tiles that lack the expression of the expressed genes as tiles to be imputed, select the top I tiles that have the expression of the expressed gene and are most similar to the embedding representation of the tile to be imputed as the similar tiles of the tile to be imputed; use the expression values of the expressed gene of the similar tiles of the tile to be imputed to re-impute the tile to be imputed.
[0012] Another aspect of the present invention provides a gene expression deletion filling system based on deep learning, the system includes a memory and a processor; the memory is used to store application programs; the processor is used to run the application programs and execute the above-mentioned gene expression deletion filling method based on deep learning.
[0013] The present invention has at least the following beneficial effects:
[0014] The present invention divides the detection points into multiple tiles and spots. By counting the expressed genes within the spots and processing the tiles, it can effectively capture the interactions between local regions, making the filling results more consistent in the local regions and improving the local consistency of the filling results. The method proposed by the present invention, through a strategy based on graph neural networks and contrastive learning, to a certain extent avoids the problem of overly complex models. While ensuring the filling accuracy, it can process data more efficiently, achieving a balance between model complexity and efficiency, and is more suitable for application in filling large-scale spatial transcriptome data. Through the calculation of population feature embeddings of positive and negative samples and contrastive learning training, the present invention can more accurately extract the embedding representation of the tiles, thereby selecting more appropriate similar tiles to re-impute the tiles to be imputed, improving the accuracy of the filling results. At the same time, fully considering the spatial information and local region interactions also enhances the spatial continuity of the filling results, solving the deficiencies of existing methods in terms of spatial continuity. BRIEF DESCRIPTION OF THE DRAWINGS
[0015] Figure 1 It is a schematic flowchart of the method of the present invention. DETAILED DESCRIPTION OF THE INVENTION
[0016] The following uses specific specific examples to illustrate the implementation manners of the present invention. Those skilled in the art can easily understand other advantages and effects of the present invention from the content disclosed in this specification. The present invention can also be implemented or applied through other different specific implementation manners. Various details in this specification can also be modified or changed based on different viewpoints and applications without departing from the spirit of the present invention. It should be noted that the diagrams provided in the following embodiments only illustrate the basic concept of the present invention in a schematic manner. Without conflict, the following embodiments and the features in the embodiments can be combined with each other.
[0017] Please refer to Figure 1 , one aspect of the present invention provides a method for filling missing gene expressions based on deep learning, including:
[0018] S1: Preprocess the original data of the detection points on the tissue section to obtain high-resolution spatial transcriptome data, where the high-resolution spatial transcriptome data includes the gene expression vectors of the detection points and the position information of the detection points; the original data includes: the original positions of the detection points, the position of the tissue section, and the reads set R.
[0019] Preferably, the preprocessing of the original data of the detection points on the tissue section includes:
[0020] S11: Use the fastqc software to check the readsSet R, filtering set R to obtain the read set collected at the detection point by filtering the adapter sequences and low-quality reads in ; In this embodiment, reads refer to short sequence fragments generated during high-throughput sequencing, and these short sequences are the results obtained by sequencing an RNA sample. Adapter sequences are non-target sequences introduced during high-throughput sequencing, and adapter sequences are used to fix RNA fragments on the sequencing chip. In the present invention, the fastqc software is used to detect whether there are adapter sequences in the reads and filter them. The determination of low-quality reads is based on the Phred quality score calculated by fastqc. The Phred score is used to quantify the probability of sequencing errors, and its calculation formula is: Q = -10 × log 10 (P), where Q is the Phred score and P is the probability of base sequencing error. A Phred score of 30 indicates a base error rate of 0.1%, that is, 10 -3 , in this method, the average Phred quality score of a single read ≥ 30 is defined as a high-quality read, while a read with a score < 30 is defined as a low-quality read and needs to be filtered. The determination conditions of the Phred score can be adjusted according to the actual situation. In this embodiment, only 30 is taken as an example.
[0021] S12: Use the STAR software to map each read in the read set to the reference genome to obtain the alignment result set ;
[0022] S13: Use the HTseq software to count the reads of each gene in the target gene set according to the alignment result set M to generate the gene expression vector of the detection point;
[0023] S14: Match the spatial information of the tissue section with the position coordinates of all detection points, and correct the position information of each detection point to obtain the final position information of the detection point;
[0024] Preferably, step S14 includes:
[0025] S141: Determine a unified spatial coordinate system according to the spatial information of the tissue section; for example, take the upper left corner of the tissue section as the coordinate origin (0, 0), define the horizontal direction as the x-axis, the vertical direction as the y-axis, and establish a two-dimensional spatial coordinate system. If it is three-dimensional tissue section data, the z-axis also needs to be defined.
[0026] S142: Map the position coordinates of the detection points into this unified spatial coordinate system; for example, convert the relative coordinates based on the sequencing technology into the absolute coordinates based on the tissue section image.
[0027] S143: Use image processing techniques and algorithms to match the positions of the detection points with the spatial information of the tissue section; for example, an image registration algorithm can be used to match the position information of the detection points with specific features on the tissue section image. Specifically, by identifying markers in the tissue section image (such as cell structures, specific stained regions, etc.), the positions of the detection points can be associated with the positions of these markers.
[0028] S144: For each detection point, find the corresponding position in the spatial information of the tissue section;
[0029] S145: According to the results of the position matching, correct the position information of each detection point. If there are deviations in the original coordinates of the detection points, adjust the coordinate values of the detection points through matching with the spatial information of the tissue section. For example, if there are differences between the actual position of the detection point in the image and the original coordinates, the coordinates of the detection point can be corrected according to the coordinate system of the image. Considering factors such as possible deformation of the tissue section during preparation and detection, further optimize and correct the position information of the detection points. For example, an elastic registration algorithm can be used to dynamically adjust the positions of the detection points to adapt to the deformation of the tissue section.
[0030] S146: Verify the corrected position information of the detection points to ensure its accuracy and reliability. It can be compared with the known spatial information of the tissue section to check whether the positions of the detection points are reasonable. Adopt certain evaluation indicators, such as position error, matching degree, etc., to evaluate the corrected position information. If the evaluation results are not satisfactory, return to the previous steps to adjust and optimize the position matching and correction process.
[0031] S15: Standardize the gene expression vectors of the detection points using the TPM method to obtain the gene expression vectors of the detection points after standardization processing.
[0032] Preferably, the step S15 includes:
[0033] S151: For each gene in the target gene set , obtain its gene length information; the gene length can be obtained by querying the gene database or referring to the genome annotation file. Generally, the gene length refers to the nucleotide sequence length from the start codon to the stop codon of the gene, in base pairs (bp).
[0034] S152: According to the results of counting the reads of each gene in the target gene set using the HTseq software in step S13 obtain the number of reads of each gene . Here, represents the number of times the gene is sequenced at this detection point .
[0035] S153: For each gene , calculate its reads per kilobase (RPK) value , and the calculation formula is:
[0036]
[0037] where is the length of the gene in kilobases; The RPK value reflects the expression level of the gene per unit length, eliminating the influence of gene length on read counting.
[0038] S154: Add up the RPK values of all genes to obtain the total RPK value at this detection point , where is the total number of genes in the target gene set Genes.
[0039] S155: For each gene , calculate its TPM value:
[0040]
[0041] where represents the TPM value of the gene ;
[0042] The TPM value represents the number of transcripts of this gene per million transcripts, which is a standardized gene expression index and can be used to compare the gene expression levels between different detection points or different samples.
[0043] Through the above steps, the gene expression vector at the detection point can be standardized using the TPM method to obtain the gene expression vector at the detection point after standardization, making the expression levels between different genes comparable and providing a unified standard for subsequent data analysis and comparison.
[0044] In this embodiment, through step S11, the fastqc software is used to comprehensively detect the raw data, accurately filter out low-quality reads and adapter sequences, effectively eliminate interference factors such as sequencing errors and contamination, and obtain a pure reads set R1, which builds a reliable data foundation for subsequent analysis; through step S12, the STAR software is used to accurately align the reads in R1 with the reference genome Gref, establish the correspondence between the sequencing data and the genome, and clarify the positions of the reads in the genome. This not only provides a basis for subsequent gene expression level calculation but also can mine gene structure and variation information, enhancing the biological significance of the data. Through step S13, the HTseq software is used to count the reads of the target gene set Genes, convert the raw sequencing data into an intuitive gene expression vector, quantify the expression level of each gene at the detection points, and provide core data for gene expression pattern analysis; through step S14, by matching the spatial information of tissue sections with the coordinates of the detection points, the position information is corrected to accurately associate the gene expression data with the tissue spatial structure, laying a foundation for studying the spatial distribution law of gene expression and analyzing the tissue microenvironment; through step S15, the TPM method is used to standardize the gene expression vector, eliminate the bias caused by differences in sequencing depths of different samples, ensure the comparability of gene expression data among different detection points, and enhance the analysis value of the data.
[0045] After this series of preprocessing, the finally obtained high-resolution spatial transcriptome data contains both accurate gene expression information and integrated reliable position information, providing high-quality data support for subsequent operations such as dividing tiles and spots, counting expressed genes, constructing positive and negative samples, and implementing contrastive learning training. It not only significantly improves the accuracy of filling gene expression missing values but also enhances the spatial continuity of the filling results. At the same time, it effectively balances the model complexity and computational efficiency, creating favorable conditions for in-depth analysis and application of large-scale spatial transcriptome data.
[0046] S2: Divide the detection points into multiple tiles and spots. Among them, each spot contains multiple tiles, and each tile contains multiple detection points; where a tile refers to a square detection area containing N×N detection points, and a spot refers to a square area containing M×M tile areas, N≥10; M≥10;
[0047] Preferably, the dividing the detection points into multiple tiles and spots includes: the detection points are distributed in an array on the tissue section to form a two-dimensional detection point grid area of M×M, and each grid represents a detection point; the two-dimensional grid area of M×M is divided into N×N tiles of size A×A, where N = M÷A; each tile is taken as a grid to form a two-dimensional tile grid area of N×N, and the two-dimensional tile grid area of N×N is divided into K×K spots of size B×B; where K = N÷B, and M, N, A, B, and K are all positive integers; the gene expression vectors of all detection points within each tile are added to obtain the gene expression vector of the tile; the position of each tile takes the central value of the positions of all detection points within the tile.
[0048] In this embodiment, first, the M×M two-dimensional grid region is further divided into N×N tiles of size A×A, where N = M÷A. This division method is equivalent to a preliminary "grid-based subdivision" of the entire detection region, enabling the originally scattered detection points to be aggregated in groups of A×A points. Each tile becomes a basic unit containing multiple detection points. For example, if M = 100 and A = 10, then 10×10 tiles will be formed, and each tile covers 10×10 = 100 detection points. Such a division facilitates the subsequent integration and analysis of gene expression characteristics in local regions. Next, each tile is regarded as a new "grid", and an N×N two-dimensional tile grid region is constructed, and on this basis, it is further divided into K×K spots of size B×B, where K = N÷B. This step is a further aggregation of tile units, forming a set of larger-scale local regions. Continuing with the previous example, if B = 2, then N = 10 and K = 5, and finally 5×5 spots will be obtained, and each spot contains 2×2 tiles. Through this hierarchical division strategy, from detection points to tiles, and then to spots, a spatial region system from small to large and from micro to macro is constructed, realizing the structured decomposition and integration of the spatial information of tissue sections. The gene expression vectors of all detection points within each tile are added to obtain the gene expression vector of the tile. This operation effectively aggregates the gene expression information of each detection point within the tile. On the one hand, through the summation operation, the overall trend of gene expression in this local region can be highlighted, reducing the impact of data fluctuations at individual detection points on the overall characteristics; on the other hand, compressing the gene expression vectors of multiple detection points into the gene expression vector of one tile greatly reduces the data dimension, reduces the complexity of subsequent calculations, and improves the model processing efficiency. At the same time, the position of each tile takes the central value of the positions of all detection points within the tile. This setting gives the tile a clear and representative spatial coordinate, enabling the accurate consideration of the spatial relationship between tiles in subsequent analysis and providing a precise position reference for filling missing gene expression values using spatial information. In terms of effect, such a division method brings many advantages.First, by dividing the detection points into tiles and spots, the spatial location information of the spatial transcriptome data is fully utilized, enabling the analysis and calculation based on the spatial structure characteristics of the local area during the subsequent filling process of missing gene expression values, and improving the spatial accuracy of the filling results. Second, the constructed tile and spot hierarchical system strengthens the interaction between local areas. When performing gene expression analysis and missing value imputation, the cooperative relationship within and between local areas can be better considered, ensuring the consistency of the filling results in the local area. Third, through the aggregation of gene expression vectors and the regularization of position information, the model complexity and efficiency are balanced to a certain extent. It can not only retain key spatial and gene expression characteristics but also reduce data redundancy, enabling the model based on graph neural networks and contrastive learning to process data more efficiently and be more suitable for the analysis and application of large-scale spatial transcriptome datasets.
[0049] S3: Statistically analyze the genes with expression frequencies higher than the set threshold within the spot as the expressed genes to construct an expressed gene set; perform preliminary imputation on the tiles within the spot that lack the expression of the expressed genes through the KNN algorithm.
[0050] Preferably, the preliminary imputation of the tiles within the spot that lack the expression of the expressed genes through the KNN algorithm includes: regarding the tiles that lack the expression of the expressed genes as the tiles to be imputed; selecting the nearest K0 tiles with the expression of the expressed genes to the tile to be imputed according to the position of the tile to be imputed; and using the mean value of the expression levels of the expressed genes of the nearest K0 tiles with the expression of the expressed genes to the tile to be imputed as the expression value of the tile to be imputed for the expressed genes, and performing preliminary imputation on the tile to be imputed. Among them, the value range of K0 is: 0.05×M×M~0.1×M×M, and M×M is the number of tiles contained within the spot.
[0051] In this embodiment, within each spot, the gene expression frequencies are statistically analyzed and the genes higher than the set threshold are screened to construct an expressed gene set. This operation focuses on identifying the genes with active expression characteristics in the local area. The purpose of setting the threshold is to filter out the genes with low expression, which may be caused by experimental noise or accidental factors, and retain the genes that truly play a role in this area. For example, if the set threshold is 0.8, it means that within a spot, only when a certain gene is expressed in 80% or more of the tiles will it be included in the expressed gene set. This screening mechanism effectively reduces data redundancy, highlights the key gene expression information, enables the subsequent analysis to focus on the core genes, and avoids the influence of secondary or interfering gene expression data on the filling results. At the same time, constructing the expressed gene set clarifies the target gene range for the subsequent imputation operation, making the imputation process more targeted and improving the calculation efficiency.
[0052] After marking the tiles in the spot that lack the expression of the expressed gene as the tiles to be imputed, the KNN algorithm takes effect based on the spatial position information. According to the position of the tile to be imputed, the algorithm searches for the nearest K0 tiles with the expression of the corresponding expressed gene within the same spot. Here, the "distance" is usually calculated using metrics such as the Euclidean distance to quantify the spatial proximity between the tiles. For example, if K0 is set to 5, then the 5 tiles with the expression of the target gene that are closest to the tile to be imputed are selected. Subsequently, the average expression level of the expressed gene of these K0 tiles is used as the expression value of the tile to be imputed for preliminary imputation. This method is based on the assumption of spatial proximity, believing that tiles that are close in spatial position have similar gene expression patterns, and uses the expression information of the surrounding tiles to provide a reference value for the tile to be imputed, which can initially fill in the missing gene expression data.
[0053] Looking at the overall technical process, the dual operations in step S3 bring significant effects. At the level of data feature extraction, the construction of the expressed gene set realizes the feature screening and dimensionality reduction of the spatial transcriptome data, focusing on the core gene expression information, and provides a cleaner and more representative data input for subsequent contrast learning and model training. At the level of missing value processing, the preliminary imputation of the KNN algorithm is based on the spatial position relationship, effectively utilizing the spatial dependence and expression similarity between the tiles within the local area, and alleviating the problem of gene expression missing to a certain extent. Although the result of this preliminary imputation may have insufficient accuracy, it lays a foundation for subsequent fine-tuning through contrast learning and deep learning models, reduces the difficulty and computational amount of subsequent model training, and at the same time provides a more complete data basis for subsequent complex analyses based on tile and spot features, improving the consistency in the local area and the spatial accuracy of the filling result, and contributing to more accurate filling of missing gene expression values.
[0054] S4: Select tiles from the spot to construct positive and negative samples, and calculate the population feature embeddings of the positive and negative samples through the tile feature embedding module and the spot feature embedding module; wherein, the tile feature embedding module is used to obtain the embedding representation of the tile; the spot feature embedding module is used to obtain the population feature embedding representation of the positive and negative samples according to the embedding representation of the tile;
[0055] Preferably, the construction of the positive and negative samples includes: taking all the tiles of the spot as the first positive sample; selecting the tiles in the 80%-90% square area within the spot as the second positive sample;
[0056] Select tiles of size C×C at the four corners of the spot, and randomly select P tiles as negative samples in each tile of size C×C, where , .
[0057] Preferably, the tile feature embedding module uses a first Transformer encoder to obtain the embedding representation of the tile, including: expanding the gene expression vector of the imputed tile into a two-dimensional matrix to adapt to the input format of the first Transformer encoder, and mapping the gene expression vector expanded into a two-dimensional matrix to a high-dimensional space; encoding the position coordinates of the tile using Sinusoidal, and concatenating the position encoding features of the tile and the gene expression vector mapped to the high-dimensional space to obtain the input features of the first Transformer encoder, and encoding the input features through the first Transformer encoder to obtain the embedding representation of the tile.
[0058] Preferably, the spot feature embedding module includes: a graph structure construction module, a GAT encoding module, and a deep feature extraction module; the graph structure construction module is used to construct a graph structure according to the embedding representation and position of the tiles in the positive and negative samples; the GAT encoding module aggregates and updates the feature matrix of the graph structure using multiple GAT layers according to the adjacency matrix of the graph structure; the deep feature extraction module includes: a second Transformer encoder, an attention pooling layer, a global pooling layer, and a fully connected layer; the input feature of the second Transformer encoder is the feature matrix aggregated and updated by the GAT encoding module; the input features of the second Transformer encoder are respectively used as the input features of the attention pooling layer and the global pooling layer; the output features of the attention pooling layer and the global pooling layer are concatenated and used as the input features of the fully connected layer; the output feature of the fully connected layer is the population feature embedding representation of the positive and negative samples.
[0059] Preferably, constructing the graph structure according to the embedding representation and position of the tiles in the positive and negative samples includes:
[0060] S41: Use the tiles in the positive and negative samples as nodes, and use the embedding representation of the tiles as the feature of the tile nodes to construct the initial feature matrix of the graph structure;
[0061] S42: Connect each tile to its K1 nearest neighbor tiles, and calculate the local weight of the edge based on the positions of the tile and its connected neighbor tiles; where the value range of K1 is an integer from 1 / 4×M to 1 / 2×M; the calculation of the local weight of the edge includes: calculating the Euclidean distance between the positions of the tile and its connected neighbor tiles, and performing normalization to obtain the local weight of the edge. The formula for the local weight of the edge is: , where represents the local weight of the edge between tile i and tile j, is the positional distance between two connected tiles i and j, and max_distance is the maximum distance for normalization.
[0062] S43: Connect each tile to its K2 most similar tiles, and calculate the global weight of the edge based on the features of the tile and its connected similar tiles; where the value range of K2 is from 0.05×M×M to 0.1×M×M, and M×M is the number of tiles contained in the spot; the calculation of the global weight of the edge includes: calculating the similarity between the features of the tile and its connected similar tiles, and performing normalization to obtain the global weight of the edge. The formula for the global weight of the edge is: , where is the positional distance between two connected tiles i and j, and are the gene expression vectors (features) between them, max_distance is the maximum distance for normalization, is the cosine similarity between tile i and tile j; represents the global weight of the edge between tile i and tile j;
[0063] S44: Remove the edges with duplicate connections between tiles, and retain the maximum value of the local weight and the global weight as the final weight of the edge, and then construct the adjacency matrix of the graph structure; where the formula for the final weight of the edge is: , is the final weight of the edge between tile i and tile j; construct the adjacency matrix A of the graph structure: if there is an edge between node tiles i and j, then A ij =W ij , otherwise A ij =0; A ij represents the element in the i-th row and j-th column of the adjacency matrix A;
[0064] In this embodiment, positive and negative samples are constructed in a hierarchical and targeted selection manner. All tiles of the spot are used as the first positive samples, and tiles in 80%-90% of the square regions within the spot are selected as the second positive samples. This selection method ensures that the positive samples cover the main gene expression information and spatial structure features within the spot and can represent the common patterns of gene expression within the spot. Specific-sized tiles are selected from the four corners of the spot and negative samples are randomly drawn. By leveraging the spatial differences between the corner regions and the central region, samples with different expression patterns are introduced, ensuring an effective distinction between negative samples and positive samples in terms of spatial location and expression characteristics. The reasonable construction of positive and negative samples provides rich and contrasting data samples for contrastive learning, enabling the model to learn the differential features of different gene expression patterns and spatial structures, thereby enhancing the model's ability to distinguish the characteristics of spatial transcriptome data.
[0065] In this embodiment, the tile feature embedding module uses a first Transformer encoder, and its processing process fully integrates gene expression information and spatial location information. First, the gene expression vector of the imputed tile is extended into a two-dimensional matrix. This operation not only adapts to the input format of the Transformer encoder but also provides a structured basis for subsequent exploration of the potential features of gene expression data. Then, by mapping the gene expression vector into a high-dimensional space, complex non-linear relationships between gene expressions can be captured, and richer expression features can be mined. The position coordinates of the tile are encoded using Sinusoidal, converting the position information into computable features and concatenating them with the gene expression features, enabling the model to consider both gene expression and spatial location information during the learning process and avoiding the defect of existing methods that ignore spatial information. Finally, after encoding by the first Transformer encoder, the obtained tile embedding representation contains dual features of gene expression and spatial location, providing a more comprehensive and accurate feature input for the subsequent calculation of the spot feature embedding module and contrastive learning.
[0066] In this embodiment, the spot feature embedding module works in cooperation with the graph structure construction module, the GAT encoding module, and the deep feature extraction module. The graph structure construction module constructs an initial feature matrix by using tiles as nodes and embedding representations as node features, connects the nodes based on position and features, and calculates edge weights. The constructed graph structure not only reflects the spatial proximity relationship between tiles but also embodies the similarity relationship of gene expression features. The GAT encoding module aggregates and updates the feature matrix using multiple GAT layers based on the adjacency matrix of the graph structure, effectively capturing the interactions between tiles in local regions and making the feature representation more locally consistent. In the deep feature extraction module, the second Transformer encoder further extracts features, combines the advantages of the attention pooling layer and the global pooling layer, screens and integrates the features, and finally outputs the population feature embedding representations of positive and negative samples through a fully connected layer. This multi-level feature extraction and integration mechanism fully explores the features of the tile population in positive and negative samples, retains the detailed features of local regions, and extracts the overall global features, providing high-quality feature representations for contrastive learning, helping the model learn the essential differences between positive and negative samples, and enhancing the model's understanding and processing ability of spatial transcriptome data features.
[0067] By constructing positive and negative samples and calculating feature embeddings, a solid foundation for contrastive learning is laid at both the data sample and feature representation levels. At the data level, the reasonably constructed positive and negative samples provide rich and discriminative training materials for contrastive learning, guiding the model to learn the feature differences of different gene expression patterns and spatial structures. At the feature level, the tile and spot feature embedding modules deeply fuse spatial position and gene expression information, and the extracted population feature embedding representations accurately reflect the spatial and expression features of the samples, enabling the model to more accurately learn the correlations and differences between features in subsequent contrastive learning training, effectively solving the problems of insufficient utilization of spatial information and inadequate capture of local region interactions in existing methods, providing strong support for improving the accuracy and spatial continuity of gene expression missing value filling results, and at the same time balancing the model complexity and efficiency, making it applicable to the processing of large-scale spatial transcriptome data.
[0068] S5: Train the tile feature embedding module and the spot feature embedding module through the contrastive learning module according to the population feature embedding representations of positive and negative samples; extract the embedding representation of the tile through the trained tile feature embedding module;
[0069] Preferably, the contrastive learning module uses the noise contrastive estimation function InfoNCE Loss:
[0070]
[0071] Among them, log represents the natural logarithm, exp represents the exponential function, L represents the loss function, q represents the population feature embedding representation of the first positive sample, k+ represents the population feature embedding representation of the second positive sample, k− represents the population feature embedding representation of the negative sample, and τ is the temperature parameter; sim represents the cosine similarity function.
[0072] S6: Take the tile lacking the expression of the expressed gene as the tile to be imputed, and select the top I tiles that have the expression of the expressed gene and are most similar to the embedding representation of the tile to be imputed as the similar tiles of the tile to be imputed; use the expression values of the similar tiles of the tile to be imputed for the expressed gene to re-impute the tile to be imputed. The value range of I is 0.05×M×M~0.1×M×M, and M×M is the number of tiles contained in the spot.
[0073] In the previous steps of this embodiment, the detection points on the tissue section have been divided to form tiles and spots, and the set of expressed genes has been counted. In this step, the tiles lacking the expression of the expressed gene are marked as tiles to be imputed. These tiles may not have detected the expression of certain expressed genes in the original data due to technical limitations, experimental errors, etc. To impute the tiles to be imputed, similar tiles need to be found. Here, the embedding representation of the tiles is used to measure the similarity. In step S4, the embedding representation of each tile has been calculated through the tile feature embedding module and the spot feature embedding module. These embedding representations contain various features such as the gene expression information and spatial position information of the tiles. From all the tiles that have the expression of the expressed gene, select the top I tiles that are most similar to the embedding representation of the tile to be imputed as the similar tiles of the tile to be imputed. The similarity can be measured using common distance metrics such as the Euclidean distance and cosine similarity. For example, if the cosine similarity is used, the higher the similarity, the more similar the embedding representations of the two tiles. Use the expression values of the similar tiles of the tile to be imputed for the expressed gene to re-impute the tile to be imputed. A common imputation method is to calculate the mean of the expression values of these I similar tiles for the expressed gene, and use this mean as the expression value of the tile to be imputed for the expressed gene.
[0074] By selecting similar tiles for imputation through tile-based embedded representation, both gene expression information and spatial location information are fully utilized. The embedded representation is carefully calculated through the previous steps and can comprehensively reflect various features of the tile. Compared with traditional imputation methods based only on gene expression values or spatial locations, this approach can more accurately find tiles that are similar to the tile to be imputed in both gene expression patterns and spatial distributions, thereby improving the accuracy of imputation. Since the embedded representation contains spatial location information, the selected similar tiles are usually also spatially close to the tile to be imputed. In this way, the imputed result will be more continuous spatially, avoiding obvious discontinuities or mutations. For example, in tissue sections, the gene expression values of adjacent tiles after imputation will be more in line with the actual biological laws, making the gene expression distribution of the entire tissue section more reasonable. Step S6 is carried out based on the previous steps. The previous steps, such as data preprocessing, tile and spot division, construction of the expressed gene set, and calculation of feature embedding, all provide strong support for this step of imputation. Through this progressive approach, the quality of filling in missing gene expression values is gradually improved, so that the final filling result can better reflect the true gene expression situation of the tissue section. The number I of similar tiles selected can be adjusted according to the specific dataset and analysis requirements. When the data volume is large and the gene expression pattern is relatively complex, the value of I can be appropriately increased to obtain more similar information; when the data volume is small or more attention is paid to local features, the value of I can be decreased so that the imputation result focuses more on the few tiles that are most similar to the tile to be imputed. This flexibility enables the method to adapt to different types of spatial transcriptome data.
[0075] Another aspect of the present invention provides a gene expression missing value filling system based on deep learning. The system includes a memory and a processor; the memory is used to store application programs; the processor is used to run the application programs and execute the described gene expression missing value filling method based on deep learning.
[0076] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, storage, database, or other medium used in the various embodiments provided in the present application can include non-volatile and / or volatile memories. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory can include random access memory (RAM) or an external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and Rambus dynamic RAM (RDRAM), etc.
[0077] In summary, the present invention divides the detection points into multiple tiles and spots. By counting the expressed genes within the spots and processing the tiles, it can effectively capture the interactions between local regions, making the filling results more consistent in the local regions and improving the local consistency of the filling results. The method proposed by the present invention, through a strategy based on graph neural networks and contrastive learning, to a certain extent avoids the problem of overly complex models. While ensuring the filling accuracy, it can process data more efficiently, achieving a balance between model complexity and efficiency, and is more suitable for application to large-scale spatial transcriptome data filling. Through the calculation of population feature embeddings of positive and negative samples and contrastive learning training, the present invention can more accurately extract the embedding representation of the tiles, thereby selecting more appropriate similar tiles to re-fill the tiles to be interpolated, improving the accuracy of the filling results. At the same time, fully considering the spatial information and local region interactions also enhances the spatial continuity of the filling results, solving the deficiencies of existing methods in terms of spatial continuity.
[0078] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit them. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that the technical solutions of the present invention can be modified or equivalently replaced without departing from the purpose and scope of the present technical solution, and they should all be covered within the scope of the claims of the present invention.
Claims
1. A method for filling in missing gene expression based on deep learning, characterized in that, Including: S1: Preprocess the original data of the detection points on the tissue section to obtain high-resolution spatial transcriptome data, where the high-resolution spatial transcriptome data includes the gene expression vector of the detection point and the position information of the detection point; S2: Divide the detection points into multiple tiles and spots, where each spot contains multiple tiles, and each tile contains multiple detection points; S3: Count the genes with expression frequencies higher than the set threshold within the spot as expressed genes to construct a set of expressed genes; perform preliminary imputation on the tiles lacking the expression of the expressed genes within the spot through the KNN algorithm; S4: Select tiles from the spot to construct positive and negative samples, and calculate the population feature embeddings of the positive and negative samples through the tile feature embedding module and the spot feature embedding module; where the tile feature embedding module is used to obtain the embedding representation of the tile; the spot feature embedding module is used to obtain the population feature embedding representation of the positive and negative samples according to the embedding representation of the tile; S5: Train the tile feature embedding module and the spot feature embedding module through the contrast learning module according to the population feature embedding representations of the positive and negative samples; extract the embedding representation of the tile through the trained tile feature embedding module; S6: Take the tiles lacking the expression of the expressed genes as the tiles to be imputed, select the top I tiles with the expression of the expressed gene and the most similar embedding representation to the tile to be imputed as the similar tiles of the tile to be imputed; use the similar tiles of the tile to be imputed to re-impute the expression value of the expressed gene for the tile to be imputed.
2. The method for filling in missing gene expression based on deep learning according to claim 1, characterized in that, The preprocessing of the original data of the detection points on the tissue section includes: S11: Use the fastqc software to check the reads set R, filter the R adapter sequences and low-quality reads in the reads set to obtain the set collected at the detection point; S12: Use the STAR software to align each read in the reads set to the reference genome and obtain the alignment result set ; ; S13: According to the comparison result set M, use the HTseq software to count the reads of each gene in the target gene set to generate a gene expression vector for the detection point for each gene in the set; S14: Match the spatial information of the tissue section with the position coordinates of all detection points, and correct the position information of each detection point to obtain the final position information of the detection point; S15: Standardize the gene expression vector of the detection point using the TPM method to obtain the standardized gene expression vector of the detection point.
3. The method for filling in missing gene expression based on deep learning according to claim 1, characterized in that, The division of the detection points into multiple tiles and spots includes: The detection points are distributed in an array on the tissue section to form a two-dimensional detection point grid area of M×M, and each grid represents a detection point; divide the two-dimensional grid area of M×M into N×N tiles of size A×A, where N = M÷A; take each tile as a grid to form a two-dimensional tile grid area of N×N, and divide the two-dimensional tile grid area of N×N into K×K spots of size B×B; where K = N÷B, and M, N, A, B, and K are all positive integers; add the gene expression vectors of all detection points within each tile to obtain the gene expression vector of the tile; take the central value of the positions of all detection points within each tile as the position of the tile.
4. The method for filling in missing gene expression based on deep learning according to claim 3, characterized in that, The construction of positive and negative samples includes: Taking all the tiles of the spot as the first positive sample; Selecting the tiles in the 80%-90% square area within the spot as the second positive sample; Select tiles of size C×C at the four corners of the spot, and randomly select P tiles as negative samples in each tile of size C×C, where, , .
5. The method for filling in missing gene expression based on deep learning according to claim 4, characterized in that, The contrastive learning module uses the InfoNCE Loss based on the noise contrastive estimation function: where log represents the natural logarithm, exp represents the exponential function, L represents the loss function, represents the group feature embedding representation of the first positive sample, k+ represents the group feature embedding representation of the second positive sample, k− represents the group feature embedding representation of the negative sample, and τ is the temperature parameter; sim represents the cosine similarity function.
6. The method for filling in missing gene expression based on deep learning according to claim 5, characterized in that, The preliminary imputation of the tiles lacking the expression of the target gene within the spot by the KNN algorithm includes: taking the tiles lacking the expression of the target gene as the tiles to be imputed; selecting the K0 tiles with the expression of the target gene closest to the tile to be imputed according to the position of the tile to be imputed; and using the mean value of the expression levels of the target gene of the K0 tiles closest to the tile to be imputed as the expression value of the target gene of the tile to be imputed, and performing preliminary imputation on the tile to be imputed.
7. The method for filling in missing gene expression based on deep learning according to claim 6, characterized in that, The tile feature embedding module uses the first Transformer encoder to obtain the embedding representation of the tile, including: expanding the gene expression vector of the imputed tile into a two-dimensional matrix to adapt to the input format of the first Transformer encoder, and mapping the gene expression vector expanded into a two-dimensional matrix to a high-dimensional space; encoding the position coordinates of the tile using Sinusoidal, and concatenating the position encoding features of the tile and the gene expression vector mapped to the high-dimensional space to obtain the input features of the first Transformer encoder, and encoding the input features through the first Transformer encoder to obtain the embedding representation of the tile.
8. The method for filling in missing gene expression based on deep learning according to claim 1, characterized in that, The spot feature embedding module includes: a graph structure construction module, a GAT encoding module, and a deep feature extraction module; the graph structure construction module is used to construct a graph structure according to the embedding representation and position of the tiles in the positive and negative samples; the GAT encoding module aggregates and updates the feature matrix of the graph structure using multiple GAT layers according to the adjacency matrix of the graph structure; the deep feature extraction module includes: a second Transformer encoder, an attention pooling layer, a global pooling layer, and a fully connected layer; the input feature of the second Transformer encoder is the feature matrix aggregated and updated by the GAT encoding module; the input feature of the second Transformer encoder is respectively used as the input feature of the attention pooling layer and the global pooling layer; the output features of the attention pooling layer and the global pooling layer are concatenated and used as the input feature of the fully connected layer; the output feature of the fully connected layer is the population feature embedding representation of the positive and negative samples.
9. For the method for filling in missing gene expression based on deep learning according to claim 8, the construction of the graph structure according to the embedded representations and positions of tiles in positive and negative samples includes: S41: Taking the tiles in the positive and negative samples as nodes, and using the embedding representation of the tiles as the initial feature matrix of the graph structure of the tile nodes; S42: Connecting each tile to its K1 nearest neighbor tiles, and calculating the local weight of the edge based on the positions of the tile and its connected neighbor tiles; S43: Connecting each tile to its K2 most similar tiles, and calculating the global weight of the edge based on the features of the tile and its connected similar tiles; S44: Remove the edges that are repeatedly connected between tiles, and retain the maximum value of the local weight and the global weight as the final weight of the edge, and then construct the adjacency matrix of the graph structure of the graph.
10. A gene expression deletion filling system based on deep learning, the system comprising a memory and a processor; the memory is used for storing application programs; the processor is used for running the application programs and executing a gene expression deletion filling method according to any one of claims 1 to 9.
Citation Information
Patent Citations
Space-time crack data inference method based on graph and attention
CN116304511A
Spatial transcriptome spot region clustering method fusing image gene data
CN116312782A
Spatial transcriptome spatial domain identification method and system based on neighbor contrast learning
CN117976057A
Non-intrusive load decomposition method based on GAT and Transform
CN118820904A
Defect detection method based on deep contrast learning
CN118967690A
Cited By
Space transcriptome gene expression prediction method and system, terminal and storage medium
CN121354663A