Idle data feature extraction method, spatial domain identification method and system

Through idle data feature extraction method and multi-scale hypergraph autoencoder, the problem of low recognition accuracy in spatial transcriptome data analysis is solved, and high-precision spatial domain recognition and biological information capture are achieved.

CN120388612AActive Publication Date: 2025-07-29SHANDONG UNIV

Patent Information

Application Number
CN202510872837.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-27
Publication Date
2025-07-29
Estimated Expiration
2045-06-27

AI Technical Summary

Technical Problem

In the analysis of spatial transcriptome data, it is difficult to accurately identify the spatial domain, which is affected by high-dimensionality, sparseness, sequencing technology uncertainty and experimental conditions, resulting in low recognition accuracy.

Method used

Idle data feature extraction methods are adopted, including denoising processing, position encoding splicing and mask autoencoder encoding, combined with multi-scale hypergraph autoencoder, to build a loss function training model and extract low-order potential representations.

Benefits of technology

It improves the accuracy of spatial domain identification, captures multi-level biological information, and provides a good upstream analysis foundation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120388612A_ABST
    Figure CN120388612A_ABST
Patent Text Reader

Abstract

The invention belongs to the field of spatial transcriptome spatial domain recognition, and provides an idling data feature extraction method, a spatial domain recognition method and a spatial domain recognition system in order to solve the problem of poor recognition accuracy of a spatial transcriptome spatial domain. The idling data feature extraction method comprises the following steps: acquiring idling data of a pathological section, wherein the idling data comprises an initial gene expression matrix and spatial position information; performing denoising processing on the initial gene expression matrix to obtain a denoised gene expression matrix; calculating a position code for each site according to the spatial position information, and splicing the position code with the denoised gene expression matrix to obtain an enhanced gene expression matrix; a mask auto-encoder is utilized to encode an enhanced gene expression matrix, low-order potential representation is obtained and serves as extracted idle data features, accurate recognition of a spatial domain can be achieved, and a good upstream analysis basis is provided for downstream tasks.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of spatial domain recognition of spatial transcriptomics, and particularly relates to a method for extracting features of spatial transcriptomics data, a method and a system for spatial domain recognition. Background Art

[0002] The statements in this part only provide background technical information related to the present invention and do not necessarily constitute prior art.

[0003] In the analysis of spatial transcriptomics data, one of the core tasks is the recognition of spatial domains. A spatial domain refers to a region in spatial organization with similar gene expression patterns, and its recognition result is of great significance for studying tissue functions and disease mechanisms. However, since spatial transcriptomics data usually contains the expression levels of thousands of genes at thousands of loci ( ), it has high dimensionality and high sparsity. Coupled with the uncertainty of sequencing technology and significant noise introduced by differences in experimental conditions, the accuracy of analysis is affected; in addition, the gene expression relationships between cells or are often non-linear and high-order, and it is difficult to capture complex long-range interactions only relying on simple Euclidean distances or local neighborhood relationships, ultimately affecting the recognition accuracy of spatial domains in spatial transcriptomics. Summary of the Invention

[0004] In order to solve the technical problems existing in the above background art, the present invention provides a method for extracting features of spatial transcriptomics data, a method and a system for spatial domain recognition, which can achieve accurate recognition of spatial domains and provide a good upstream analysis basis for downstream tasks.

[0005] In order to achieve the above object, the present invention adopts the following technical solutions: The first aspect of the present invention provides a method for extracting features of spatial transcriptomics data.

[0006] A method for extracting features of spatial transcriptomics data includes: Obtaining spatial transcriptomics data of a pathological section, which includes an initial gene expression matrix and spatial position information; Performing denoising processing on the initial gene expression matrix to obtain a denoised gene expression matrix; Calculating position encoding for each locus according to the spatial position information and splicing it with the denoised gene expression matrix to obtain an enhanced gene expression matrix; Encoding the enhanced gene expression matrix by using a masked autoencoder to obtain a low-order latent representation and taking it as the extracted features of spatial transcriptomics data.

[0007] As an implementation, during the process of training the masked autoencoder, an alignment loss function is constructed using the adjacency matrix reconstructed from the low-order latent representation of the gene expression matrix; a spatial consistency loss function is constructed using the high-order and low-order latent representations of the gene expression matrix; a second reconstruction loss function is constructed using the reconstructed enhanced gene expression matrix and the enhanced gene expression matrix before reconstruction. Then, the total loss function of the masked autoencoder is obtained from the alignment loss function, the spatial consistency loss function, and the second reconstruction loss function, and further, the masked autoencoder is trained.

[0008] As an implementation, the reconstructed enhanced gene expression matrix is obtained by decoding the low-order latent representation output by the masked autoencoder; the high-order latent representation of the gene expression matrix is obtained by encoding the denoised gene expression matrix using a multi-scale hypergraph autoencoder.

[0009] As an implementation, the multi-scale hypergraph autoencoder is used to learn the high-order latent representation of the denoised gene expression matrix at different scales and reconstruct the denoised gene expression matrix. Then, a first reconstruction loss function is constructed to train the multi-scale hypergraph autoencoder.

[0010] As an implementation, the spatial consistency loss function is: ; ; ; where, is the spatial consistency loss function; is the smoothing term; is the low-order latent representation of the gene expression matrix; is the high-order latent representation of the gene expression matrix; and are both intermediate parameters; , are respectively the th site corresponding vectors in the low-order and high-order latent representations of the gene expression matrix respectively; is the total number of sites; is the normalization exponential function.

[0011] As an implementation, the alignment loss function is: ; where, is the alignment loss function; , are respectively the , Laplacian matrices of; The adjacency matrix reconstructed for the low-order latent representation of the gene expression matrix; The adjacency matrix of the gene expression matrix; Denotes the F norm.

[0012] The second aspect of the present invention provides a system for extracting features of idling data.

[0013] A system for extracting features of idling data, comprising: An idling data acquisition module, which is used to acquire the idling data of the pathological section, including the initial gene expression matrix and spatial position information; A data denoising module, which is used to perform denoising processing on the initial gene expression matrix to obtain a denoised gene expression matrix; A position encoding module, which is used to calculate the position encoding for each site according to the spatial position information and splice it with the denoised gene expression matrix to obtain an enhanced gene expression matrix; A low-order latent representation module, which is used to encode the enhanced gene expression matrix by using a masked autoencoder to obtain a low-order latent representation and use it as the extracted idling data feature.

[0014] The third aspect of the present invention provides a method for identifying spatial domains of spatial transcriptomics.

[0015] A method for identifying spatial domains of spatial transcriptomics, comprising: Adopt the steps in the above-mentioned method for extracting features of idling data to extract the low-order latent representation of the idling data of the pathological section and use it as the idling data feature; Cluster the sites of the idling data feature to obtain the spatial domain identification result.

[0016] The fourth aspect of the present invention provides a system for identifying spatial domains of spatial transcriptomics.

[0017] A system for identifying spatial domains of spatial transcriptomics, comprising: A feature extraction module, which is used to adopt the steps in the above-mentioned method for extracting features of idling data to extract the low-order latent representation of the idling data of the pathological section and use it as the idling data feature; A clustering module, which is used to cluster the sites of the idling data feature to obtain the spatial domain identification result.

[0018] The fifth aspect of the present invention provides an electronic device, which includes a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, it implements the steps in the above-mentioned method for identifying spatial domains of spatial transcriptomics based on a hypergraph autoencoder.

[0019] Compared with the prior art, the beneficial effects of the present invention are: (1) The present invention performs denoising processing on the initial gene expression matrix to obtain a denoised gene expression matrix; calculates position encoding for each site according to spatial position information, and splices it with the denoised gene expression matrix to obtain an enhanced gene expression matrix. Then, a masked autoencoder is used to encode the enhanced gene expression matrix to obtain a low-order latent representation, which is used as the extracted feature of the spatial transcriptomics data. Combining the position encoding calculated for the sites with the denoised gene expression matrix improves the accuracy of feature extraction of the spatial transcriptomics data.

[0020] (2) The present invention combines the graph construction of the multi-scale hypergraph with the multi-scale hypergraph autoencoder to solve the problem that it is difficult to model the high-order relationships between sites at multiple levels, achieving the effect of capturing high-order biological information; moreover, it uses spatial consistency loss and alignment loss to explicitly fuse the high-order latent representation and the low-order latent representation, solving the problem of low interpretability of the implicit fusion of the low-order and high-order latent representations, and achieving the effect of capturing multi-level biological information.

[0021] The advantages of the additional aspects of the present invention will be partly given in the following description, partly will become obvious from the following description, or will be understood through the practice of the present invention. Description of the Drawings

[0022] The specification drawings forming a part of the present invention are used to provide a further understanding of the present invention. The schematic embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute an improper limitation of the present invention.

[0023] Figure 1 It is a flowchart of a method for extracting features of spatial transcriptomics data according to an embodiment of the present invention; Figure 2 It is the training process of the masked autoencoder and the multi-scale hypergraph autoencoder according to an embodiment of the present invention; Figure 3 It is a schematic structural diagram of a system for extracting features of spatial transcriptomics data according to an embodiment of the present invention; Figure 4 It is an architecture diagram of the multi-scale hypergraph autoencoder according to an embodiment of the present invention; Figure 5 It is a flowchart of a method for identifying the spatial domain of spatial transcriptomics according to an embodiment of the present invention; Figure 6 It is a schematic structural diagram of a system for identifying the spatial domain of spatial transcriptomics according to an embodiment of the present invention. Detailed Embodiments

[0024] The present invention will be further described below in conjunction with the drawings and embodiments.

[0025] It should be noted that the following detailed description is illustrative and is intended to provide further explanation of the present invention. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the technical field to which the present invention belongs.

[0026] It should be noted that the terms used herein are for the purpose of describing particular embodiments only and are not intended to limit the exemplary embodiments according to the present invention. As used herein, unless the context clearly dictates otherwise, the singular forms are also intended to include the plural forms. In addition, it should be understood that when the terms "comprises" and / or "comprising" are used in this specification, they specify the presence of the stated features, steps, operations, devices, components, and / or combinations thereof.

[0027] The rapid development of high-throughput sequencing technology has greatly promoted research in the fields of genomics and transcriptomics. In particular, the emergence of spatial transcriptomics technology has enabled scholars in medicine and bioinformatics to explore gene expression based on tissue spatial context. This technology can not only reveal the spatial distribution of gene expression in tissues, but also provide a new perspective for humans to understand tissue structure, gene function, and pathogenic mechanisms, and has gradually become an important component in the prevention and treatment of clinical diseases.

[0028] Term Explanation: (1) Spatial domain identification: Identifying regions with similar gene expression patterns in spatial tissues.

[0029] (2) Evaluation criteria for spatial domain identification: Assume is the number of sites belonging to the same spatial domain and the same true label class quantity, is the number of belonging to the same spatial domain but not belonging to the same true label class quantity, is the number of not belonging to the same spatial domain but belonging to the same true label class quantity.

[0030] (3) Adjusted Rand index : Definition: , where Rand sparse , is the expected Rand index.

[0031] The role of

[0032] (4) Normalized mutual information : Definition: , where are the cross-entropy of the spatial domain recognition distribution and the true label class distribution, respectively, .

[0033] Function of : Evaluate the similarity between the spatial domain recognition result and the true label class.

[0034] (5) (Fowlkes–Mallows index): Definition: ; Function of : Evaluate the consistency between the clustering result and the true label.

[0035] In one or more embodiments, as Figure 1 shown, a method for extracting idling data features according to an embodiment of the present invention includes: S101: Obtain the idling data of the pathological section, which includes a gene expression matrix and spatial position information.

[0036] The complete spatial transcriptome sequencing data corresponding to a pathological section includes a gene expression matrix and spatial position information .

[0037] S102: Perform data denoising processing on the gene expression matrix to obtain a denoised gene expression matrix; Before performing data denoising processing on the gene expression matrix, it further includes: Perform normalization and logarithmic transformation on the gene expression matrix to obtain a preprocessed gene expression matrix; Before performing normalization and logarithmic transformation on the gene expression matrix, it further includes: Screen the gene expression matrix according to the set screening conditions to achieve quality control of the gene expression matrix.

[0038] From the initial gene expression matrix screen out the gene expression matrix that meets the conditions, and obtain the output gene expression matrix that meets the conditions , where the screening conditions are as follows: Gene is expressed in at least 1% of the ; The total expression count of gene is not less than the set value, such as 200.

[0039] For the gene expression matrix that meets the conditions Normalization and logarithmic transformation are performed to obtain the normalized gene expression matrix .

[0040] Then, the normalized gene expression matrix is processed for data denoising using graph Fourier transform.

[0041] Distance metric: Based on , calculate the Euclidean distance between any two points ; Construct a graph : is a set of (denoted as ). When and only when is 's nearest neighbor or is 's nearest neighbor, and there is an edge between them. The set composed of constitutes , and is obtained; Construct the adjacency matrix , diagonal matrix : ; Diagonal matrix , where is 's degree; Fourier transform and eigen - decomposition: Calculate the Laplacian matrix . The implementation of Fourier transform in the graph is based on 's eigen - decomposition , and the eigenvalues , eigen - vectors are obtained, where the eigenvalues represent frequencies; Graph Fourier transform: Use to project the original data into the frequency space. The graph Fourier transform of the normalized gene expression matrix is ; Low - pass filter: The low - pass filter weights each frequency component, and the expression is , where , , is the smoothing parameter, ; Inverse transform of the filtered signal: Inverse - transform the filtered signal from the frequency space back to the original space ; , the output is the denoised gene expression matrix and the adjacency matrix .

[0042] S103: Calculate the positional encoding for each site according to the spatial position information, and splice it with the denoised gene expression matrix to obtain the enhanced gene expression matrix.

[0043] The goal of the positional encoding is to provide a unique representation for each position in , and this encoding is based on sine and cosine functions: ; where is the position index, is the index of the positional encoding dimension (sine is used for even dimensions and cosine is used for odd dimensions), that is, the th site; is the total dimension of the positional encoding.

[0044] Splice with to obtain the enhanced gene expression matrix: ; represents the concatenation function.

[0045] S104: Encode the enhanced gene expression matrix using a masked autoencoder to obtain a low-order latent representation and use it as the extracted quiescent data feature.

[0046] Combine Figure 2 , when training the masked autoencoder, decode the low-order latent representation output by the masked autoencoder to obtain the reconstructed enhanced gene expression matrix; use the adjacency matrix reconstructed from the low-order latent representation of the gene expression matrix to construct an alignment loss function, use the high-order latent representation and the low-order latent representation to construct a spatial consistency loss function, use the reconstructed enhanced gene expression matrix and the enhanced gene expression matrix before reconstruction to construct a second reconstruction loss function, and finally obtain the total loss function of the masked autoencoder.

[0047] Among them, the reconstructed enhanced gene expression matrix is obtained by decoding the low-order latent representation output by the masked autoencoder; the high-order latent representation of the gene expression matrix is obtained by encoding the denoised gene expression matrix using a multi-scale hypergraph autoencoder.

[0048] In the specific implementation process, a multi-scale hypergraph autoencoder is used to learn the high-order latent representation of the denoised gene expression matrix at different scales, reconstruct the denoised gene expression matrix, and then construct a first reconstruction loss function to train the multi-scale hypergraph autoencoder.

[0049] Distance metric: Based on , calculate the Euclidean distance between any two ; The process of constructing a multi-scale hypergraph is as follows: First, taking the multi-scale hypergraph as an example, illustrate how to construct a hypergraph: For each , select the nearest to form a hyperedge (these and together form the hyperedge), which can be represented by the multi-scale hypergraph incidence matrix : ; ; is an element in the multi-scale hypergraph incidence matrix ; The hyperedge set can be represented as , where , obtaining the multi-scale hypergraph .

[0050] is adaptively determined by the data to be the number of , generating hypergraph incidence matrices, and concatenating them by column to obtain the multi-scale hypergraph incidence matrix , the multi-scale hypergraph . . , ; , ; The diagonal matrices of are respectively denoted as ; is an element in the multi-scale hypergraph incidence matrix ; is a preset weight function; is the edge between sites.

[0051] Hyperedge convolutional layer: The hyperedge convolutional layer can be expressed as: ; where is the representation of the hypergraph at the th layer, is the feature dimension, , is the diagonal matrix of hyperedge weights, is a parameter that can be learned during the training process, is the bias vector; is the diagonal matrix of the site set .

[0052] Combined with Figure 4 , the structure of the multi-scale hypergraph autoencoder includes an encoder and a decoder; The encoder has 3 layers, namely the hyperedge convolution layer, the activation layer, and the hyperedge convolution layer in sequence, and can be expressed as: ; where , are learnable weight matrices; represents the hyperedge convolution layer.

[0053] The decoder has the same structure as the encoder and can be expressed as: ; where , are learnable weight matrices. The module output is the high-order latent representation learned at the high-order feature level and the reconstructed .

[0054] The loss function of the multi-scale hypergraph autoencoder is the first reconstruction loss function.

[0055] The first reconstruction loss function is constructed as follows: In , the corresponding vectors in are , respectively, then the loss of the multi-scale hypergraph autoencoder module is: .

[0056] The masked autoencoder ( ) consists of a masking mechanism, an encoder, and a decoder; Masking mechanism: Mask in accordance with a set ratio to obtain .

[0057] Encoder: Encode the data input to the encoder to obtain the low-order latent representation , which can be expressed as: ; , is the weight matrix, and is the bias vector.

[0058] Decoder: The decoder structure is consistent with the masker and can be expressed as: ; and are learnable weights, and are bias vectors. The module output is and .

[0059] The total loss function of the masked autoencoder ( ) consists of three parts: the second reconstruction loss function, the alignment loss function, and the spatial consistency loss function.

[0060] Second reconstruction loss function: The corresponding vectors in and are and respectively. Then the second reconstruction loss function is: ; Alignment loss function : Taking as nodes, construct the reconstruction adjacency matrix by. When is 's nearest neighbor , otherwise . The module alignment loss is: ; where is the alignment loss function; and are the Laplacian matrices of and respectively; is the adjacency matrix reconstructed from the low-order latent representation of the gene expression matrix; is the adjacency matrix of the gene expression matrix; represents the F-norm. The alignment loss encourages the module to generate embeddings with a similar graph structure to its input.

[0061] The spatial consistency loss function is: ; ; ; Among them, is the spatial consistency loss function; is the smoothing term; is the low-order latent representation of the gene expression matrix; is the high-order latent representation of the gene expression matrix; and are both intermediate parameters; , are respectively the th site corresponding vectors in the low-order and high-order latent representations of the gene expression matrix respectively; is the total number of sites; is the normalization exponential function.

[0062] The spatial consistency loss function is used to integrate the high-order latent representation into the low-order latent representation, so that the trained masked autoencoder obtains the low-order latent representation integrated with the high-order latent representation.

[0063] The total loss function of the masked autoencoder is: ; Among them, and are constant coefficients.

[0064] Corresponding to the above method, as Figure 3 shown, a no-load data feature extraction system is also provided, including: A no-load data acquisition module 301, which is used to acquire the no-load data of the pathological section, including the initial gene expression matrix and spatial position information; A data denoising module 302, which is used to perform denoising processing on the initial gene expression matrix to obtain the denoised gene expression matrix; A position encoding module 303, which is used to calculate the position encoding for each site according to the spatial position information and splice it with the denoised gene expression matrix to obtain an enhanced gene expression matrix; A low-order latent representation module 304, which is used to encode the enhanced gene expression matrix by using a masked autoencoder to obtain the low-order latent representation and use it as the extracted no-load data feature.

[0065] Among them, during the process of training the masked autoencoder, an alignment loss function is constructed using the reconstructed adjacency matrix of the low-order latent representation of the gene expression matrix; a spatial consistency loss function is constructed using the high-order and low-order latent representations of the gene expression matrix; a second reconstruction loss function is constructed using the reconstructed enhanced gene expression matrix and the enhanced gene expression matrix before reconstruction. Then, the total loss function of the masked autoencoder is obtained from the alignment loss function, the spatial consistency loss function, and the second reconstruction loss function, and further, the masked autoencoder is trained.

[0066] Specifically, the reconstructed enhanced gene expression matrix is decoded from the low-order latent representation output by the masked autoencoder; the high-order latent representation of the gene expression matrix is encoded by the multi-scale hypergraph autoencoder for the denoised gene expression matrix.

[0067] The multi-scale hypergraph autoencoder is used to learn the high-order latent representation of the denoised gene expression matrix at different scales and reconstruct the denoised gene expression matrix. Then, a first reconstruction loss function is constructed to train the multi-scale hypergraph autoencoder.

[0068] It should be noted here that the specific implementation processes of the various modules in the embodiments of the present invention correspond one by one to the steps in the above-mentioned idle data feature extraction method, and their specific implementation processes are the same, so they will not be elaborated here.

[0069] As Figure 5 shown, the embodiments of the present invention provide a method for identifying the spatial domain of spatial transcriptomics, including: S501: Extract the low-order latent representation of the idle data of the pathological section using the steps in the above-mentioned idle data feature extraction method as the idle data feature; S502: Cluster the sites of the idle data feature to obtain the spatial domain identification result.

[0070] As Figure 6 shown, in one or more embodiments, a system for identifying the spatial domain of spatial transcriptomics is also provided, including: A feature extraction module 601, which is used to extract the low-order latent representation of the idle data of the pathological section using the steps in the above-mentioned idle data feature extraction method as the idle data feature; A clustering module 602, which is used to cluster the sites of the idle data feature to obtain the spatial domain identification result.

[0071] The input of the clustering module 602 is , and in (K-means clustering), (model-based clustering), and (Gaussian mixture model), select one clustering method to perform clustering on Perform clustering, and the clustering labels are the recognition results in the spatial domain.

[0072] Model evaluation: Use 、 、 to evaluate the consistency between the clustering results and the true labels. For better understanding and to better demonstrate the model performance, visualizations of the recognition results in the spatial domain and the true labels are provided.

[0073] Downstream tasks: Based on the recognition results in the spatial domain, perform biological function analyses such as gene enrichment analysis and signaling pathway prediction.

[0074] In particular, according to the embodiments of the present application, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, the embodiments of the present application include a computer program product that includes a computer program carried on a computer-readable medium, and the computer program contains program codes for executing the methods shown in the flowcharts. In such an embodiment, the computer program can be downloaded and installed from the network through the communication part, and / or installed from a removable medium. When the computer program is executed by the central processing unit, various functions defined in the device of the present application are executed.

[0075] The present invention is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products of the embodiments of the present invention. It should be understood that each process and / or block in the flowcharts and / or block diagrams, and the combination of processes and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to the processors of general-purpose computers, special-purpose computers, embedded processors, or other programmable data processing devices to generate a machine, such that the instructions executed by the processors of the computer or other programmable data processing devices generate means for implementing the specified functions in Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks.

[0076] The above are only the preferred embodiments of the present invention and are not used to limit the present invention. For those skilled in the art, the present invention can have various changes and modifications. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.

Claims

1. A method for extracting idling data features, characterized in that, Including: Obtaining the idle running data of the pathological section, which includes the initial gene expression matrix and spatial position information; Performing denoising processing on the initial gene expression matrix to obtain the denoised gene expression matrix; Calculating the position encoding for each site according to the spatial position information and splicing it with the denoised gene expression matrix to obtain the enhanced gene expression matrix; Encoding the enhanced gene expression matrix using a masked autoencoder to obtain a low-order latent representation and using it as the extracted feature of the idle running data.

2. The no-load data feature extraction method according to claim 1, characterized in that, During the process of training the masked autoencoder, constructing an alignment loss function using the adjacency matrix reconstructed from the low-order latent representation of the gene expression matrix; constructing a spatial consistency loss function using the high-order latent representation and low-order latent representation of the gene expression matrix; constructing a second reconstruction loss function using the reconstructed enhanced gene expression matrix and the enhanced gene expression matrix before reconstruction, and then obtaining the total loss function of the masked autoencoder from the alignment loss function, spatial consistency loss function, and second reconstruction loss function, and further training the masked autoencoder.

3. The no-load data feature extraction method according to claim 2, wherein The reconstructed enhanced gene expression matrix is obtained by decoding the low-order latent representation output by the masked autoencoder; the high-order latent representation of the gene expression matrix is obtained by encoding the denoised gene expression matrix using a multi-scale hypergraph autoencoder.

4. The no-load data feature extraction method according to claim 3, wherein Using the multi-scale hypergraph autoencoder to learn the high-order latent representation of the denoised gene expression matrix at different scales and reconstructing the denoised gene expression matrix, and then constructing a first reconstruction loss function to train the multi-scale hypergraph autoencoder.

5. The no-load data feature extraction method according to claim 2, wherein The spatial consistency loss function is: ; ; ; Among them, is the spatial consistency loss function; is the smoothing term; is the low-order latent representation of the gene expression matrix; is the high-order latent representation of the gene expression matrix; and are both intermediate parameters; 、 are respectively the th site corresponding vectors in the low-order latent representation and high-order latent representation of the gene expression matrix respectively; is the total number of sites; is the normalized exponential function.

6. The method for extracting no-load data features according to claim 2, wherein The alignment loss function is as follows: ; Among them, is the alignment loss function; , are respectively , 's Laplacian matrices; is the adjacency matrix reconstructed from the low-order latent representation of the gene expression matrix; is the adjacency matrix of the gene expression matrix; represents the Frobenius norm.

7. An idling data feature extraction system, characterized in that, Including: An idle running data acquisition module, which is used to obtain the idle running data of the pathological section, which includes the initial gene expression matrix and spatial position information; A data denoising module, which is used to perform denoising processing on the initial gene expression matrix to obtain the denoised gene expression matrix; A position encoding module, which is used to calculate the position encoding for each site according to the spatial position information and splice it with the denoised gene expression matrix to obtain the enhanced gene expression matrix; A low-order latent representation module, which is used to encode the enhanced gene expression matrix using a masked autoencoder to obtain a low-order latent representation and using it as the extracted feature of the idle running data.

8. A method for identifying spatial domains in spatial transcriptomics, characterized in that, Including: Extracting the low-order latent representation of the idle running data of the pathological section using the steps in the method for extracting features of idle running data described in any one of claims 1-6, and using it as the feature of the idle running data; Clustering the sites of the feature of the idle running data to obtain the spatial domain recognition result.

9. A spatial transcriptome spatial domain recognition system, characterized in that, Including: A feature extraction module, which is used to extract the low-order latent representation of the idle running data of the pathological section using the steps in the method for extracting features of idle running data described in any one of claims 1-6, and using it as the feature of the idle running data; A clustering module, which is used to cluster the sites of the feature of the idle running data to obtain the spatial domain recognition result.

10. An electronic device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, When the processor executes the program, it implements the steps in the method for extracting features of idle running data described in any one of claims 1-6; Or when the processor executes the program, it implements the steps in the method for recognizing the spatial domain of spatial transcriptomics described in claim 8.

Citation Information

Patent Citations

  • Spatial domain identification method based on spatial transcriptomics data feature extraction

    CN116189785A

  • Space transcriptome data processing method and system based on hypergraph

    CN117457081A

  • Spatial domain identification method integrating spatial transcriptome multi-modal information

    CN118016149A

  • Method for carrying out spatial domain division on spatial transcriptomics data

    CN118588159A

  • Histological image-based space gene expression level prediction method, device and equipment

    CN118629498A

Cited By

  • Breast cancer detection method and system based on morphological image and space transcriptome cross-graph collaborative learning

    CN121582228A