Transcriptome annotation method and system based on large language model

By fusing single-cell spatial coordinates with gene expression values into pseudo-image modal data, topological features are extracted and cross-modal alignment is carried out to construct functional semantic space, the problems of extensive topological modeling, low cross-modal alignment accuracy and limited non-modal species annotation in single-cell spatial transcriptome annotation are solved, and high-precision cross-species functional semantic inference is achieved.

CN120356514AActive Publication Date: 2025-07-22INSTITUTE OF CROP SCIENCE CHINESE ACADEMY OF AGRICULTURAL SCIENCES +1
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202510817866.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-18
Publication Date
2025-07-22
Estimated Expiration
2045-06-18

AI Technical Summary

Technical Problem

The prior art has problems such as extensive spatial topology modeling, low cross-modal alignment accuracy, limited non-modal species annotation, and rigid semantic mapping in single-cell spatial transcriptome annotation.

Method used

The single-cell spatial coordinates and gene expression values are fused into pseudo-image modal data, spatial topological features are extracted, and functional semantic space is constructed through cross-modal alignment, the semantic similarity of gene expression embedding is calculated, and semantic mapping is used for large language models.

Benefits of technology

It improves the microenvironment analytical accuracy of single-cell spatial transcriptome annotation, reduces noise interference, breaks through model species dependence, realizes semantic equivalence inference across species functions, and improves annotation accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120356514A_ABST
    Figure CN120356514A_ABST
Patent Text Reader

Abstract

The invention relates to the crossing field of bioinformatics and computational biology, and discloses a transcriptome annotation method and system based on a large language model.The method comprises the steps that single cell space coordinates and gene expression values are fused into pseudo-image modal data, and spatial topological features of the pseudo-image modal data are extracted; performing cross-modal alignment between the spatial topological features and a preset medical database, and analyzing a cell type probability of a cross-modal embedded vector; constructing a function semantic space of the function description text, and projecting the non-mode species into the function semantic space; and calculating semantic similarity between gene expression embedding and homologous genes of reference species, and converting gene expression embedding into a semantic mapping relationship by using a large language model. According to the method, the core problems of extensive spatial topology modeling, low cross-modal alignment precision, limited non-modal species annotation, semantic mapping stiffness and the like in single cell spatial transcriptome annotation are reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a method and system for transcriptome annotation based on a large language model, belonging to the cross - field of bioinformatics and computational biology. Background Art

[0002] Currently, existing methods such as heatmaps or scatter plots only statically associate single - cell spatial coordinates with gene expression values, lacking the ability to dynamically model spatial topological structures. For example, traditional techniques cannot grid discrete cell positions and convert them into image - modality data, resulting in the neglect of key features such as spatial proximity and hole structures, limiting the analytical accuracy of complex tissue microenvironments. The integration of medical databases such as known cell - type marker gene libraries and single - cell spatial features mostly relies on simple linear regression or shallow neural networks, making it difficult to capture the non - linear associations between cross - modality data. Current gene annotation tools such as DAVID and GO enrichment analysis mainly infer the functions of homologous genes based on model species such as humans and mice, and their semantic similarity calculation only relies on the static hierarchical structure of Gene Ontology (GO) terms, without combining the semantic space modeling of the protein function description texts of all species. This leads to functional biases or "no homologous annotation" blind spots when annotating genes of non - model species. Traditional methods calculate similarity through gene sequence alignment or conserved domain matching, but cannot effectively utilize the semantic information of functional description texts, such as the natural language descriptions in UniProtKB. In addition, existing technologies such as BLAST lack a dynamic expansion mechanism when there is no reliable homologous annotation, resulting in low annotation coverage. For the gene function annotation of non - model species, manual alignment or low - throughput experimental verification is mostly used, which is time - consuming and costly. Existing algorithms support gene family clustering, but do not combine the semantic reasoning ability of large language models, making it difficult to achieve intelligent mapping of cross - species functional families.

[0003] Therefore, the existing technologies face core problems such as rough spatial topological modeling, low cross - modality alignment accuracy, limited annotation of non - model species, and rigid semantic mapping in single - cell spatial transcriptome annotation. Summary of the Invention

[0004] The present invention provides a method and system for transcriptome annotation based on a large language model, and its main purpose is to reduce the core problems such as rough spatial topological modeling, low cross - modality alignment accuracy, limited annotation of non - model species, and rigid semantic mapping faced in single - cell spatial transcriptome annotation.

[0005] To achieve the above object, a method for transcriptome annotation based on a large language model provided by the present invention includes: Fusing single - cell spatial coordinates and gene expression values into pseudo - image - modality data, and extracting the spatial topological features of the pseudo - image - modality data; Perform cross-modal alignment between the spatial topological features and a preset medical database to obtain cross-modal embedded vectors after alignment, and analyze the cell type probabilities corresponding to the cross-modal embedded vectors; Collect the functional description texts of all-species proteins in UniProtKB, construct a functional semantic space for the functional description texts, project non-model species into the functional semantic space, and obtain gene expression embeddings in the functional semantic space; Calculate the semantic similarity between the gene expression embeddings and homologous genes of a reference species, and based on the semantic similarity, use a large language model to convert the gene expression embeddings into semantic mapping relationships; Determine the transcriptome annotation results using the cell type probabilities and the semantic mapping relationships.

[0006] Optionally, the fusion of single-cell spatial coordinates and gene expression values into pseudo-image modal data includes: Obtain the coordinate space corresponding to the single-cell spatial coordinates; Grid the coordinate space to obtain a gridded space; Query the grid corresponding to the single-cell spatial coordinates in the gridded space to discretize the single-cell spatial coordinates into pixel points; Fill the gene expression values into the pixel points to obtain pixel values; Determine pseudo-image modal data from the pixel points and the pixel values.

[0007] Optionally, the extraction of the spatial topological features of the pseudo-image modal data includes: Convert the grid space formed by the pixel points in the pseudo-image modal data into a cubical complex; Use the pixel values to construct a filtration model for the cubical complex; Based on the filtration model, determine the filtration values of the pixel points; Sort the pixel points from low to high based on the filtration values to obtain a subcomplex; Calculate the homology groups of the subcomplex at different filtration values, where the homology groups include connectivity, the number of holes, and the number of cavities; Record the birth filtration value when the topological feature is generated and the death filtration value when the topological feature disappears; Construct a persistence diagram of the topological feature with the birth filtration value as the abscissa and the death filtration value as the ordinate; Use the persistence diagram as the spatial topological features of the pseudo-image modal data.

[0008] Optionally, the cross-modal alignment between the spatial topological features and a preset medical database to obtain cross-modal embedded vectors after alignment includes: Convert the spatial topological features into a first feature vector; Convert the cell type information in the medical database into a second feature vector; Use the first feature vector and the second feature vector as a query vector and a key-value pair vector respectively; Calculate the dot product between the query vector and the key vector in the key-value pair vector to obtain an attention score; Normalize the attention score into a weight value through an activation function; Perform weighted summation on the value vectors in the key-value pair vector using the weight value to obtain a weighted feature vector; Perform a non-linear transformation on the weighted feature vector to obtain an aligned cross-modal embedding vector.

[0009] Optionally, analyzing the cell type probability corresponding to the cross-modal embedding vector includes: Collect cross-modal embedding vector samples within a historical period; Use a classifier to identify the cell type probability samples corresponding to the cross-modal embedding vector samples; Use the cell type probability samples as output vectors, and use the single-cell spatial coordinate samples corresponding to the cross-modal embedding vector samples and the cross-modal embedding vector samples as input vectors; Define a radial basis function of the input vector and a Gaussian process regression model between the output vector and the input vector; When optimizing the parameters of the radial basis function using a preset marginal likelihood function, perform parameter fitting on the Gaussian process regression model based on the output vector and the input vector to obtain a well-fitted regression model; Use the well-fitted regression model to output the cell type probability corresponding to the cross-modal embedding vector.

[0010] Optionally, constructing the functional semantic space of the functional description text includes: Perform text embedding on the functional description text using a natural language processing model to obtain a semantic vector; Model the gene-function relationships in UniProtKB as a knowledge graph; Wherein, the knowledge graph includes gene nodes, function nodes, and gene-function relationship edges; After using the semantic vector as the initial vector of the function nodes in the knowledge graph, perform graph embedding learning on the knowledge graph through a graph convolutional network to obtain a functional semantic space.

[0011] Optionally, projecting the non-model species into the functional semantic space to obtain the gene expression embedding of the functional semantic space includes: Projecting the gene expression values of the non-model species into the functional semantic space through an autoencoder to obtain the gene expression embedding of the functional semantic space.

[0012] Optionally, calculating the semantic similarity between the gene expression embedding and the homologous genes of the reference species includes: Obtaining the gene ontology between the gene expression embedding and the homologous genes; Querying the hierarchical structure of the gene ontology; Querying the first functional node and the second functional node corresponding to the gene expression embedding and the homologous genes in the hierarchical structure respectively; Querying the lowest common ancestor of the first functional node and the second functional node in the hierarchical structure; According to the hierarchical structure, calculating the information amount of the lowest common ancestor by using the following formula: ; Wherein, represents the information amount, represents the frequency of occurrence of the lowest common ancestor in the hierarchical structure; Taking the information amount as the semantic similarity.

[0013] Optionally, based on the semantic similarity, using a large language model to convert the gene expression embedding into a semantic mapping relationship includes: When the semantic similarity is less than a preset similarity threshold, it is determined that there is no reliable homologous annotation for the gene expression embedding in the reference species; After determining that there is no reliable homologous annotation for the gene expression embedding in the reference species, using a domain adaptation method and the large language model to map the gene expression embedding into a preset gene-functional family to obtain an updated target gene; Performing gene set expansion on the updated target gene to obtain an expanded gene set; Extracting the semantic mapping relationship from the expanded gene set.

[0014] To solve the above problems, the present invention also provides a transcriptome annotation system based on a large language model, and the system includes: A feature extraction module, configured to fuse single-cell spatial coordinates and gene expression values into pseudo-image modal data, and extract the spatial topological features of the pseudo-image modal data; A probability analysis module, configured to perform cross-modal alignment between the spatial topological features and a preset medical database to obtain cross-modal embedded vectors after alignment, and analyze the cell type probabilities corresponding to the cross-modal embedded vectors; A species projection module, configured to collect functional description texts of all-species proteins in UniProtKB, construct a functional semantic space of the functional description texts, and project non-model species into the functional semantic space to obtain gene expression embeddings of the functional semantic space; A gene conversion module, configured to calculate the semantic similarity between the gene expression embeddings and homologous genes of a reference species, and based on the semantic similarity, use a large language model to convert the gene expression embeddings into semantic mapping relationships; A result determination module, configured to determine a transcriptome annotation result by using the cell type probabilities and the semantic mapping relationships.

[0015] Compared with the problems in the background art, in the embodiments of the present invention, by fusing single-cell spatial coordinates and gene expression values into pseudo-image modal data, generating pseudo-image modal data with grid-like spatial coordinates, and combining subsequent persistent homology to extract topological features, the spatial structure can be dynamically captured, and the microenvironment analysis accuracy can be improved. Further, in the embodiments of the present invention, by performing cross-modal alignment between the spatial topological features and a preset medical database, the medical database and the spatial topological features are weighted and aligned by using an attention mechanism, and the cell type probabilities are modeled by Gaussian process regression subsequently, so as to reduce noise interference and improve the annotation accuracy. Further, in the embodiments of the present invention, by constructing a functional semantic space of the functional description texts and projecting non-model species into the functional semantic space, the dependence on model species is broken through, and cross-species functional semantic equivalence inference is realized. Further, in the embodiments of the present invention, by calculating the semantic similarity between the gene expression embeddings and homologous genes of a reference species, the biological rationality of functional similarity determination is enhanced. Therefore, the transcriptome annotation method and system based on a large language model provided by the embodiments of the present invention can reduce the core problems faced in single-cell spatial transcriptome annotation, such as rough spatial topological modeling, low cross-modal alignment accuracy, limited non-model species annotation, and rigid semantic mapping. Description of the Drawings

[0016] Figure 1 It is a schematic flowchart of a transcriptome annotation method based on a large language model provided by an embodiment of the present invention; Figure 2 It is a schematic module diagram of a system for implementing the transcriptome annotation system based on a large language model provided by an embodiment of the present invention.

[0017] The implementation, functional features, and advantages of the present invention will be further described with reference to the embodiments and the accompanying drawings. Detailed Embodiments

[0018] It should be understood that the specific embodiments described herein are merely used to explain the present invention and are not intended to limit the present invention.

[0019] An embodiment of the present application provides a transcriptome annotation method based on a large language model. The execution subject of the transcriptome annotation method based on the large language model includes, but is not limited to, at least one of electronic devices such as a server, a terminal, etc. that can be configured to execute the method provided by the embodiment of the present application. In other words, the transcriptome annotation method based on the large language model can be executed by software or hardware installed on a terminal device or a server device. The server includes, but is not limited to: a single server, a server cluster, a cloud server, or a cloud server cluster, etc.

[0020] Example 1: Refer to Figure 1 As shown, it is a flowchart of a transcriptome annotation method based on a large language model provided by an embodiment of the present invention. In this embodiment, the transcriptome annotation method based on the large language model includes: S1. Fuse the single-cell spatial coordinates and gene expression values into pseudo-image modal data, and extract the spatial topological features of the pseudo-image modal data.

[0021] In an embodiment of the present invention, by fusing the single-cell spatial coordinates and gene expression values into pseudo-image modal data, generating pseudo-image modal data with a grid-based spatial coordinate, and combining subsequent persistent homology to extract topological features, the spatial structure can be dynamically captured, and the microenvironment analysis accuracy can be improved.

[0022] Among them, the single-cell spatial coordinates refer to the physical position coordinates of a single cell in a tissue or sample, usually obtained through spatial transcriptomics technology, etc., and the gene expression value refers to the expression level of a gene in a specific cell, usually obtained through technologies such as transcriptome sequencing, and can be used to reflect the activity degree of the gene.

[0023] In an embodiment of the present invention, the fusing of the single-cell spatial coordinates and gene expression values into pseudo-image modal data includes: obtaining the coordinate space corresponding to the single-cell spatial coordinates; gridifying the coordinate space to obtain a gridified space; querying the grid corresponding to the single-cell spatial coordinates in the gridified space to discretize the single-cell spatial coordinates into pixel points; filling the gene expression values into the pixel points to obtain pixel values; and determining the pseudo-image modal data from the pixel points and the pixel values.

[0024] Optionally, the details of fusing the single-cell spatial coordinates and gene expression values into pseudo-image modal data are as follows: Collect the spatial coordinate data of single cells, and determine the coordinate range and coordinate system where they are located; Divide the coordinate space into regular grid cells, and discretize the continuous spatial coordinates into grid coordinates; Find the grid cell where each single-cell spatial coordinate is located, and represent the position information of the cell as a pixel point in the grid; Associate the gene expression value with the corresponding pixel point, and use the gene expression value as the pixel value of the pixel point; Integrate all pixel points and their pixel values together to form a data form similar to an image, that is, pseudo-image modal data.

[0025] In one embodiment of the present invention, extracting the spatial topological features of the pseudo-image modal data includes: converting the grid space composed of pixel points in the pseudo-image modal data into a cubical complex; constructing a filtration model of the cubical complex using pixel values; determining the filtration value of the pixel points based on the filtration model; sorting the pixel points from low to high based on the filtration value to obtain a subcomplex; calculating the homology groups of the subcomplex at different filtration values, where the homology groups include connectivity, the number of holes, and the number of cavities; recording the birth filtration value when the topological feature is generated and the death filtration value when it disappears; constructing a persistence diagram of the topological feature with the birth filtration value as the abscissa and the death filtration value as the ordinate; and taking the persistence diagram as the spatial topological feature of the pseudo-image modal data.

[0026] Optionally, the details of extracting the spatial topological features of the pseudo-image modal data are as follows: Regarding the grid space composed of pixel points in the pseudo-image modal data as a cubical complex, each pixel point corresponds to a vertex of a cube, and geometric structures such as edges, faces, and cubes are formed between adjacent pixel points, thereby establishing a topological structure representation of the entire pseudo-image. Define a filtration function, usually a function related to the research objective such as gene expression value, and sort each pixel point in the cubical complex according to the value of the filtration function to form a filtration process. For example, sort according to the gene expression value from low to high, and gradually construct a series of nested subcomplexes. As the filtration process progresses, at different filtration value thresholds, calculate the homology groups of the corresponding subcomplexes, including topological features such as connectivity (0-dimensional homology group), the number of holes (1-dimensional homology group), and the number of cavities (2-dimensional homology group). For example, when the filtration value is small, only a few pixel points may meet the conditions, and the connectivity is low at this time. As the filtration value increases, the number of pixel points that meet the conditions increases, and the connectivity gradually increases, and structures such as holes may also gradually appear. Record the birth and death filtration values of each topological feature, that is, its survival interval, and plot it in the plane coordinate system in the form of points to generate a persistence diagram. The abscissa represents the filtration value at which the topological feature is born, and the ordinate represents the filtration value at which it disappears. The coordinates (b, d) of the point indicate that the topological feature is born at the filtration value b and disappears at the filtration value d, and the persistence is , each point in the persistence diagram represents a topological feature, and its position and density reflect the stability and distribution of different topological structures in the pseudo-image.

[0027] S2. Perform cross-modal alignment between the spatial topological feature and a preset medical database to obtain an aligned cross-modal embedding vector, and analyze the cell type probability corresponding to the cross-modal embedding vector.

[0028] In an embodiment of the present invention, by performing cross-modal alignment between the spatial topological feature and a preset medical database, the attention mechanism is used to weight-align the medical database and the spatial topological feature, and then the cell type probability is modeled by Gaussian process regression, thereby reducing noise interference and improving the annotation accuracy.

[0029] Among them, the medical database refers to a database storing medical-related data, such as a cell type marker gene library, which can be used to provide reference information for comparative analysis.

[0030] In an embodiment of the present invention, the performing cross-modal alignment between the spatial topological feature and a preset medical database to obtain an aligned cross-modal embedding vector includes: converting the spatial topological feature into a first feature vector; converting the cell type information in the medical database into a second feature vector; using the first feature vector and the second feature vector as a query vector and a key-value pair vector respectively; calculating the dot product between the query vector and the key vector in the key-value pair vector to obtain an attention score; normalizing the attention score into a weight value through an activation function; performing weighted summation on the value vector in the key-value pair vector using the weight value to obtain a weighted feature vector; and performing a non-linear transformation on the weighted feature vector to obtain an aligned cross-modal embedding vector.

[0031] Optionally, the details of extracting the spatial topological features of the pseudo-image modality data are as follows: Vectorize the spatial topological features (i.e., the persistence diagram) to obtain the first feature vector. Model the persistence diagram as a graph structure, where the nodes represent the points in the persistence diagram, and the weights of the edges can be defined based on the distance or similarity between the points. Then apply a graph embedding algorithm to learn the embedding vectors of each node, and aggregate the embedding vectors of all nodes to obtain the feature vector of the entire persistence diagram. For the cell type information, one-hot encoding can be used, or word embedding techniques such as bag of words, TF-IDF and other text representation methods can be used, or the cell type description text can be input into a natural language processing model (such as BERT) to generate semantic vectors. Use the weight value to perform a weighted sum on the value vectors in the key-value pair vectors, and the weighted feature vector is obtained as follows: Multiply the value vectors in each key-value pair vector by the corresponding weight value, and then sum all the weighted value vectors to obtain the weighted feature vector. Perform a non-linear transformation on the weighted feature vector to obtain the aligned cross-modal embedding vector as follows: Use a non-linear transformation model such as a multi-layer perceptron (MLP) to perform a non-linear transformation on the weighted feature vector.

[0032] In one embodiment of the present invention, analyzing the cell type probability corresponding to the cross-modal embedding vector includes: collecting cross-modal embedding vector samples within a historical period; using a classifier to identify the cell type probability samples corresponding to the cross-modal embedding vector samples; taking the cell type probability samples as the output vectors, and taking the single-cell spatial coordinate samples corresponding to the cross-modal embedding vector samples and the cross-modal embedding vector samples as the input vectors; defining a radial basis function of the input vectors and a Gaussian process regression model between the output vectors and the input vectors; when optimizing the parameters of the radial basis function using a preset marginal likelihood function, performing parameter fitting on the Gaussian process regression model based on the output vectors and the input vectors to obtain a well-fitted regression model; using the well-fitted regression model to output the cell type probability corresponding to the cross-modal embedding vector.

[0033] Optionally, the details of the cell type probabilities corresponding to the cross-modal embedding vectors described above are as follows: Collect cross-modal embedding vector samples within a certain past time range. These samples will be used to train and establish a regression model to provide a data basis for subsequent cell type probability prediction. Use a classifier (such as a support vector machine, random forest, neural network, etc.) to classify the cross-modal embedding vector samples to obtain corresponding cell type probability samples. The classifier can predict the probability of each sample belonging to different cell types by learning the mapping relationship between the cross-modal embedding vectors and the cell types. Take the cell type probability samples as the output vectors, and take the single-cell spatial coordinate samples and cross-modal embedding vector samples as the input vectors to construct the input-output pairs of the regression model to provide a data format for subsequent regression analysis. Define the radial basis function (RBF) of the input vectors. RBF is a commonly used kernel function for measuring the similarity between various input vectors. Establish a Gaussian process regression model between the output vector (cell type probability) and the input vectors (single-cell spatial coordinates and cross-modal embedding vectors). The Gaussian process regression model can model and predict the input-output relationship and output the probability distribution of the predicted values. Use the marginal likelihood function to optimize the parameters of the radial basis function. The marginal likelihood function is a function for measuring the degree of fit of the model to the data. By optimizing the marginal likelihood function, the optimal parameters of the radial basis function can be found. At the same time, fit the parameters of the Gaussian process regression model to obtain a well-fitted regression model so that the model can better predict the cell type probability. Use the well-fitted regression model to predict new cross-modal embedding vectors and output the corresponding cell type probabilities.

[0034] S3. Collect the functional description texts of all-species proteins in UniProtKB, construct the functional semantic space of the functional description texts, and project non-model species into the functional semantic space to obtain the gene expression embedding of the functional semantic space.

[0035] In the embodiment of the present invention, by constructing the functional semantic space of the functional description texts and projecting non-model species into the functional semantic space, the dependence on model species is broken through, and cross-species functional semantic equivalence inference is realized.

[0036] Among them, the UniProtKB refers to a database that provides protein sequence and function information, including a large amount of information such as functional descriptions of known proteins, gene-protein relationships, etc. The functional description text refers to the text description of the function of a gene or protein, such as Gene Ontology (GO) annotation, domain information, etc., which can be used to understand the biological function of a gene or protein.

[0037] In one embodiment of the present invention, constructing the functional semantic space of the functional description text includes: performing text embedding on the functional description text using a natural language processing model to obtain semantic vectors; modeling the gene-function relationships in UniProtKB as a knowledge graph; wherein the knowledge graph includes gene nodes, function nodes, and gene-function relationship edges; after using the semantic vectors as the initial vectors of the function nodes in the knowledge graph, performing graph embedding learning on the knowledge graph through a graph convolutional network to obtain the functional semantic space.

[0038] Among them, the functional semantic space is constructed based on the gene function description text, and each dimension corresponds to a gene function.

[0039] Optionally, the details of constructing the functional semantic space of the functional description text are as follows: using a natural language processing model (such as BERT, Word2Vec, GloVe, etc.) to perform embedding processing on the functional description text to obtain semantic vectors, modeling the gene-function relationships in UniProtKB as a knowledge graph, the knowledge graph includes gene nodes, function nodes, and gene-function relationship edges, the gene nodes represent different genes, the function nodes represent different functions, and the gene-function relationship edges represent the associations between genes and functions. Using the semantic vectors as the initial vectors of the function nodes, and using a graph convolutional network (GCN) to perform graph embedding learning on the knowledge graph. GCN updates the embedding vectors of the nodes by aggregating the feature information of the nodes and their neighbor nodes, and finally obtains the functional semantic space. The distance between the nodes reflects the similarity between gene functions. The embedding vectors of all function nodes obtained through the above graph embedding learning process together constitute the functional semantic space. In this space, the vector representation of each function node not only contains the semantic information of the original functional description, but also integrates the structural information in the gene-function relationship graph, enabling the use of vector operations to analyze and compare the similarities between different gene functions.

[0040] Furthermore, in the embodiment of the present invention, the non-model species refer to those species that have not been widely selected by the research community for in-depth research. These species usually lack the easily studied characteristics of model organisms, such as being unable to grow in the laboratory, having a long life cycle, a low reproduction rate, or a complex genetic background, etc.

[0041] In one embodiment of the present invention, projecting the non-model species into the functional semantic space to obtain the gene expression embedding of the functional semantic space includes: projecting the gene expression values of the non-model species into the functional semantic space through an autoencoder to obtain the gene expression embedding of the functional semantic space.

[0042] Among them, the gene expression embedding is the representation of gene expression values in the functional semantic space, inheriting the characteristics of the functional semantic space. Each dimension in the functional semantic space is related to gene functions. Therefore, the gene expression embedding also contains the connection between gene expression and functions.

[0043] Optionally, the details of projecting the non-model species into the functional semantic space are as follows: An autoencoder is a neural network model composed of an encoder and a decoder. The encoder maps gene expression values to low-dimensional embedding vectors in the functional semantic space, and the decoder restores the low-dimensional embedding vectors to the original gene expression values. By training the autoencoder, the decoder can restore the input data as accurately as possible, thereby obtaining gene expression embeddings. The gene expression embedding vectors not only contain the expression patterns of genes but also indirectly reflect the functional characteristics of genes through the mapping of the functional semantic space.

[0044] S4. Calculate the semantic similarity between the gene expression embedding and the homologous genes of the reference species, and based on the semantic similarity, use a large language model to convert the gene expression embedding into a semantic mapping relationship.

[0045] In an embodiment of the present invention, by calculating the semantic similarity between the gene expression embedding and the homologous genes of the reference species, the biological rationality of functional similarity determination is enhanced.

[0046] Among them, the homologous genes of the reference species refer to the genes in the reference species that have similar sequences and functions to the genes corresponding to the gene expression embedding. They are differentiated from a common ancestor gene through the evolutionary process. Therefore, they perform similar biological functions in different species. When calculating the semantic similarity between the gene expression embedding and the homologous genes of the reference species, the homologous genes of the reference species provide known functional annotation information for the target genes corresponding to the gene expression embedding.

[0047] In an embodiment of the present invention, calculating the semantic similarity between the gene expression embedding and the homologous genes of the reference species includes: obtaining the gene ontology between the gene expression embedding and the homologous genes; querying the hierarchical structure of the gene ontology; respectively querying the first functional node and the second functional node corresponding to the gene expression embedding and the homologous genes in the hierarchical structure; querying the lowest common ancestor of the first functional node and the second functional node in the hierarchical structure; according to the hierarchical structure, using the following formula to calculate the information content of the lowest common ancestor: ; Among them, represents the information content, represents the frequency of occurrence of the lowest common ancestor in the hierarchical structure; Taking the information content as the semantic similarity.

[0048] Optionally, the details of calculating the semantic similarity between the gene expression embedding and the homologous gene of the reference species are as follows: The Gene Ontology (GO) is a dynamic and hierarchical knowledge base that provides standard descriptions of gene functions. First, obtain the hierarchical structure of GO, which is like a family tree. Each node on the tree represents a gene function, and the branches represent the hierarchical relationship between functions. At the same time, it is also necessary to obtain the annotation information of the gene expression embedding and the reference gene in GO, that is, their respective corresponding functional nodes. For a pair of functional nodes of the target gene and the reference gene, find their lowest common ancestor (LCA, that is, the smallest common parent node of the two nodes) in the GO hierarchical structure. The more frequently a functional node appears in gene annotation, the less information it contains. On the contrary, the less frequently it appears, the more information it contains. That is, the information content of the rare lowest common ancestor is greater. Take the information content of this LCA node as the similarity between the first functional node and the second functional node. For example, if the functional node of the gene expression embedding is "cellular respiration" and the functional node of the reference gene is "mitochondrial respiration", their LCA may be "respiration process", then the information content is the information content of the "respiration process" node. The following is an example: Suppose the gene expression embedding vector obtained after the autoencoder processing of gene A is E, and the homologous gene B of the reference species has been annotated to the functional node F2 in the Gene Ontology (GO). The gene expression embedding vector E of gene A is analyzed to be related to the functional node F1 "cellular respiration", and the corresponding functional node of the homologous gene B is F2 "mitochondrial respiration". Among them, the process of analyzing that the gene expression embedding vector E of gene A is related to the functional node F1 "cellular respiration" is: calculate the similarity between the gene expression embedding vector E and the embedding vector EF1 of the functional node F1, that is, use methods such as cosine similarity to calculate the similarity between the gene expression embedding vector E and the embedding vector EF1 of the functional node F1. If the similarity is higher than the preset threshold, it is considered that the gene expression embedding vector E is related to the functional node F1.

[0049] Furthermore, in the embodiment of the present invention, based on the semantic similarity, a large language model is used to convert the gene expression embedding into a semantic mapping relationship, solve the "annotation blind spot", and improve the annotation coverage rate of non-model species.

[0050] In one embodiment of the present invention, the conversion of the gene expression embedding into a semantic mapping relationship based on the semantic similarity includes: when the semantic similarity is less than a preset similarity threshold, it is determined that there is no reliable homologous annotation for the gene expression embedding in the reference species; after determining that there is no reliable homologous annotation for the gene expression embedding in the reference species, a domain adaptation method and the large language model are used to map the gene expression embedding into a preset gene - function family to obtain an updated target gene; gene set expansion is performed on the updated target gene to obtain an expanded gene set; and semantic mapping relationships are extracted from the expanded gene set.

[0051] Optionally, the details of converting the gene expression embedding into a semantic mapping relationship based on the semantic similarity are as follows: when the similarity is lower than the preset threshold of 0.7, for gene expression embeddings without reliable homologous annotations, a domain adaptation method is used to map them into a preset gene - function family. The domain adaptation method can adopt strategies such as adversarial training and feature alignment to make the feature distributions of the source domain (reference species) and the target domain (non - model species) close. At the same time, the large language model is used to provide richer semantic information for the gene expression embedding. For example, the gene expression embedding and the functional description text of the reference gene are input into the large language model together. Through the generation ability of the model, possible functional descriptions are generated for the gene expression embedding, or through the embedding layer of the model, the gene expression embedding is transformed into the same functional semantic space as the reference gene. Gene set expansion for the updated target gene can be achieved through methods such as functional similarity search and literature mining to find other genes with similar functions or related relationships to the genes corresponding to the gene expression embedding, expanding the coverage of the gene set and enhancing the comprehensiveness of functional annotation. Extracting semantic mapping relationships from the expanded gene set means determining the semantic associations between each gene and its corresponding function.

[0052] S5. Determine the transcriptome annotation result by using the cell type probability and the semantic mapping relationship.

[0053] It should be noted that the final annotation result includes the annotation of the cell type corresponding to the cell type probability and the updated semantic mapping relationship. In spatial transcriptome analysis, the cell type probability distribution provides the probability values of belonging to different cell types for each spatial position. After analysis and interpretation, these probability values can be converted into specific cell type annotations. The semantic mapping relationship can more accurately describe the function of genes, especially in cross - species annotation, providing a more reliable basis for gene function annotation.

[0054] Compared with the problems described in the background art, in the embodiments of the present invention, the single-cell spatial coordinates and gene expression values are fused into pseudo-image modal data, the pseudo-image modal data is generated with gridded spatial coordinates, and the subsequent persistent homology is combined to extract topological features, so as to dynamically capture the spatial structure and improve the accuracy of microenvironment analysis. Further, in the embodiments of the present invention, cross-modal alignment is performed between the spatial topological features and a preset medical database, and an attention mechanism is used to weight-align the medical database and the spatial topological features, and then the cell type probability is modeled by Gaussian process regression, so as to reduce noise interference and improve the annotation accuracy. Further, in the embodiments of the present invention, the functional semantic space of the functional description text is constructed and non-model species are projected into the functional semantic space to break through the dependence on model species and realize the inference of cross-species functional semantic equivalence. Further, in the embodiments of the present invention, the semantic similarity between the gene expression embedding and the homologous genes of the reference species is calculated to enhance the biological rationality of functional similarity determination. Therefore, the transcriptome annotation method and system based on the large language model provided by the embodiments of the present invention can reduce the core problems faced in single-cell spatial transcriptome annotation, such as rough spatial topological modeling, low cross-modal alignment accuracy, limited annotation of non-model species, and rigid semantic mapping.

[0055] Embodiment 2: As Figure 2 shown, it is a functional module diagram of a transcriptome annotation system based on a large language model of the present invention.

[0056] The transcriptome annotation system 200 based on the large language model of the present invention can be installed in an electronic device. According to the functions achieved, the transcriptome annotation system based on the large language model may include a feature extraction module 201, a probability analysis module 202, a species projection module 203, a gene conversion module 204, and a result determination module 205. The modules of the present invention may also be referred to as units, which refer to a series of computer program segments that can be executed by a processor of an electronic device and can complete fixed functions, and are stored in the memory of the electronic device.

[0057] In the embodiments of the present invention, the functions of each module / unit are as follows: The feature extraction module 201 is used to fuse single-cell spatial coordinates and gene expression values into pseudo-image modal data and extract the spatial topological features of the pseudo-image modal data; The probability analysis module 202 is used to perform cross-modal alignment between the spatial topological features and a preset medical database to obtain an aligned cross-modal embedding vector and analyze the cell type probability corresponding to the cross-modal embedding vector; The species projection module 203 is used to collect the functional description texts of proteins of all species in UniProtKB, construct the functional semantic space of the functional description texts, project non-model species into the functional semantic space, and obtain the gene expression embedding of the functional semantic space; The gene conversion module 204 is used to calculate the semantic similarity between the gene expression embedding and the homologous genes of the reference species, and based on the semantic similarity, use a large language model to convert the gene expression embedding into a semantic mapping relationship; The result determination module 205 is used to determine the transcriptome annotation result by using the cell type probability and the semantic mapping relationship.

[0058] Specifically, each module in the transcriptome annotation system 200 based on a large language model in the embodiments of the present invention adopts the same technical means as those in the above-mentioned Figure 1 transcriptome annotation method based on a large language model, and can produce the same technical effects, which will not be elaborated here.

[0059] For those skilled in the art, it is obvious that the present invention is not limited to the details of the above exemplary embodiments, and without departing from the spirit or basic characteristics of the present invention, the present invention can be implemented in other specific forms.

[0060] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit them. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that the technical solutions of the present invention can be modified or equivalently replaced without departing from the spirit and scope of the technical solutions of the present invention.

Claims

1. A transcriptome annotation method based on large language models, characterized in that, The method includes: Fusing the single-cell spatial coordinates and gene expression values into pseudo-image modal data, and extracting the spatial topological features of the pseudo-image modal data; Performing cross-modal alignment between the spatial topological features and a preset medical database to obtain an aligned cross-modal embedding vector, and analyzing the cell type probability corresponding to the cross-modal embedding vector; Collecting the functional description texts of all-species proteins in UniProtKB, constructing a functional semantic space for the functional description texts, and projecting non-model species into the functional semantic space to obtain gene expression embeddings in the functional semantic space; Calculating the semantic similarity between the gene expression embeddings and homologous genes of a reference species, and based on the semantic similarity, using a large language model to convert the gene expression embeddings into semantic mapping relationships; Determining the transcriptome annotation results using the cell type probability and the semantic mapping relationships.

2. The transcriptome annotation method based on a large language model according to claim 1, wherein, The fusing of the single-cell spatial coordinates and gene expression values into pseudo-image modal data includes: Obtaining the coordinate space corresponding to the single-cell spatial coordinates; Gridifying the coordinate space to obtain a gridified space; Querying the grid corresponding to the single-cell spatial coordinates in the gridified space to discretize the single-cell spatial coordinates into pixel points; Filling the gene expression values into the pixel points to obtain pixel values; Determining the pseudo-image modal data from the pixel points and the pixel values.

3. The transcriptome annotation method based on a large language model according to claim 1, wherein The extracting of the spatial topological features of the pseudo-image modal data includes: Converting the grid space formed by the pixel points in the pseudo-image modal data into a cubical complex; Using the pixel values to construct a filtration model for the cubical complex; Based on the filtration model, determining the filtration values of the pixel points; Sorting the pixel points from low to high based on the filtration values to obtain a subcomplex; Calculating the homology groups of the subcomplex at different filtration values, where the homology groups include connectivity, the number of holes, and the number of cavities; Recording the birth filtration value when the topological feature is generated and the death filtration value when the topological feature disappears; Constructing a persistence diagram of the topological feature with the birth filtration value as the abscissa and the death filtration value as the ordinate; Taking the persistence diagram as the spatial topological feature of the pseudo-image modal data.

4. The transcriptome annotation method based on a large language model according to claim 1, wherein The performing of cross-modal alignment between the spatial topological features and a preset medical database to obtain an aligned cross-modal embedding vector includes: Converting the spatial topological features into a first feature vector; Converting the cell type information in the medical database into a second feature vector; Using the first feature vector and the second feature vector as the query vector and the key-value pair vector respectively; Calculating the dot product between the query vector and the key vectors in the key-value pair vector to obtain attention scores; Normalizing the attention scores into weight values through an activation function; Using the weight values to perform weighted summation on the value vectors in the key-value pair vector to obtain a weighted feature vector; Performing a non-linear transformation on the weighted feature vector to obtain an aligned cross-modal embedding vector.

5. The transcriptome annotation method based on a large language model according to claim 1, wherein, The analyzing of the cell type probability corresponding to the cross-modal embedding vector includes: Collect cross-modal embedding vector samples within a historical time period; Use a classifier to identify the cell type probability samples corresponding to the cross-modal embedding vector samples; Take the cell type probability samples as output vectors, and take the single-cell spatial coordinate samples corresponding to the cross-modal embedding vector samples and the cross-modal embedding vector samples as input vectors; Define the radial basis function of the input vector and the Gaussian process regression model between the output vector and the input vector; When optimizing the parameters of the radial basis function using a preset marginal likelihood function, perform parameter fitting on the Gaussian process regression model based on the output vector and the input vector to obtain a well-fitted regression model; Use the well-fitted regression model to output the cell type probability corresponding to the cross-modal embedding vector.

6. The transcriptome annotation method based on a large language model according to claim 1, wherein, Constructing the functional semantic space of the functional description text includes: Using a natural language processing model to perform text embedding on the functional description text to obtain semantic vectors; Model the gene-function relationships in UniProtKB as a knowledge graph; Wherein, the knowledge graph includes gene nodes, function nodes, and gene-function relationship edges; After taking the semantic vectors as the initial vectors of the function nodes in the knowledge graph, perform graph embedding learning on the knowledge graph through a graph convolutional network to obtain a functional semantic space.

7. The transcriptome annotation method based on a large language model according to claim 1, characterized in that Projecting the non-model species into the functional semantic space to obtain the gene expression embedding of the functional semantic space includes: Projecting the gene expression values of the non-model species into the functional semantic space through an autoencoder to obtain the gene expression embedding of the functional semantic space.

8. The transcriptome annotation method based on a large language model according to claim 1, wherein Calculating the semantic similarity between the gene expression embedding and the homologous genes of the reference species includes: Obtain the gene ontology between the gene expression embedding and the homologous genes; Query the hierarchical structure of the gene ontology; Query the first function node and the second function node corresponding to the gene expression embedding and the homologous genes in the hierarchical structure respectively; Query the lowest common ancestor of the first function node and the second function node in the hierarchical structure; According to the hierarchical structure, use the following formula to calculate the information content of the lowest common ancestor: ; Among them, represents the amount of information, represents the frequency at which the lowest common ancestor appears in the hierarchy; Take the information content as the semantic similarity.

9. The transcriptome annotation method based on a large language model according to claim 1, wherein Based on the semantic similarity, use a large language model to convert the gene expression embedding into a semantic mapping relationship, including: When the semantic similarity is less than a preset similarity threshold, determine that there is no reliable homologous annotation for the gene expression embedding in the reference species; After determining that there is no reliable homologous annotation for the gene expression embedding in the reference species, use a domain adaptation method and the large language model to map the gene expression embedding into a preset gene-function family to obtain an updated target gene; Perform gene set expansion on the updated target gene to obtain an expanded gene set; Extract semantic mapping relationships from the expanded gene set.

10. A transcriptome annotation system based on a large language model, wherein the system applies a transcriptome annotation method based on a large language model as described in any one of claims 1-9, characterized in that The system includes: A feature extraction module for fusing single-cell spatial coordinates and gene expression values into pseudo-image modal data and extracting the spatial topological features of the pseudo-image modal data; A probability analysis module for performing cross-modal alignment between the spatial topological features and a preset medical database to obtain cross-modal embedded vectors after alignment, and analyzing the cell type probabilities corresponding to the cross-modal embedded vectors; A species projection module for collecting functional description texts of all-species proteins in UniProtKB, constructing a functional semantic space of the functional description texts, and projecting non-model species into the functional semantic space to obtain gene expression embeddings of the functional semantic space; A gene conversion module for calculating the semantic similarity between the gene expression embeddings and homologous genes of a reference species, and based on the semantic similarity, using a large language model to convert the gene expression embeddings into semantic mapping relationships; A result determination module for determining transcriptome annotation results using the cell type probabilities and the semantic mapping relationships.

Citation Information

Patent Citations

  • Cross-species single cell annotation method

    CN118298926A

  • Annotation model training method, cell type annotation method and related equipment

    CN119108025A

  • Single-cell transcriptome cell annotation method and system fused with large language model

    CN119601094A

  • Methods and systems for identifying gene regulatory elements and altering gene regulation and expression

    WO2024178321A1

  • Graph neural network construction method and apparatus, and electronic device and storage medium

    WO2025007301A1