Space transcriptome data analysis method and device based on PDX model and storage medium

By using coordinate identifiers and molecular identifiers for preliminary classification and ratio calculation of sequencing sequences in PDX models, the problem of inaccurate gene expression assignment was solved, achieving more accurate spatial transcriptome data analysis and higher sequencing sequence utilization.

CN120808886AActive Publication Date: 2025-10-17HANGZHOU LC BIOTECH
View PDF 9 Cites 0 Cited by

Patent Information

Application Number
CN202511301841.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-09-12
Publication Date
2025-10-17
Estimated Expiration
2045-09-12

AI Technical Summary

Technical Problem

When using PDX models for spatial transcriptome data analysis, existing technologies have problems with inaccurate gene expression distribution and expression loss, especially due to the inaccurate gene expression distribution caused by homologous genes between humans and mice.

Method used

By extracting the coordinate identifiers and molecular identifiers from the sequencing sequences, they were preliminarily classified into the first sequence, second sequence, uncertain sequence and discarded sequence, the first ratio of each bin was calculated, and the bins were classified into the first bin and the second bin according to the ratio. The combined analysis obtained an accurate spatial gene expression profile.

Benefits of technology

The accuracy of spatial transcriptome gene expression distribution and the utilization rate of sequencing sequences were improved, and accurate classification and combined analysis of homologous genes were achieved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120808886A_ABST
    Figure CN120808886A_ABST
Patent Text Reader

Abstract

The invention provides a space transcriptome data analysis method and device based on a PDX model and a storage medium, and relates to the technical field of biological information analysis. Comprising the following steps: extracting a coordinate identifier, a molecular identifier and a sample transcript sequence in a sequencing sequence; analyzing a sample transcript sequence, and preliminarily classifying a sequencing sequence into a first sequence, a second sequence, an uncertain sequence and a discarded sequence; calculating a first ratio under each sub-box; classifying the sub-boxes into a first sub-box and a second sub-box according to the first ratio; re-classifying the uncertain sequences in the first sub-box into a first sequence, and re-classifying the uncertain sequences in the second sub-box into a second sequence; and S104, respectively merging and analyzing all the first sequences and the second sequences to obtain a first space gene expression profile and a second space gene expression profile. According to the method provided by the invention, the distribution accuracy of the gene expression quantity of the space transcriptome and the utilization rate of the sequencing sequence are improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of bio-information analysis, in particular to a spatial transcriptome data analysis method based on a PDX model, a device and a storage medium. BACKGROUND

[0002] The birth of spatial transcriptome technology effectively makes up for the defects of traditional transcriptome and single-cell transcriptome in losing spatial information, enabling researchers to study the spatial characteristics of gene expression while preserving the spatial structure of the tissue. At present, a variety of spatial transcriptome technologies have emerged, such as 10X Visium and Stereo-seq. These technologies achieve spatial localization of gene expression through different ways, such as using oligonucleotide arrays with spatial barcodes to capture mRNA. However, they differ in resolution, for example, the minimum distance between the center points of the minimum spatial capture units of 10X Visium is 100 μm, which is much larger than the size of a single cell, while the minimum distance between the minimum spatial capture units of Stereo-seq is 0.5 μm, which is much smaller than the size of a single cell, so it can be called high-definition spatial transcriptome.

[0003] Patient-Derived Xenograft (PDX) model is a tumor research method that transplants tumor tissue or primary cells from patients directly into immunodeficient mice, and uses the microenvironment provided by mice to make the tumor grow and pass on. Since this step does not require in vitro culture, it avoids genetic and phenotypic changes that occur when cell lines adapt to the culture environment in vitro. In terms of data analysis, although PDX models can generate a large amount of tumor-related data, and when analyzing, additional labels can be added to the genes to construct a reference genome that can distinguish between two species, but due to the presence of homologous genes in the genomes of humans and mice, it can lead to inaccurate gene expression allocation, or loss of expression due to multiple alignments. SUMMARY

[0004] The present application provides a spatial transcriptome data analysis method based on a PDX model, a device and a storage medium to at least solve the above technical problems existing in the prior art.

[0005] According to a first aspect of the present application, a spatial transcriptome data analysis method based on a PDX model is provided, comprising the following steps: S100, extracting a coordinate identifier for recording coordinates, a molecular identifier for distinguishing different transcripts, and a captured sample transcript sequence in the sequencing sequence; S101, analyzing the sample transcript sequence, and preliminarily classifying the sequencing sequence into a first sequence, a second sequence, an uncertain sequence, and a discarded sequence; S102, analyzing the coordinate identifier, calculating a first ratio under each bin, the first ratio being a ratio of a number of sequencing sequences classified as the first sequence and a total number of sequencing sequences classified as the first sequence and the second sequence; S103, according to the first ratio, classifying the bins as first bins and second bins; reclassifying sequencing sequences classified as the uncertain sequence in the first bins as the first sequence, and reclassifying sequencing sequences classified as the uncertain sequence in the second bins as the second sequence; S104, combining and analyzing all the first sequences to obtain a first spatial gene expression profile; combining and analyzing all the second sequences to obtain a second spatial gene expression profile.

[0006] In some embodiments of the first aspect of the present application, in S101, the method for analyzing the sample transcript sequence and initially classifying the sequencing sequences into the first sequence, the second sequence, the uncertain sequence and the discarded sequence is as follows: For each sequencing sequence, the sample transcript sequence therein is used to align with the first reference genome and the second reference genome, and according to the alignment result, the sequencing sequence is classified into the following cases: If the sequencing sequence aligns to the first reference genome but does not align to the second reference genome, the sequencing sequence is classified as the first sequence; If the sequencing sequence does not align to the first reference genome but aligns to the second reference genome, the sequencing sequence is classified as the second sequence; If the sequencing sequence aligns to both the first reference genome and the second reference genome, the sequencing sequence is classified as the uncertain sequence; If the sequencing sequence does not align to both the first reference genome and the second reference genome, the sequencing sequence is classified as the discarded sequence.

[0007] In some embodiments of the first aspect of the present application, in S102, the method for analyzing the coordinate identifier and calculating the first ratio under each bin is as follows: S1021, according to the coordinate identifier, attaching a coordinate to the sequencing sequence; S1022, in combination with the coordinate, obtaining the bin corresponding to the sequencing sequence; S1023, in units of bins, counting the number of sequencing sequences classified as the first sequence under each bin, marked as the first number; and the number of sequencing sequences classified as the second sequence, marked as the second number; S1024, dividing the first number by the sum of the first number and the second number to obtain the first ratio under the bin.

[0008] In some embodiments of the first aspect of the present application, in S102, under each bin, the sequencing sequences are de-duplicated.

[0009] In some embodiments of the first aspect of the application, the deduplication method of the sequencing sequences is as follows: When the coordinate identifier, the molecular identifier, and the sample transcript sequence of a plurality of sequencing sequences are all duplicated, it is determined that the plurality of sequencing sequences are duplicated, and when the number of sequencing sequences is counted in S1023, the plurality of sequencing sequences are only counted as one.

[0010] In some embodiments of the first aspect of the application, in S103, the method of classifying the bins into first bins and second bins according to the first ratio is as follows: The first ratio is compared with a set threshold value, if the first ratio is greater than or equal to the threshold value, the bin is classified as a first bin; if the first ratio is less than the threshold value, the bin is classified as a second bin.

[0011] In some embodiments of the first aspect of the application, the threshold value is 0.5.

[0012] In some embodiments of the first aspect of the application, in S104, after the genes of the first spatial gene expression profile and the second spatial gene expression profile are respectively labeled with species tags, the bins are combined into a comprehensive spatial gene expression profile.

[0013] According to the second aspect of the application, a spatial transcriptome data analysis system based on a PDX model is provided, comprising: The extraction module extracts the coordinate identifier for recording coordinates, the molecular identifier for distinguishing different transcripts, and the captured sample transcript sequence in the sequencing sequence; The preliminary classification module analyzes the sample transcript sequence and classifies the sequencing sequence into a first sequence, a second sequence, an uncertain sequence, and a discarded sequence; The proportion calculation module analyzes the coordinate identifier and calculates a first ratio under each bin, wherein the first ratio is the ratio of the number of sequencing sequences classified as the first sequence to the total number of sequencing sequences classified as the first sequence and the second sequence; The final classification module classifies the bins into first bins and second bins according to the first ratio; reclassifies the sequencing sequences classified as the uncertain sequence in the first bin as the first sequence, and reclassifies the sequencing sequences classified as the uncertain sequence in the second bin as the second sequence; The expression profile generation module combines and analyzes all the first sequences to obtain a first spatial gene expression profile, and combines and analyzes all the second sequences to obtain a second spatial gene expression profile.

[0014] In some embodiments of the second aspect of the application, the method of classifying the sequencing sequences into a first sequence, a second sequence, an uncertain sequence, and a discarded sequence is as follows: For each sequencing sequence, the sample transcript sequence therein is used to align the first reference genome and the second reference genome, and according to the alignment result, the sequencing sequence is classified into the following cases: If the sequencing sequence is aligned to the first reference genome but not aligned to the second reference genome, the sequencing sequence is classified as a first sequence; If the sequencing sequence is not aligned to the first reference genome but aligned to the second reference genome, the sequencing sequence is classified as a second sequence; If the sequencing sequence is aligned to both the first reference genome and the second reference genome, the sequencing sequence is classified as an uncertain sequence; If the sequencing sequence is not aligned to both the first reference genome and the second reference genome, the sequencing sequence is classified as a discarded sequence.

[0015] In some embodiments of the second aspect of the present application, the method for calculating the first ratio under each bin is as follows: According to the coordinate identifier, a coordinate is attached to the sequencing sequence; In combination with the coordinate, a bin corresponding to the sequencing sequence is obtained; In units of bins, the number of sequencing sequences classified as the first sequence under each bin is counted and marked as a first number; and the number of sequencing sequences classified as the second sequence is counted and marked as a second number; The first ratio under the bin is obtained by dividing the first number by the sum of the first number and the second number.

[0016] In some embodiments of the second aspect of the present application, in each bin, the sequencing sequences are de-duplicated.

[0017] In some embodiments of the second aspect of the present application, the de-duplication method of the sequencing sequences is as follows: When the coordinate identifier, the molecular identifier and the sample transcript sequence of a plurality of sequencing sequences are all repeated, it is determined that the plurality of sequencing sequences are repeated, and in the counting of the number of sequencing sequences, the plurality of sequencing sequences are only counted as one.

[0018] In some embodiments of the second aspect of the present application, according to the first ratio, the method for classifying the bins into first bins and second bins is as follows: The first ratio is compared with a set threshold value, if the first ratio is greater than or equal to the threshold value, the bin is classified as a first bin; if the first ratio is less than the threshold value, the bin is classified as a second bin.

[0019] According to a third aspect of the present application, an electronic device is provided, comprising: at least one processor; and a memory connected in communication with the at least one processor; wherein, The memory stores instructions executable by the at least one processor, the instructions being executed by the at least one processor to enable the at least one processor to perform the method described in the present application.

[0020] According to a fourth aspect of the present application, a non-transitory computer readable storage medium storing computer instructions is provided, the computer instructions being used to cause the computer to perform the method described in the present application.

[0021] Compared with the prior art, the present application has the following beneficial effects: By classifying all sequencing sequences and calculating the first ratio of different transcripts on each bin, the bins are further classified, and then the transcripts of homologous genes are classified again according to the classification results of the bins, so as to realize accurate classification of the sequencing sequences; finally, according to the classification results, the analysis is combined to obtain the corresponding spatial gene expression profile. Compared with the prior art, the method provided in the present application improves the accuracy of spatial transcriptome gene expression distribution and the utilization rate of sequencing sequences.

[0022] It should be understood that the content described in this part is not intended to identify key or important features of the embodiments of the present application, nor is it used to limit the scope of the present application. Other features of the present application will become apparent through the following description. BRIEF DESCRIPTION OF DRAWINGS

[0023] The above and other objects, features and advantages of the example embodiments of the present application will be more apparent from the following detailed description taken in conjunction with the accompanying drawings, in which: In the drawings, the same or corresponding reference numerals indicate the same or corresponding parts.

[0024] Figure 1 An implementation flowchart of the embodiment one of the present application is shown.

[0025] Figure 2 A first ratio calculation method flowchart of the embodiment one of the present application is shown.

[0026] Figure 3 A system structure schematic diagram of the embodiment two of the present application is shown.

[0027] Figure 4 A composition structure schematic diagram of an electronic device of the embodiment three of the present application is shown. DETAILED DESCRIPTION

[0028] In order to make the purposes, characteristics and advantages of the present application more obvious and easy to understand, the technical solutions in the embodiments of the present application will be described clearly and completely in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, but not all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor are within the scope of protection of the present application.

[0029] Embodiment one: spatial transcriptome data analysis method based on PDX model The present embodiment one provides a spatial transcriptome data analysis method based on PDX model, please refer to Figure 1 , comprising the following steps: S100, extracting coordinate identifiers for recording coordinates, molecular identifiers for distinguishing different transcripts and captured sample transcript sequences in the sequencing sequences.

[0030] Stereo-seq technology: In the Stereo-seq spatial transcriptome technology platform, single-stranded linear spherical DNA nanoballs (DNA NanoBall, DNB) arranged in a regular array are used to capture mRNA-each DNB carries coordinate identifier CID (Coordinate ID) and molecular identifier MID (Molecular ID), which are used to record coordinates and distinguish different transcripts respectively. When the tissue section is placed on the chip, the sample transcript sequences (mRNA) in the tissue can be combined with the capture probes or random primers with poly-T on the chip, and the cDNA is formed by reverse transcription, so as to be captured and carry the coordinate information. Through subsequent sequencing and data analysis, the gene expression information can be accurately restored to the original coordinates in the tissue.

[0031] Stereo-seq also performs nuclear staining on the section, which is aligned with the gene expression information in subsequent analysis, and can locate the gene expression amount on the area of a single cell. The analysis method in the unit of gene expression of a single cell is called cellbin. Of course, the chip can also be cut into continuous square regions according to a certain diameter, and the gene expression amount in each region is combined into square bin, which is usually called Bin* according to the diameter of the square region, such as Bin20 which is the analysis unit combined by 20*20 DNBs in series.

[0032] S101, analyzing sample transcript sequences, and preliminarily classifying the sequencing sequences into first sequences, second sequences, uncertain sequences and discarded sequences.

[0033] The specific method of this step S101 is as follows: For each sequencing sequence, the sample transcript sequence therein is used to align the first reference genome and the second reference genome, and according to the alignment result, the sequencing sequence is classified into the following cases: If the sequencing sequence is aligned to the first reference genome but not to the second reference genome, the sequencing sequence is classified as a first sequence; If the sequencing sequence is not aligned to the first reference genome but is aligned to the second reference genome, the sequencing sequence is classified as a second sequence; If the sequencing sequence is aligned to both the first reference genome and the second reference genome, the sequencing sequence is classified as an uncertain sequence; If the sequencing sequence is not aligned to both the first reference genome and the second reference genome, the sequencing sequence is classified as a discarded sequence.

[0034] Through the specific method of S101, the species of the sequencing sequence is preliminarily distinguished, that is, the first sequence corresponds to the first species, and the second sequence corresponds to the second species. The uncertain sequence will be further distinguished in the subsequent steps. It is worth mentioning that the generation of the uncertain sequence is often due to the existence of homologous genes between two species.

[0035] This embodiment takes two species of human and mouse as an example for illustration.

[0036] The sequencing sequence is aligned to the human reference genome and the mouse reference genome respectively, and the sequence is classified into four categories: only aligned to the human Rhuman, only aligned to the mouse Rmouse, aligned to both reference genomes Rboth, and not aligned to both reference genomes Rneither. Specifically, on the Stereo-seq platform, each sequencing sequence is composed of three sequences, namely the coordinate identifier CID (Coordinate ID) for recording the coordinate position, the molecular identifier MID (Molecular ID) for distinguishing different transcripts, and the captured sample transcript sequence (mRNA).

[0037] For each sequencing sequence, only the mRNA is used to align the human reference genome and the mouse reference genome - if the sequencing sequence is aligned to the human reference genome but not to the mouse reference genome, the sequence is classified as a "human sequence" and denoted as Rhuman; if aligned to the mouse reference genome but not to the human reference genome, the sequence is classified as a "mouse sequence" and denoted as Rmouse; if aligned to both the human reference genome and the mouse reference genome, the sequence is classified as an "uncertain sequence" and denoted as Rboth; if not aligned to both the human reference genome and the mouse reference genome, the sequence is classified as a "discarded sequence" and denoted as Rneither. Table 1 shows the sequencing sequence and its classification result.

[0038] Table 1: Sequencing sequences and their classification results

[0039] S102, analyze the coordinate identifier, calculate the first ratio under each bin, the first ratio being the ratio of the number of sequencing sequences classified as the first sequence and the total number of sequencing sequences classified as the first sequence and the second sequence.

[0040] Please refer to Figure 2 The method for calculating the first ratio is as follows: S1021, according to the coordinate identifier, attach coordinates to the sequencing sequence; S1022, in combination with the coordinates, obtain the bin corresponding to the sequencing sequence; S1023, in units of bins, count the number of sequencing sequences classified as the first sequence under each bin, marked as the first number; and the number of sequencing sequences classified as the second sequence, marked as the second number; S1024, divide the first number by the sum of the first number and the second number to obtain the first ratio under the bin.

[0041] Still taking the two species of humans and mice as examples.

[0042] Use the official software SAW to analyze the coordinate identifier CID in the sequencing sequence, attach coordinates (x, y) to each CID, and further obtain the corresponding bin in combination with the coordinates.

[0043] In the Stereo-seq spatial transcriptome technology platform, three types of information are obtained after the experiment, one is the sequencing sequence information (recorded as reads file), the second is the corresponding information of the coordinate identifier CID region sequence on the sequence and the chip spatial coordinates (recorded as mask file, each coordinate identifier CID sequence corresponds to a pair of x and y, the corresponding relationship is determined when the experimental consumables are generated), and the third is the high-definition photo of the experimental tissue section cell nucleus staining (recorded as image file). In the official software SAW analysis process, the coordinate identifier CID of the sequencing sequence is first compared with the mask file to record the spatial coordinates (x and y coordinates) where each sequence is located, then the transcript region in the sequencing sequence is compared to the reference genome to obtain the gene expression, and the MID sequence is used to filter the repeated captured transcripts.

[0044] In the analysis process, there are two binning methods: One is to bin by square region, that is, to split all DNBs into square bins according to a set diameter size (such as 20 DNBs), that is, each bin contains the gene expression of adjacent 20*20 DNBs, simply referred to as Bin20.

[0045] The other is analysis by cell region, which uses high-definition photos of cell nucleus staining obtained experimentally to automatically identify and predict cell boundaries, and merge all DNBs and their gene expressions in the boundaries of a single cell into a cell bin, referred to as cellbin.

[0046] In this embodiment, the binning is either square binning or cell binning. Regardless of whether the binning is square binning or cell binning, the median x and y values ​​of the DNBs within the same bin are used as the bin coordinates. Through the SAW analysis process, an expression matrix of bin 20 (or cell bin)*genes can be obtained, and the spatial coordinates of each bin can be obtained.

[0047] Taking cell bins as an example, Table 2 shows the sequencing sequences and their corresponding bins.

[0048] Table 2: Sequencing sequences and their corresponding bins

[0049] The number of sequences marked as human sequence Rhuman and mouse sequence Rmouse in each bin was counted. During this process, the sequencing sequences were deduplicated. When the complete CID+MID+mRNA sequences of multiple sequencing sequences were repeated, the number of sequencing sequences was counted as only one.

[0050] The number of human sequences and mouse sequences after deduplication is marked as Nhuman (i.e., the first number) and Nmouse (i.e., the second number), respectively. The first ratio under this bin is calculated based on the following formula: Phuman=Nhuman / (Nhuman+Nmouse) Among them, Phuman is the first ratio, representing the proportion of human transcripts; Table 3 shows the proportion of human transcripts in the bins of this embodiment.

[0051] Table 3: Percentage of human transcripts in bins

[0052] S103: Classify the bins into a first bin and a second bin according to the first ratio; reclassify the sequencing sequences classified as uncertain sequences in the first bin as first sequences, and reclassify the sequencing sequences classified as uncertain sequences in the second bin as second sequences.

[0053] Specifically, the first ratio is compared with a set threshold, and the threshold is preferably 0.5; if the first ratio is greater than or equal to the threshold, the bin is classified as the first bin; if the first ratio is less than the threshold, the bin is classified as the second bin.

[0054] In the case of both human and mouse species, according to the first ratio Phuman (i.e. the proportion of human transcripts), as shown in Table 3, the species to which the bins belong is classified into two categories, human and mouse, and the sequencing sequences classified as uncertain sequences Rboth are divided into two categories (Rboth-human, Rboth-mouse) according to the bins. If the Phuman of a sequencing sequence is not less than 0.5, the species to which the bin belongs is marked as human; on the contrary, if the Phuman of the sequencing sequence is less than 0.5, the species to which the bin belongs is marked as mouse. The uncertain sequences Rboth belonging to human in the bin are reclassified as Rboth-human; similarly, the uncertain sequences Rboth belonging to mouse are reclassified as Rboth-mouse.

[0055] S104, all the first sequences are combined and analyzed, such as directly using the official software SAW and the human reference genome for analysis, to obtain the first spatial gene expression profile (i.e. the spatial gene expression profile of human) as shown in Table 4; similarly, all the second sequences are combined and analyzed to obtain the second spatial gene expression profile (i.e. the spatial gene expression profile of mouse) as shown in Table 5.

[0056] Table 4: Spatial gene expression profile of human

[0057] Table 5: Spatial gene expression profile of mouse

[0058] Finally, after the genes of the first spatial gene expression profile and the second spatial expression profile are respectively labeled with species tags, the bins are combined into a comprehensive spatial gene expression profile as shown in Table 6, i.e. the spatial gene expression profile of PDX model can be obtained for other analysis.

[0059] Table 6: Comprehensive spatial gene expression profile

[0060] The multiple classification steps of the embodiment can accurately realize the allocation of gene expression even if homologous genes exist in two species.

[0061] Example Two: Spatial transcriptome data analysis system based on PDX model The second embodiment provides a spatial transcriptome data analysis system based on PDX model, please refer to Figure 3 , which comprises the following functional modules: extraction module, preliminary classification module, proportion calculation module, final classification module, and expression profile generation module.

[0062] The extraction module extracts a coordinate identifier for recording coordinates, a molecular identifier for distinguishing different transcripts, and a captured sample transcript sequence from the sequencing sequence.

[0063] The preliminary classification module analyzes the sample transcript sequence and classifies the sequencing sequence into a first sequence, a second sequence, an uncertain sequence, and a discarded sequence.

[0064] Specifically, the method for classifying the sequencing sequence into a first sequence, a second sequence, an uncertain sequence, and a discarded sequence is as follows: For each sequencing sequence, the sample transcript sequence is aligned with a first reference genome and a second reference genome, and the alignment result is classified into the following cases: If the sequencing sequence is aligned to the first reference genome but not to the second reference genome, the sequencing sequence is classified as a first sequence; If the sequencing sequence is not aligned to the first reference genome but is aligned to the second reference genome, the sequencing sequence is classified as a second sequence; If the sequencing sequence is aligned to both the first reference genome and the second reference genome, the sequencing sequence is classified as an uncertain sequence; If the sequencing sequence is not aligned to both the first reference genome and the second reference genome, the sequencing sequence is classified as a discarded sequence.

[0065] The proportion calculation module analyzes the coordinate identifier and calculates a first ratio under each bin, where the first ratio is the ratio of the number of sequencing sequences classified as a first sequence to the total number of sequencing sequences classified as a first sequence and a second sequence.

[0066] Specifically, the method for calculating the first ratio under each bin is as follows: According to the coordinate identifier, the sequencing sequence is attached with coordinates; In combination with the coordinates, the bin corresponding to the sequencing sequence is obtained; In units of bins, the number of sequencing sequences classified as a first sequence under each bin is counted and marked as a first number, and the number of sequencing sequences classified as a second sequence is counted and marked as a second number; The first ratio under the bin is obtained by dividing the first number by the sum of the first number and the second number.

[0067] It is worth mentioning that, in each bin, the sequencing sequences are de-duplicated.

[0068] The de-duplication method of the sequencing sequence is as follows: When the coordinate identifier, the molecular identifier, and the sample transcript sequence of multiple sequencing sequences are repeated, it is determined that the multiple sequencing sequences are repeated, and in the counting of the number of sequencing sequences, the multiple sequencing sequences are only counted as one.

[0069] a final categorizing module, categorizing the bins into a first bin and a second bin according to the first ratio; re-categorizing the sequencing sequences categorized as uncertain sequences in the first bin into the first sequences, and re-categorizing the sequencing sequences categorized as uncertain sequences in the second bin into the second sequences.

[0070] The method of categorizing the bins into a first bin and a second bin according to the first ratio is as follows: comparing the first ratio with a set threshold value, if the first ratio is greater than or equal to the threshold value, categorizing the bin into the first bin; if the first ratio is less than the threshold value, categorizing the bin into the second bin; the threshold value is preferably 0.5.

[0071] a profile generating module, combining and analyzing all the first sequences to obtain a first spatial gene expression profile, and combining and analyzing all the second sequences to obtain a second spatial gene expression profile.

[0072] Finally, the genes in the first spatial gene expression profile and the second spatial gene expression profile are labeled with species tags respectively, and then combined into a comprehensive spatial gene expression profile according to the bins.

[0073] The specific execution methods and principles of the various functional modules in this embodiment two correspond to steps S100 to S104 in embodiment one one by one, and will not be repeated again.

[0074] Embodiment three: electronic device and readable storage medium This embodiment three also provides an electronic device and a readable storage medium.

[0075] Figure 4 A schematic block diagram of an example electronic device that can be used to implement embodiments of the application is shown. The electronic device is intended to represent various forms of digital computers, such as laptops, desktops, tablets, personal digital assistants, servers, blade servers, mainframes, and other appropriate computers. The electronic device can also represent various forms of mobile devices, such as personal digital assistants, cellular telephones, smartphones, wearable devices, and other similar computing devices. The components shown here, their connections and relationships, and their functions, are meant to be examples only, and are not meant to limit implementations of the application described and / or claimed in this document.

[0076] As Figure 4As shown, the device includes a computing unit that can perform various appropriate actions and processes in accordance with a computer program stored in a read-only memory (ROM) or a computer program loaded from a storage unit into a random access memory (RAM). In the RAM, various programs and data required for the operation of the device can also be stored. The computing unit, the ROM, and the RAM are connected to each other through a bus. An input / output (I / O) interface is also connected to the bus.

[0077] A plurality of components in the device are connected to the I / O interface, including: an input unit such as a keyboard, a mouse, etc.; an output unit such as various types of displays, a speaker, etc.; a storage unit such as a magnetic disk, an optical disk, etc.; and a communication unit such as a network card, a modem, a wireless communication transceiver, etc. The communication unit allows the device to exchange information / data with other devices through a computer network such as the Internet and / or various telecommunication networks.

[0078] The computing unit can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various specialized artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any appropriate processor, controller, microcontroller, etc. The computing unit performs various methods and processes described above, such as the PDX model-based spatial transcriptomic data analysis method described in Embodiment I. For example, in some embodiments, the PDX model-based spatial transcriptomic data analysis method can be implemented as a computer software program that is tangibly embodied in a machine-readable medium, such as the storage unit. In some embodiments, part or all of the computer program can be loaded and / or installed on the device via the ROM and / or the communication unit. When the computer program is loaded into the RAM and executed by the computing unit, one or more steps of the PDX model-based spatial transcriptomic data analysis method described above can be performed. Alternatively, in other embodiments, the computing unit can be configured to perform the PDX model-based spatial transcriptomic data analysis method by any other appropriate means (e.g., by means of firmware).

[0079] Various embodiments of the systems and techniques described above can be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), system-on-chip systems (SOCs), programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include being implemented in one or more computer programs that are executable and / or interpreted on a programmable system that includes at least one programmable processor, which can be a special purpose or general purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit data and instructions to the storage system, the at least one input device, and the at least one output device.

[0080] The program code for implementing the methods of the present application can be written in any combination of one or more programming languages. Such program code can be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device, so that when the program code is executed by the processor or controller, the functions / operations specified in the flow charts and / or block diagrams are implemented. The program code can be executed entirely on the machine, partially on the machine, as a stand-alone software package, partially on the machine and partially on a remote machine, or entirely on a remote machine or server.

[0081] In the context of this application, a machine-readable medium may be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, device, or apparatus. A machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium may include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing. More specific examples of machine-readable storage media may include an electrical connection based on one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), optical fibers, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0082] To provide for interaction with a user, the systems and techniques described here can be implemented on a computer having a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the computer. Other kinds of devices can be used to provide for interaction with a user as well; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form, including acoustic, speech, or tactile input.

[0083] The systems and techniques described here can be implemented in a computing system that includes a back end component (e.g., as a data server), or that includes a middleware component (e.g., an application server), or that includes a front end component (e.g., a user computer having a graphical user interface or a Web browser through which a user can interact with an implementation of the systems and techniques described here), or any combination of such back end, middleware, or front end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a local area network (LAN), a wide area network (WAN), and the Internet.

[0084] The computer system can include clients and servers. A client and server are generally remote from each other and typically interact through a communication network. The relationship of client and server is generally established by computer programs running on the respective computers and having a client-server relationship to each other. The servers can be cloud servers, servers of a distributed system, or servers combined with a blockchain.

[0085] It should be understood that the various forms of flow shown above can be re-ordered, added to, or deleted from without departing from the scope of the present disclosure. For example, the steps recited in the present disclosure can be executed in parallel, executed in series, or executed in different orders, as long as the desired results of the present disclosure are achieved, and the present disclosure is not limited herein.

[0086] In addition, the terms "first", "second", etc., are used herein only to describe different instances, and do not imply or suggest relative importance or an implied number of the indicated technical features. Thus, the features defined with "first", "second" can explicitly or implicitly include at least one of the features. In the description of the present disclosure, the meaning of "a plurality of" is two or more, unless otherwise specifically limited.

[0087] The above merely provides the specific implementation of the present application, but the protection scope of the present application is not limited to this. Any person skilled in the art can easily think of the changes or replacements within the technical range disclosed by the present application, which should be covered in the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.

Claims

1. A spatial transcriptome data analysis method based on PDX model, characterized in that: The following steps are involved: S100, extracting coordinate identifiers used to record coordinates in the sequencing sequence, molecular identifiers used to distinguish different transcripts, and captured sample transcript sequences; S101, analyze the sample transcript sequence and preliminarily classify the sequence into the first sequence, the second sequence, the uncertain sequence and the discarded sequence; S102, analyzing the coordinate identifiers and calculating a first ratio for each bin, where the first ratio is the ratio of the number of sequencing sequences classified as the first sequence to the total number of sequencing sequences classified as the first sequence and the second sequence; S103, classifying the bins into a first bin and a second bin according to the first ratio; reclassifying the sequencing sequences classified as uncertain sequences in the first bin as first sequences, and reclassifying the sequencing sequences classified as uncertain sequences in the second bin as second sequences; S104, all first sequences are combined and analyzed to obtain a first spatial gene expression profile; all second sequences are combined and analyzed to obtain a second spatial gene expression profile.

2. The spatial transcriptome data analysis method based on the PDX model according to claim 1, characterized in that In S101, the sample transcript sequences are analyzed and the sequencing sequences are preliminarily classified into first sequences, second sequences, uncertain sequences, and discarded sequences as follows: For each sequencing sequence, the sample transcript sequence is used to align the first reference genome with the second reference genome. The alignment results are divided into the following situations: If the sequence is aligned to the first reference genome but not to the second reference genome, the sequence is classified as the first sequence; If the sequenced sequence is not aligned to the first reference genome but is aligned to the second reference genome, the sequenced sequence is classified as the second sequence; If the sequence is aligned to both the first reference genome and the second reference genome, the sequence is classified as an undefined sequence. If the sequencing sequence is not aligned to the first reference genome and the second reference genome, the sequencing sequence is classified as a non-discarded sequence.

3. The spatial transcriptome data analysis method based on the PDX model according to claim 1, characterized in that In S102, the method for analyzing the coordinate identifier and calculating the first ratio in each bin is as follows: S1021, attaching coordinates to the sequencing sequence according to the coordinate identifier; S1022, obtaining the bins corresponding to the sequencing sequences based on the coordinates; S1023, counting the number of sequencing sequences classified as the first sequence in each bin, marked as a first number; and the number of sequencing sequences classified as the second sequence, marked as a second number; S1024: Divide the first quantity by the sum of the first quantity and the second quantity to obtain a first ratio for the bin.

4. The spatial transcriptome data analysis method based on the PDX model according to claim 3, characterized in that In S102, duplicate sequencing sequences are removed in each bin.

5. The spatial transcriptome data analysis method based on the PDX model according to claim 4, characterized in that The method for deduplication of the sequencing sequence is as follows: When the coordinate identifiers, molecular identifiers, and sample transcript sequences of multiple sequencing sequences are all repeated, the multiple sequencing sequences are determined to be repeated. When counting the number of sequencing sequences in S1023, the multiple sequencing sequences are only counted as one.

6. The spatial transcriptome data analysis method based on the PDX model according to claim 1, characterized in that In S103, the method of classifying the bins into the first bin and the second bin according to the first ratio is as follows: The first ratio is compared with a set threshold value. If the first ratio is greater than or equal to the threshold value, the bin is classified as the first bin; if the first ratio is less than the threshold value, the bin is classified as the second bin.

7. The method for analyzing spatial transcriptome data based on a PDX model according to claim 6, wherein: The threshold is set to 0.

5.

8. The method for analyzing spatial transcriptome data based on a PDX model according to claim 6, wherein: In S104, after the genes in the first spatial gene expression profile and the second spatial expression profile are respectively labeled with species labels, they are merged into a comprehensive spatial gene expression profile according to bins.

9. An electronic device, characterized in that: include: at least one processor; as well as a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method according to any one of claims 1 to 8.

10. A non-transitory computer-readable storage medium storing computer instructions, characterized in that: The computer instructions are used to cause the computer to execute the method according to any one of claims 1 to 8.

Citation Information

Patent Citations

  • Methods for histological diagnosis and treatment of diseases

    CN108350507A

  • Systems and methods for spatial analysis of analytes using fiducial alignment

    CN115023734A

  • Method and device for analyzing sequencing data of space transcriptome chip

    CN115331733A

  • PDX-based single cell transcriptome data analysis method, system, equipment and medium

    CN116153401A

  • PDX model single cell transcriptome data analysis method, device and medium

    CN117912552A