Single-cell multi-omics data analysis system, method and equipment and storage medium

By designing a single-cell multi-omics data analysis system, the problem of poor compatibility among single-cell multi-omics analysis tools was solved, cross-platform data compatibility, seamless integration of multi-language tools, and efficient access to deep learning frameworks were achieved, improving the efficiency and scalability of the analysis process.

CN120766779APending Publication Date: 2025-10-10XI AN JIAOTONG UNIV

Patent Information

Application Number
CN202510910837.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-02
Publication Date
2025-10-10

AI Technical Summary

Technical Problem

Existing single-cell multi-omics analysis tools have poor compatibility with each other, complex cross-language operations, high platform heterogeneity barriers, lack of unified interfaces, and are unable to quickly adapt to new data types and algorithms. The analysis process is fragmented and inefficient.

Method used

A single-cell multi-omics data analysis system was designed, including a data import module, a cross-language analysis module, a deep learning optimization module, and an interaction and extension module. This system enables cross-platform data compatibility, seamless integration of multi-language tools, and efficient access to deep learning frameworks. The system unifies data formats through S4 objects and supports mixed calls of R and Python languages ​​and deep learning model training.

Benefits of technology

It achieves data compatibility across sequencing platforms, reduces the complexity of data preprocessing, avoids cross-language data conversion losses, supports custom plug-in development, quickly responds to technology iteration needs, and improves the efficiency and scalability of the analysis process.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120766779A_ABST
    Figure CN120766779A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of single-cell multi-omics, and discloses a single-cell multi-omics data analysis system, method and device and a storage medium, a data import module is used for reading sequencing data from different sequencing platforms and storing the sequencing data as S4 objects; the cross-language analysis module is used for calling a single-cell multi-omics analysis tool based on an R language and a Python language to analyze an S4 object according to an analysis process in a unified framework; the deep learning optimization module is used for training a deep learning model based on the cross-language interaction interface and integrating an analysis result of the deep learning model to an object S4; and the interaction and extension module is used for obtaining an analysis process and visualizing an analysis result, constructing a user-defined single-cell multi-omics analysis tool in a plug-in form and loading the tool to the S4 object for analysis. Cross-sequencing platform data compatibility, multi-language tool seamless integration and deep learning framework efficient access are achieved, and the problems of data isomerism, language ecological splitting and algorithm expansibility in the prior art are solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of single-cell multi-omics technology, and relates to a single-cell multi-omics data analysis system, method, equipment and storage medium. Background Art

[0002] Single-cell transcriptomics and epigenomics provide unprecedented resolution for revealing cellular heterogeneity, gene regulatory networks, and tissue microenvironments by analyzing gene expression, chromatin accessibility, and cellular spatial localization information. scRNA-seq (single-cell RNA sequencing) and scATAC-seq (single-cell ATAC sequencing) characterize cellular states at the transcriptome and epigenomic levels, respectively, while spatial transcriptomics further associates gene expression with cellular spatial location, promoting research on developmental biology, tumor microenvironment, and disease mechanisms. Current single-cell sequencing technologies are highly diverse, with significant differences in spatial resolution, molecular capture type, and data output format, making it difficult for analytical tools to be unified and compatible. However, the explosive growth of multi-omics data and the diversity of platform technologies pose severe challenges to data processing and analysis.

[0003] At present, single-cell multi-omics analysis mainly relies on the following technical frameworks: (1) Single-omics-specific analysis tools. Tools represented by Seurat (R language), Scanpy (Python language) and ArchR (R language) provide standardized analysis processes for transcriptome and epigenomic data, respectively. For example, Seurat supports dimensionality reduction clustering and differential expression analysis of scRNA-seq data, and ArchR focuses on peak identification and chromatin accessibility visualization of ATAC-seq data. Although such tools perform well in single-omics analysis, their cross-omics data integration capabilities are limited. For example, Seurat is not compatible with spatial ATAC-seq data, and ArchR cannot directly process spatial transcriptomics data. In addition, there are many analysis tools such as CellChat and Monocle3, but they each use different data structures and target different analysis contents (such as cell communication and trajectory inference). Currently, there is a lack of a unified platform to systematically integrate these analysis tools. (2) Multi-language hybrid analysis ecosystem. Mainstream tools have functional fragmentation due to the differentiation of programming languages. For example, users need to manually export ATAC-seq peak analysis results from ArchR and then import them into Scanpy for spatial neighborhood enrichment analysis. This cross-language operation not only increases learning costs but also leads to information loss or format errors due to data conversion.

[0004] Based on the above analysis, the existing single-cell multi-omics analysis tools have the following defects: 1. Multi-omics analysis tools are redundant and diverse and have poor mutual compatibility: transcriptomics and epigenomics each have independent analysis tools. These single-omics tools focus on specific analysis directions, and their data structures are incompatible with each other. Users need to frequently convert data formats to adapt to different tools, resulting in fragmented analysis processes. 2. Platform heterogeneity barriers: The data formats generated by different sequencing platforms vary significantly, and there is a lack of unified interface support, resulting in inefficient data preprocessing and import. 3. Cross-language operation redundancy: The separation of the R language and Python language tool chains forces users to switch between multiple environments, resulting in poor reproducibility of the analysis process, and the lack of compatibility between the deep learning framework and the R language ecosystem further exacerbates the complexity of the technology stack. 4. Lack of dynamic adaptability: Existing tools have limited scalability for emerging technologies and cannot quickly adapt to new data types or algorithms through modular design. Summary of the Invention

[0005] The purpose of the present invention is to overcome the shortcomings of the above-mentioned prior art and provide a single-cell multi-omics data analysis system, method, device and storage medium.

[0006] In order to achieve the above object, the present invention adopts the following technical solutions:

[0007] In a first aspect, the present invention provides a single-cell multi-omics data analysis system, comprising a data import module, a cross-language analysis module, a deep learning optimization module, and an interaction and extension module; the data import module is used to read sequencing data from different sequencing platforms and save it as an S4 object; the cross-language analysis module is used to call single-cell multi-omics analysis tools based on R language and Python language according to the analysis process within a unified framework to analyze the S4 object; the deep learning optimization module is used to train a deep learning model on the S4 object based on a cross-language interaction interface, and integrate the analysis results of the deep learning model on the S4 object into the S4 object; the interaction and extension module is used to obtain the analysis process and visualize the analysis results, as well as to build a custom single-cell multi-omics analysis tool in the form of a plug-in and load it into the S4 object for analysis.

[0008] Optionally, the data import module is specifically used to: obtain sequencing files from different sequencing platforms, and read the sequencing data of sequencing files from different sequencing platforms based on dynamic adapters pre-designed based on different sequencing file formats, and convert the read sequencing data into a unified format based on the data storage architecture in Parquet format and save it to an S4 object.

[0009] Optionally, the data import module is further configured to: obtain SRA ID, and download epigenome sequencing data from NCBI according to the SRA ID, and use 10XCellRanger-ATAC tool to identify open chromatin regions, motif annotation, and differential accessibility analysis, and generate fragments.tsv.gz file for downstream analysis; obtain user-uploaded custom single-cell multi-omics sequencing file, and read custom single-cell multi-omics sequencing data using preset corresponding mode according to file form of the custom single-cell multi-omics sequencing file.

[0010] Optionally, when the analysis process calls the single-cell multi-omics analysis tool based on R language and Python language to analyze the S4 object, the data exchange between the R language and the Python language is implemented by using the.parquet file, and the analysis result is saved to the S4 object and dynamically displayed by using Echarts.

[0011] Optionally, the deep learning optimization module is specifically configured to: according to a Pytorch script or a tensorflow script written by a user in a specified format, build a communication bridge of an R language and a PyTorch interface or a TensorFlow interface by calling a python interpreter in a current environment, convert the S4 object to Python language for deep learning model training, and convert analysis result of the deep learning model on the S4 object to R language and integrate the analysis result into the S4 object.

[0012] Optionally, the deep learning optimization module is further configured to: embed a plurality of deep learning models for single-cell multi-omics sequencing data analysis, and directly call the embedded deep learning models on the S4 object through an R language script to obtain analysis result of the deep learning models on the S4 object and integrate the analysis result into the S4 object.

[0013] Optionally, the interaction and expansion module is specifically configured to: obtain a Pytorch script or a tensorflow script written by a user in a specified format, call a python interpreter in a current environment based on a script path, encapsulate a custom single-cell multi-omics analysis tool written by a user according to a standardized API interface into an independent plug-in, and dynamically load the independent plug-in to the S4 object for analysis through modular design.

[0014] In a second aspect, the present application provides a single-cell multi-omics data analysis method, comprising: reading sequencing data from different sequencing platforms and saving as S4 objects through a data import module; performing analysis of S4 objects based on single-cell multi-omics analysis tools based on R language and Python language according to an analysis process in a unified framework through a cross-language analysis module; training a deep learning model on S4 objects based on a cross-language interaction interface through a deep learning optimization module, and integrating analysis results of S4 objects by the deep learning model to S4 objects; obtaining the analysis process and visualizing the analysis results through an interaction and expansion module, and constructing a plug-in form of a custom single-cell multi-omics analysis tool and loading to S4 objects for analysis.

[0015] In a third aspect, the present application provides a computer device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the steps of the single-cell multi-omics data analysis method when executing the computer program.

[0016] In a fourth aspect, the present application provides a computer readable storage medium, which stores a computer program, wherein the computer program implements the steps of the single-cell multi-omics data analysis method when executed by a processor.

[0017] Compared with the prior art, the present application has the following beneficial effects:

[0018] The single-cell multi-omics data analysis system of the present application realizes reading of sequencing data from different sequencing platforms through the design of the data import module, completes compatible adaptation of different sequencing platforms, and significantly reduces the complexity of data preprocessing. The cross-language analysis module breaks through the ecological barrier of the R language and Python language tool chain, and can call functions based on the R language and the Python language such as Seurat and Scanpy in the same framework, avoiding loss and operation redundancy of cross-language data conversion. The deep learning optimization module overcomes the compatibility bottleneck of the R language and the deep learning framework, and can directly complete efficient deep learning model training and result analysis on S4 objects. The interaction and expansion module supports custom plug-in development and adaptation of emerging sequencing technologies, and quickly responds to technical iteration requirements through modular design. The single-cell multi-omics data analysis system of the present application realizes cross-sequence platform data compatibility, seamless integration of multi-language tools, and efficient access of deep learning frameworks, effectively solving the problems of data heterogeneity, language ecological fragmentation, and algorithm expandability in the prior art. BRIEF DESCRIPTION OF DRAWINGS

[0019] Figure 1 FIG. 1 is a structural block diagram of the single-cell multi-omics data analysis system of the present application.

[0020] Figure 2A single-cell multi-omics data analysis system use tutorial interface diagram of an embodiment of the present application.

[0021] Figure 3 A transcriptomics analysis webpage interface diagram of an embodiment of the present application.

[0022] Figure 4 An epigenomics analysis webpage interface diagram of an embodiment of the present application.

[0023] Figure 5 A single-cell multi-omics data analysis method flowchart of an embodiment of the present application. DETAILED DESCRIPTION

[0024] In order to enable those skilled in the art to better understand the present application, the technical solutions in the embodiments of the present application will be described clearly and completely below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative work should fall within the scope of protection of the present application.

[0025] It should be noted that the terms "first", "second", and the like in the specification and claims of the present application and the above-described drawings are used to distinguish similar objects, and do not necessarily indicate a specific order or a chronological sequence. It should be understood that the data thus used can be interchanged under appropriate circumstances, so that the embodiments of the present application described herein can be implemented in an order other than that illustrated or described herein. In addition, the terms "include" and "have" and any variations thereof are intended to cover non-exclusive inclusion, for example, a process, method, system, product or device that includes a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but can include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.

[0026] The present application will be described in further detail below with reference to the drawings:

[0027] Referring to Figure 1 In an embodiment of the present application, a single-cell multi-omics data analysis system is provided, specifically a single-cell transcriptomics and epigenomics data processing and analysis system based on Python language and R language, which realizes cross-sequencing platform data compatibility, seamless integration of multi-language tools and efficient access to deep learning frameworks, to solve the core bottlenecks of existing technologies in data heterogeneity, language ecological fragmentation and algorithm scalability.

[0028] Specifically, the single-cell multi-omics data analysis system comprises a data import module, a cross-language analysis module, a deep learning optimization module, and an interaction and expansion module. The data import module is configured to read sequencing data from different sequencing platforms and save the sequencing data as S4 objects. The cross-language analysis module is configured to call single-cell multi-omics analysis tools based on R language and Python language to analyze the S4 objects according to an analysis process in a unified framework. The deep learning optimization module is configured to train a deep learning model on the S4 objects based on a cross-language interaction interface, and integrate analysis results of the S4 objects by the deep learning model into the S4 objects. The interaction and expansion module is configured to obtain the analysis process and visualize the analysis results, and construct a plug-in form of a custom single-cell multi-omics analysis tool and load the tool into the S4 objects for analysis.

[0029] Illustratively, the S4 object is an implementation of object-oriented programming in R language, which allows developers to encapsulate data and related functions together to achieve higher-level data processing and analysis.

[0030] The single-cell multi-omics data analysis system can read sequencing data from different sequencing platforms by designing the data import module, complete compatibility and adaptation of different sequencing platforms, and significantly reduce the complexity of data preprocessing. The cross-language analysis module can break through the ecological barriers of R language and Python language tool chains, and can call functions based on R language and Python language in the same framework, such as Seurat (an R package for analyzing single-cell transcriptome data) and Scanpy (a Python library for single-cell gene expression data analysis), to avoid loss and operation redundancy of cross-language data conversion. The deep learning optimization module can overcome the compatibility bottleneck of R language and deep learning framework, and can directly complete efficient deep learning model training and result analysis on the S4 objects. The interaction and expansion module can support custom plug-in development and adaptation of emerging sequencing technologies, and can quickly respond to technical iteration requirements through modular design. The single-cell multi-omics data analysis system can realize cross-sequencing platform data compatibility, seamless integration of multi-language tools, and efficient access of deep learning framework, and effectively solve the problems of data heterogeneity, language ecological fragmentation, and algorithm expandability in the prior art.

[0031] In a possible implementation, the data import module is specifically configured to obtain sequencing files of different sequencing platforms, read sequencing data of the sequencing files of different sequencing platforms based on pre-designed dynamic adapters of different sequencing file forms, and convert the read sequencing data into a unified format based on a Parquet format data storage architecture and save the data into S4 objects.

[0032] Optionally, the data import module is further configured to: obtain an SRA ID (a unique identifier in an SRA database maintained by NCBI), and download epigenomic sequencing data from NCBI (the National Center for Biotechnology Information) according to the SRA ID, and use a 10X CellRanger-ATAC tool to identify open chromatin regions, motif annotation, and differential accessibility analysis, and generate a fragments.tsv.gz file for downstream analysis; and obtain a user-uploaded custom single-cell multi-omics sequencing file, and read custom single-cell multi-omics sequencing data according to a file form of the custom single-cell multi-omics sequencing file using a preset corresponding mode.

[0033] Illustratively, the data import module is configured to read raw data from different sequencing platforms (such as 10XVisium, Stereo-seq, spatial ATAC-seq, and Nanostring CosMx, etc.), and customize reading modes for various different formats of data, and save required information (such as a count matrix and spatial coordinates, etc.) to a unified S4 object. The data import module supports diversified input formats (including a gene expression matrix, spatial coordinates, and epigenomic peak values, etc.), and eliminates differences in data structures between platforms by automatically analyzing data and integrating standardized downstream analysis processes such as dimensionality reduction, clustering, and high-variable gene screening. In addition, the data import module supports importing raw epigenomic sequencing data by SRA ID.

[0034] Specifically, based on a high-efficiency data storage architecture of Parquet format, raw sequencing data (such as FASTQ, BAM, and HDF5 formats, etc.) is converted into a unified format with light weight and low storage occupancy. According to characteristics of different sequencing platforms, a dynamic adapter (such as a 10XVisium adapter for analyzing spatial coordinates and a gene expression matrix, and a Stereo-seq adapter for processing high-resolution nanopore data) is designed according to the form of the sequencing file to ensure data reading accuracy and efficiency. In addition, when importing raw epigenomic sequencing data by SRA ID, data will be automatically downloaded from NCBI and analyzed using a 10X CellRanger-ATAC tool to identify open chromatin regions, motif annotation, differential accessibility, and generate a fragments.tsv.gz file for downstream analysis (such as Signac and ArchR, etc.). Users can upload custom data (such as CSV (comma-separated values), H5AD (Anndata format), and RDS (R data serialization format), etc.) through a web interface or a local command line, automatically identify data types according to the form of the sequencing file, read information using a corresponding mode, and save to an S4 object to generate a standardized multi-omics object.

[0035] In addition, exemplary, in view of the privacy requirement of cell sequencing data, the webpage provided only provides analysis examples based on public data sets, and the user needs to use the locally deployed webpage during actual analysis to ensure that all data will not be uploaded to the cloud.

[0036] In a possible implementation, when the analysis process calls the single-cell multi-omics analysis tool based on R language and Python language to analyze the S4 object, the.parquet file is used to realize data exchange between R language and Python language, and the analysis result is saved to the S4 object and dynamically displayed through Echarts.

[0037] Explanatorily, the cross-language analysis module supports calling mainstream analysis tools (such as Seurat, Scanpy, ArchR, CellChat and Monocle3) of R and Python ecology in a unified framework to realize whole-process analysis of transcriptomics and epigenomics data. Users can complete data preprocessing, dimensionality reduction clustering, cell communication analysis and trajectory inference without switching programming environments.

[0038] Specifically, the data of R language is saved to the.parquet file, a Python script is written to read the data and perform corresponding analysis, and then saved to the.parquet file again; during running, the python interpreter is called through the command line in R language to run the script, realizing the quick integration of R language and Python, and also meeting the future expansion needs (users can write Pytorch or tensorflow scripts according to the specified format, and directly align with the traditional analysis process), so as to realize the lossless integration between Python and R language.

[0039] Exemplary, users can call the clustering algorithm of Seurat (an R language toolkit for single-cell RNA-seq analysis) in the R environment, and then directly pass the result to the spatial neighborhood enrichment analysis of Scanpy (a single-cell analysis toolkit in Python) in the Python environment without manually exporting or converting the data format. All intermediate results and visualization charts (such as UMAP (a nonlinear dimensionality reduction algorithm) dimensionality reduction chart and cell trajectory tree) are stored in the multi-omics object, supporting cross-language conversion and Echarts (a JavaScript-based visualization library) dynamic display.

[0040] In a possible implementation, the deep learning optimization module is specifically configured to: according to a Pytorch script or a tensorflow script written by a user in a specified format, build a communication bridge between R language and a PyTorch interface or a TensorFlow interface by calling a python interpreter in a current environment, convert an S4 object to Python language for deep learning model training, and convert analysis results of the deep learning model on the S4 object to R language and integrate the analysis results into the S4 object.

[0041] Optionally, the deep learning optimization module is further configured to: internally store a plurality of deep learning models for single-cell multi-omics sequencing data analysis, and directly call the internal deep learning models on the S4 object through an R language script to obtain analysis results of the deep learning models on the S4 object and integrate the analysis results into the S4 object.

[0042] Illustratively, the deep learning optimization module internally stores a plurality of advanced models (such as SEDR and GraphST), and provides interfaces with PyTorch and TensorFlow, effectively solving the compatibility problem between the R language ecology and the deep learning framework. A user can directly train a model on an S4 object of the R language, and seamlessly integrate prediction results into downstream analysis processes such as dimension reduction, clustering, and high-variable gene screening.

[0043] Specifically, a user can write a Pytorch script or a tensorflow script in a specified format, and specify a script path in a parameter of a function. When the function is run, a python interpreter in a current environment is automatically called to run the script, thereby building a communication bridge between R language and PyTorch / TensorFlow. A user can directly call a preset deep learning model (such as STALigner for spatial feature optimization) in an R script, or train a new model through a custom Python script. The system automatically converts input data into a tensor format and returns analysis results. In addition, the module can provide GPU acceleration support, which significantly improves the training efficiency of large-scale data.

[0044] In a possible implementation, the interaction and extension module is specifically configured to: obtain a Pytorch script or a tensorflow script written by a user in a specified format, and call a python interpreter in a current environment based on a script path, encapsulate a custom single-cell multi-omics analysis tool written by the user according to a standardized API interface into an independent plug-in, and dynamically load the independent plug-in to an S4 object for analysis through modular design.

[0045] Interpretatively, the interaction and expansion module can provide a visual interface between the web and the local end, support interactive analysis, result export and user-defined algorithm extension. Users can build analysis processes through drag-and-drop operations, or integrate new sequencing technology support through plug-in mechanisms.

[0046] Specifically, the web interactive interface is developed based on the Shiny framework. Users can upload data, select analysis steps (such as normalization, clustering, and differential gene identification), and view visual results (such as heat maps and network graphs) in real time. For developers, standardized API interfaces and example templates are provided to allow users to write Pytorch or tensorflow scripts according to the specified format and specify the script path in the function parameters. The function automatically calls the python interpreter in the current environment to implement running and encapsulating new algorithms (such as new batch correction methods) as independent plug-ins, and dynamically loads them into the multi-omics object through modular design.

[0047] Interpretatively, the single-cell multi-omics data analysis system of the application can integrate most mainstream open-source analysis tools including Seurat, Scanpy, ArchR, CellChat and Monocle3 from raw data reading, preprocessing, analysis to visualization, and can encapsulate all functions used as simple functions. Users only need to write 1 line of code to complete the complex analysis process that previously required hundreds of lines of scripts.

[0048] Illustratively, the application can also support distribution through Docker, and provides an online analysis platform web page display, detailed tutorials and open datasets to reduce learning and use costs and help single-cell transcriptomics and epigenomics technology to be deeply transformed from theoretical research to clinical practice.

[0049] In general, the single-cell multi-omics data analysis system of the application has the following advantages:

[0050] First, full data compatibility across sequencing platforms is achieved. By developing dynamic adapters and unified R language S4 objects, data import for mainstream sequencing technologies such as 10X Visium, spatialATAC-seq, Stereo-seq, Slide-seq v2, Nanostring CosMx and Vizgen MERFISH can be seamlessly supported. For different platform data formats (such as Visium HDF5 matrix, Stereo-seq nanoscale coordinates and MERFISH targeted metadata), they can be automatically parsed and standardized into unified objects, effectively eliminating data heterogeneity barriers.

[0051] Second, this invention overcomes the core challenge of a fragmented multilingual ecosystem. By building a hybrid R and Python computing kernel, users can seamlessly call tool chains such as Seurat(R), Scanpy(Python), and ArchR(R) within the same framework. To meet users' exploratory analysis needs, it also supports bidirectional conversion between data objects from various platforms such as Seurat and Scanpy and the unified data objects in this invention. This design significantly reduces the amount of code required for cross-language operations and avoids the loss of feature information caused by cross-language data conversion.

[0052] Third, the present invention provides full-process analysis and extremely simple operation experience. The complete analysis process of single-cell multi-omics (including data normalization, dimensionality reduction clustering, differential gene identification, cell communication and trajectory inference, etc.) is encapsulated into several concise functions. Users only need at least 1 line of code to complete complex tasks that traditionally require hundreds of lines of code to complete the full-process analysis. At the same time, all intermediate results and visual charts (such as UMAP dimensionality reduction maps and cell space clustering maps, etc.) are stored in a unified object and support export as RDS files. In addition, a web-side interactive interface can be provided (supporting dragging sliders to adjust parameters and one-click running of the full-process analysis, and some results are interactively displayed using Echarts), which can reduce the cost of code learning as much as possible and meet the diverse needs of clinical researchers to algorithm developers.

[0053] Fourth, the present invention overcomes the compatibility bottleneck between the R language and deep learning frameworks. By saving R language data as parquet files and then automatically reading and processing them with the Python interpreter, the present invention integrates open source deep learning frameworks such as PyTorch. Users can directly call proven, reliable pre-built models or write custom deep learning scripts in R scripts, eliminating the need to manually transfer existing data to the Python environment. This significantly reduces the difficulty for R language users to call deep learning frameworks.

[0054] Fifth, the present invention is highly flexible and scalable. Its modular architecture allows developers to quickly package new algorithms, typically those based on deep learning, using standardized methods.

[0055] The following describes a possible workflow of the single-cell multi-omics data analysis system of the present invention:

[0056] Step 1 (Epigenomics, optional): Input SRA ID, the system automatically downloads epigenomics sequencing data in SRA format through NCBI and uses the CellRanger-ATAC tool provided by 10X Genomics to align the sequencing data with the reference genome, identify chromatin accessible regions, and obtain the cell x Peak generation matrix and standard files (fragments.tsv.gz) that can be used for downstream analysis.

[0057] Step 2: Import data and build S4 objects. When importing, support users to specify data types (such as "Visium"), read each data according to the official standard format open data, extract and save the count matrix, spatial coordinates and other information to a unified S4 object. In addition, it can also support the input of cell number and gene number parameters to filter the original data.

[0058] Step 3: Dimensionality reduction and clustering. For scRNA-seq data, combine Seurat and Voyager for PCA dimensionality reduction, construct KNN graph and optimize edge weight through Jaccard similarity, use Louvain or Leiden algorithm for clustering, and embed the results in UMAP and spatial atlas visualization; for scATAC-seq data, use ArchR for spatial information filtering and LSI dimensionality reduction, construct KNN graph and perform the same clustering.

[0059] Step 4: Perform GO enrichment, cell communication and cell trajectory inference analysis on S4 objects.

[0060] Among them, GO enrichment analysis imports differential genes and species database through clusterProfiler's enrichGO, calculates P value based on hypergeometric test and performs BH correction, selects significantly enriched items, calculates enrichment score with GeneRatio and background ratio, and displays the top 3 biological functions with the highest P value in each cell category.

[0061] Cell communication analysis calls CellChat to input gene matrix and cell type annotation, filters ligand-receptor pairs and calculates average expression product, constructs directed communication network and signal pathway, evaluates significance through permutation test, identifies key pathways combined with network centrality, and visualizes receptor-ligand expression with heatmap.

[0062] Cell trajectory inference uses Monocle-3 to normalize the expression matrix and filter high-variable genes, after UMAP / PCA dimensionality reduction and Leiden clustering, constructs the minimum spanning tree to form the trajectory through the reverse graph embedding, assigns pseudo-time values to reflect the differentiation order, and identifies branch regulators combined with gene dynamics.

[0063] The single-cell multi-omics data analysis system integrates R and Python environments, supports R multi-omics platform calling deep learning tools such as PyTorch, and internally builds STAligner, SANTO and SEDR spatial alignment and batch correction methods based on Pytorch. In addition, users can write.py scripts according to the paradigm to realize the deep integration of the R language analysis platform and PyTorch.

[0064] In another embodiment of the present application, the single-cell multi-omics data analysis system uses the Shiny framework and Bootstrap components and Echarts components to make front-end web pages, providing corresponding visualization solutions, providing one-key import, data preprocessing, dimensionality reduction, clustering, cell communication and cell trajectory inference analysis functions for 10XVisium, spatial ATAC-seq and other data. For R language scripts, these analyses have been packaged in a function in this platform, supporting part or complete running process. For the web side, after adjusting the parameters by dragging the slider and clicking the run button, all or part of the process analysis can be automatically completed.

[0065] Referring to Figure 2 , a usage tutorial interface of the single-cell multi-omics data analysis system in the present embodiment is shown. Referring to Figure 3 , an analysis web interface of the single-cell multi-omics data analysis system in the present embodiment for transcriptomics (taking 10XVisium analysis as an example) is shown. Referring to Figure 4 , an analysis web interface of the single-cell multi-omics data analysis system in the present embodiment for epigenomics is shown.

[0066] The following is an apparatus embodiment of the present application, which can be used to execute the method embodiments of the present application. For details not disclosed in the apparatus embodiment, please refer to the method embodiments of the present application.

[0067] Referring to Figure 5 , in another embodiment of the present application, a single-cell multi-omics data analysis method is provided, which can be implemented based on the above-mentioned single-cell multi-omics data analysis system. Specifically, the single-cell multi-omics data analysis method comprises the following steps:

[0068] S1: reading sequencing data from different sequencing platforms through a data import module and saving it as an S4 object.

[0069] S2: through a cross-language analysis module, calling single-cell multi-omics analysis tools based on R language and Python language in a unified framework according to the analysis process to analyze the S4 object.

[0070] S3: The deep learning optimization module trains a deep learning model on the S4 object based on the cross-language interaction interface and integrates the analysis result of the S4 object by the deep learning model to the S4 object.

[0071] S4: The interaction and expansion module obtains an analysis process and visualizes an analysis result, and constructs a plug-in form of a custom single-cell multi-omics analysis tool and loads the tool to the S4 object for analysis.

[0072] The foregoing embodiments of the single-cell multi-omics data analysis system involve all related contents of each step, which can be cited to the function description of the function module corresponding to the single-cell multi-omics data analysis method in the embodiments of the present application, and will not be repeated here.

[0073] The division of the modules in the embodiments of the present application is illustrative, and is only a logical function division. In actual implementation, another division mode can be used. In addition, each function module in each embodiment of the present application can be integrated in one processor, or can be physically separated, or two or more modules can be integrated in one module. The integrated module can be realized in the form of hardware or in the form of a software function module.

[0074] In still another embodiment of the present application, a computer device is provided, which includes a processor and a memory. The memory is used to store a computer program, the computer program includes program instructions, and the processor is used to execute the program instructions stored in the computer storage medium. The processor can be a central processing unit (CPU), and can also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs) or other programmable logic devices, discrete gates or transistor logic devices, discrete hardware components, etc. It is the computing core and control core of the terminal, which is suitable for implementing one or more instructions, and is specifically suitable for loading and executing one or more instructions in the computer storage medium to realize a corresponding method process or a corresponding function. The processor in the embodiments of the present application can be used for the operation of the single-cell multi-omics data analysis method.

[0075] In still another embodiment of the present application, the present application also provides a storage medium, specifically a computer readable storage medium (Memory), which is a memory device in a computer device, used for storing programs and data. It can be understood that the computer readable storage medium herein can include an internal storage medium in the computer device, and of course can also include an extended storage medium supported by the computer device. The computer readable storage medium provides a storage space, which stores an operating system of the terminal. Moreover, one or more instructions suitable for being loaded and executed by the processor are also stored in the storage space, and these instructions can be one or more computer programs (including program codes). It should be noted that the computer readable storage medium herein can be a high-speed RAM memory, or a non-volatile memory such as at least one disk memory. One or more instructions stored in the computer readable storage medium can be loaded and executed by the processor to realize the corresponding steps of the single-cell multi-omics data analysis method in the above embodiments.

[0076] Those skilled in the art should understand that embodiments of the present application can be provided as methods, systems, or computer program products. Therefore, the present application can take the form of an entirely hardware embodiment, an entirely software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROMs, optical storage, etc.) containing computer-usable program code.

[0077] The present application is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each flow and / or block in the flowcharts and / or block diagrams, and the combination of the flows and / or blocks in the flowcharts and / or block diagrams can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing apparatus to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing apparatus produce a device that implements the functions specified in the flowcharts and / or block diagrams. Figure 1 The functions specified in a flow or multiple flows and / or blocks Figure 1 The means for performing the functions specified in a flow or multiple flows and / or blocks.

[0078] These computer program instructions can also be stored in a computer readable memory capable of directing the computer or other programmable data processing apparatus to work in a specific manner, so that the instructions stored in the computer readable memory produce a manufactured product including instruction means, which implements the functions specified in the flowcharts and / or block diagrams. Figure 1 The functions specified in a flow or multiple flows and / or blocksFigure 1 the function specified in the one or more blocks.

[0079] These computer program instructions can also be loaded into computer or other programmable data processing devices, so that a series of operation steps are performed on the computer or other programmable data processing devices to generate computer-implemented processes, thus the instructions executed on the computer or other programmable data processing devices provide a process for implementing the flow Figure 1 the flow or flows and / or blocks Figure 1 the steps of the function specified in the one or more blocks.

[0080] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present application rather than limit the same, and although the present application has been described in detail with reference to the above embodiments, those of ordinary skill in the art should understand that the specific embodiments of the present application can be modified or replaced equivalently without departing from the spirit and scope of the present application, and any modification or equivalent replacement without departing from the spirit and scope of the present application should be covered within the protection scope of the claims of the present application.

Claims

1. A single-cell multi-omics data analysis system, characterized in that: Includes data import module, cross-language analysis module, deep learning optimization module and interaction and extension module; The data import module is used to read sequencing data from different sequencing platforms and save them as S4 objects; The cross-language analysis module is used to call single-cell multi-omics analysis tools based on R and Python languages ​​to analyze S4 objects according to the analysis process within a unified framework; The deep learning optimization module is used to train a deep learning model on the S4 object based on the cross-language interactive interface, and integrate the analysis results of the deep learning model on the S4 object into the S4 object; The interaction and extension module is used to obtain analysis processes and visualize analysis results, as well as to build custom single-cell multi-omics analysis tools in the form of plug-ins and load them into S4 objects for analysis.

2. The single-cell multi-omics data analysis system according to claim 1, characterized in that The data import module is specifically used for: Acquire sequencing files from different sequencing platforms, and use pre-designed dynamic adapters based on different sequencing file formats to read sequencing data from sequencing files on different sequencing platforms. Use a data storage architecture based on the Parquet format to convert the read sequencing data into a unified format and save it to an S4 object.

3. The single-cell multi-omics data analysis system according to claim 1, characterized in that The data import module is also used to: Obtain SRAID and download epigenomic sequencing data from NCBI based on SRAID. Use the 10XCellRanger-ATAC tool to identify open chromatin regions, perform motif annotation, and perform differential accessibility analysis to generate fragments.tsv.gz files for downstream analysis. Obtain the custom single-cell multi-omics sequencing file uploaded by the user, and read the custom single-cell multi-omics sequencing data using the preset corresponding method according to the file format of the custom single-cell multi-omics sequencing file.

4. The single-cell multi-omics data analysis system according to claim 1, characterized in that When the single-cell multi-omics analysis tool based on R language and Python language is called according to the analysis process to analyze the S4 object, the .parquet file is used to realize data exchange between R language and Python language, and the analysis results are saved to the S4 object and dynamically displayed through Echarts.

5. The single-cell multi-omics data analysis system according to claim 1, characterized in that The deep learning optimization module is specifically used to: Based on the Pytorch script or tensorflow script written by the user in the specified format, by calling the Python interpreter in the current environment, a communication bridge is built between the R language and the PyTorch interface or TensorFlow interface, the S4 object is converted to the Python language for deep learning model training, and the analysis results of the deep learning model on the S4 object are converted to the R language and integrated into the S4 object.

6. The single-cell multi-omics data analysis system according to claim 1, characterized in that The deep learning optimization module is also used to: It has built-in deep learning models for single-cell multi-omics sequencing data analysis, and can directly call built-in deep learning models through R language scripts to perform analysis on S4 objects, obtain the analysis results of the deep learning model on the S4 object and integrate them into the S4 object.

7. The single-cell multi-omics data analysis system according to claim 1, characterized in that The interaction and expansion module is specifically used for: Obtain the Pytorch script or tensorflow script written by the user in the specified format, and call the Python interpreter in the current environment based on the script path. Encapsulate the custom single-cell multi-omics analysis tool written by the user according to the standardized API interface into an independent plug-in, and dynamically load the independent plug-in into the S4 object for analysis through modular design.

8. A single-cell multi-omics data analysis method, characterized in that: include: Read sequencing data from different sequencing platforms through the data import module and save them as S4 objects; Through the cross-language analysis module, single-cell multi-omics analysis tools based on R and Python are called according to the analysis process within a unified framework to analyze S4 objects; Through the deep learning optimization module based on the cross-language interactive interface, the deep learning model is trained on the S4 object, and the analysis results of the deep learning model on the S4 object are integrated into the S4 object; Obtain analysis workflows and visualize analysis results through interactive and extended modules, as well as build custom single-cell multi-omics analysis tools in the form of plug-ins and load them into S4 objects for analysis.

9. A computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the computer program, the steps of the single-cell multi-omics data analysis method according to claim 8 are implemented.

10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the steps of the single-cell multi-omics data analysis method according to claim 8 are implemented.

Citation Information

Patent Citations

  • Spatial transcriptome data conversion method and system of cross-language platform

    CN115206439A

  • Cell annotation method based on single cell space transcriptome and related device

    CN120126546A

Cited By

  • Inplanatable decision-making system for single-cell multi-omics data integration analysis

    CN122117020A