System and method for integrated data processing for lc-ms / ms data analysis

KR103022339B1Active Publication Date: 2026-09-21BERTIS INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
KR1020250150342
Authority / Receiving Office
KR · KR
Patent Type
Patents
Current Assignee / Owner
Filing Date
2025-10-17
Publication Date
2026-09-21
Estimated Expiration
2045-10-17

Smart Images

  • Figure R1020250150342_ABST
    Figure R1020250150342_ABST
Patent Text Reader

Abstract

According to the present disclosure, an integrated data processing system and method for LC-MS / MS data analysis are provided. The method may include the steps of: importing search result data from a protein search engine; performing quality control (QC) on the search result data; performing protein inference based on the search result data; performing hierarchical summarization on the protein inference results; and generating analysis result data including the QC results, the protein inference results, and the hierarchical summarization results. According to the present disclosure, an integrated data processing solution can be provided that implements the entire process of LC-MS / MS data analysis, which was difficult to integrate and standardize due to fragmented analysis processes, into a unified workflow.
Need to check novelty before this filing date? Find Prior Art

Description

Technology Field

[0001] The present disclosure relates to mass spectrometry-based proteomics, and more specifically to an integrated data processing system and method for LC-MS / MS data analysis. Background Technology

[0002] In Mass Spectrometry (MS)-based proteomics, data analysis generally involves the processes of data collection, database search, result refinement, and reporting.

[0003] However, existing representative tools are developed to be limited to specific analysis stages, resulting in fragmented steps across tools with differing formats and rules. Consequently, as there is currently no solution capable of integrated, end-to-end proteomic analysis, there is an urgent need for the development of a system that enables proteomic data analysis through a unified workflow. The problem to be solved

[0004] The present disclosure aims to solve these problems by providing an integrated data processing system and method for LC-MS / MS data analysis. means of solving the problem

[0005] According to one embodiment of the present disclosure, an integrated data processing method for LC-MS / MS data analysis executable by a computing device may be provided. The method may include the steps of: importing identification data from a protein search engine; performing quality control (QC) on the identification data; performing protein inference based on the identification data; performing hierarchical summarization on the protein inference results; and generating analysis result data including the QC results, the protein inference results, and the hierarchical summarization results.

[0006] Additionally, the step of importing the data may include the step of combining the search result data with experimental metadata to generate imported search result data in a single object format.

[0007] Additionally, the step of performing the above QC may include the step of deriving a value for at least one item among QC indicators including q-value, precursor isolation purity, peptide length, charge, missed cleavage, intensity distribution, Principal Component Analysis (PCA), and correlation.

[0008] In addition, the step of performing the protein inference may include the step of classifying protein groups based on the parsimony rule method in the search result data.

[0009] In addition, the step of performing the hierarchical summary may include the step of performing independent quantitative summaries for PSM (Peptide Spectrum Match), peptides, and proteins, respectively, from the protein inference results.

[0010] In addition, the step of performing the above hierarchical summary may further include the step of performing an independent quantitative summary for post-translational modification (PTM) sites.

[0011] Additionally, the step of generating the analysis result data may include the step of generating and storing MuData in a single object format containing the analysis result data.

[0012] Additionally, the above method may further include the step of performing a statistical analysis including a permutation test on the analysis result data; and the step of performing a visualization on the analysis result data.

[0013] According to one embodiment of the present disclosure, an integrated data processing system for LC-MS / MS data analysis may be provided. The system may include: a data import module configured to import search result data of a protein search engine; a QC module configured to perform quality control (QC) on the search result data; a protein inference module configured to perform protein inference based on the search result data; a hierarchical summary module configured to perform hierarchical summarization on the protein inference results; and a result generation module configured to generate analysis result data including the QC results, the protein inference results, and the hierarchical summary results.

[0014] According to one embodiment of the present disclosure, a computer program stored on a computer-readable medium may be provided, comprising computer-executable instructions for executing an integrated data processing method for LC-MS / MS data analysis. Effects of the invention

[0015] According to the present disclosure, an integrated data processing solution can be provided that implements the entire process of LC-MS / MS data analysis, which was difficult to integrate and standardize due to fragmented analysis processes, into a unified workflow.

[0016] In addition, according to the present disclosure, the import of result data from various protein search engines can be automated to resolve manual dependency and improve reproducibility, and the objectivity and consistency of quality control can be ensured by providing integrated key QC indicators, and reliability can be improved by finally confirming all protein groups in a distinguishable state through protein inference, and hierarchical summarization can perform quantitative summarization independently of inference and ensure data stability based on top N features, and statistical reliability can be ensured and various additional tests can be supported based on Welch T-test-based permutation tests.

[0017] In addition, according to the present disclosure, mass spectrometry data and QC results can be stored collectively in a single object MuData with an h5mu (HDF5) structure to create a standardized mass spectrometry data structure, thereby facilitating data sharing among researchers and external verification. Brief explanation of the drawing

[0018] FIG. 1 is an exemplary block diagram illustrating a computing device for integrated data processing for LC-MS / MS data analysis according to one embodiment of the present disclosure. FIG. 2 is an exemplary block diagram illustrating an integrated data processing system for LC-MS / MS data analysis according to one embodiment of the present disclosure. FIG. 3 is an exemplary flowchart illustrating an integrated data processing method for LC-MS / MS data analysis according to one embodiment of the present disclosure. FIG. 4 is a diagram illustrating an exemplary workflow for integrated data processing for LC-MS / MS data analysis according to one embodiment of the present disclosure. Specific details for implementing the invention

[0019] Hereinafter, preferred embodiments of the present invention will be described in detail with reference to the accompanying drawings. First, it should be noted that in assigning reference numerals to the components of each drawing, the same components are given the same reference numeral whenever possible, even if they are shown in different drawings. Furthermore, in describing the present invention, if it is determined that a detailed description of related known components or functions could obscure the essence of the invention, such detailed description is omitted.

[0020] Various aspects of the present invention are described below. It should be understood that the inventions presented herein may be embodied in a wide variety of forms, and that any specific structure, function, or all thereof presented herein are merely illustrative. Based on the inventions presented herein, those skilled in the art will understand that any one aspect presented herein may be embodied independently of any other aspects, and that two or more such aspects may be combined in various ways. For example, an apparatus may be embodied or a method may be practiced using any number of aspects described herein. Furthermore, such an apparatus may be embodied or such a method may be practiced using structures, functions, or structures and functions other than those described herein, in addition to or other than these aspects.

[0021] Various analysis tools are used for mass spectrometry-based proteomics (LC-MS / MS) data analysis. However, in existing tools, the data analysis process is fragmented by tool, and differing formats and rules lead to the following problems.

[0022] - Because output formats and metadata combination rules vary among database (DB) search tools, the import or ingestion of search result data is dependent on manual work.

[0023] - Reproducibility is reduced because the calculation and visualization of QC metrics are not standardized, leading to varying criteria among researchers.

[0024] - Normalization / correction techniques and procedures are being applied inconsistently.

[0025] - It is difficult to trace the provenance of mass spectrometry data during the hierarchical summarization process, and results vary depending on the researcher's choice.

[0026] - The rules for generating shared peptides and protein groups are unclear because the protein inference process is not integrated.

[0027] - As result analysis and statistical tools are not integrated, data compatibility is poor and unnecessary work is required.

[0028] - QC, statistics, and visualization results are not stored in an integrated manner, making sharing and verification difficult and resulting in fragmented reports.

[0029] Specifically, the representative tools currently in use have the following problems.

[0030] - MaxQuant: It is used for DDA-based analysis, but its format is complex and it lacks compatibility.

[0031] - DIA-NN: Used for DIA-based analysis, but lacks flexibility in QC and normalization.

[0032] - MSstats: Specialized in statistical analysis, but lacks data import and QC functions.

[0033] - OpenMS: It is a general-purpose platform, but it is complex and has a high learning curve.

[0034] - Scanpy: It is transcriptomics-focused and lacks specialized proteomics features.

[0035] As such, existing tools are limited to specific stages, and there is currently no end-to-end integrated proteomics analysis tool. Accordingly, the present disclosure aims to present an integrated data processing solution that implements the entire process of mass spectrometry-based proteomics (LC-MS / MS) data analysis into a unified workflow.

[0036] FIG. 1 is an exemplary block diagram illustrating a computing device for integrated data processing for LC-MS / MS data analysis according to one embodiment of the present disclosure.

[0037] As illustrated in FIG. 1, the computing device (100) may include a processor (110), a storage medium (120), a memory (130), and a network interface (140), which may be connected to each other via a system bus (150).

[0038] An operating system (OS) (122) and a computer program (124) may be mounted on the storage medium (120). The storage medium (120) may be a data storage device such as a hard disk, an SSD (Solid State Drive), etc., capable of storing computer programs and related data. The operating system (122) may be operating system software such as Windows, iOS, Linux, etc., for operating the computing device (100). The computer program (124) may include functional modules for integrated data processing for LC-MS / MS data analysis according to the present disclosure, as well as computer-executable instructions for this purpose. Additionally, the computer program (124) may be loaded into memory (130) so that it can be executed by the processor (110). When the computer-executable instructions of the computer program (124) are executed by the processor (110), they may cause the processor (110) to perform the integrated data processing method for LC-MS / MS data analysis according to the present disclosure. The processor (110) may be configured to provide computing and control capabilities to support the execution of the entire computing device (100). The processor (110) may be a data processing device such as a CPU (Central Processing Unit), MPU (Microprocessor Unit), AP (Application Processor), etc., and may be composed of one processor or multiple processors. If composed of multiple processors, the processors (110) may operate as parallel processing processors. The network interface (140) may provide an interface that can communicate data by connecting to an external device (e.g., a database storing proteomics-related data, a display device, another wired or wireless communication device connectable via a network, etc.).

[0039] FIG. 2 is an exemplary block diagram illustrating an integrated data processing system for LC-MS / MS data analysis according to one embodiment of the present disclosure.

[0040] The integrated data processing system (200) for LC-MS / MS data analysis illustrated in FIG. 2 can be implemented through the computing device (100) illustrated in FIG. 1. This system (200) may include a data import module (210), a quality control (QC) module (220), a protein inference module (230), a hierarchical summarization module (240), a normalization / correction module (250), a result generation module (260), a statistical analysis module (270), and a visualization module (280).

[0041] The data import module (210) may be configured to import identification data from various protein search engines (e.g., SAGE, DIA-NN, etc.). The identification data may refer to data for which identification has been completed through a database (DB) search on mass spectrometry (MS) data obtained for a sample to be analyzed. The data import module (210) may combine this identification data with experimental metadata to generate imported identification data in a single object format (e.g., h5mu), and the imported identification data may be utilized for subsequent analysis of the system (200).

[0042] The quality control (QC) module (220) may be configured to perform QC on the search result data. QC may refer to a process of analyzing various items on the search result data to obtain accurate and reliable analysis results, and selectively excluding data that may lower the reliability of the analysis results based on this analysis. The QC module (220) may be configured to derive values ​​for at least one item among QC indicators including q-value, precursor isolation purity, peptide length, charge, missed cleavage, intensity distribution, Principal Component Analysis (PCA), and correlation.

[0043] The protein inference module (230) may be configured to perform protein inference based on search result data. Protein inference may be a process of predicting proteins present in a sample to be analyzed using identified peptide sequences. The protein inference module (230) may be configured to classify protein groups based on the parsimony rule method in the search result data.

[0044] The hierarchical summary module (240) may be configured to perform a hierarchical summary on the protein inference results. The hierarchical summary may include summaries at least at the feature (i.e., PSM (Peptide Spectrum Match)) level, peptide level, and protein level. The hierarchical summary module (240) may be configured to perform independent quantitative summaries for PSM, peptide, and protein, respectively, from the protein inference results. Additionally, the hierarchical summary may include summaries at the post-translational modification (PTM) level, and to this end, the hierarchical summary module (240) may be configured to perform independent quantitative summaries for PTM sites.

[0045] The normalization / correction module (250) may be configured to perform normalization and / or batch correction on the analysis data.

[0046] The result generation module (260) may be configured to generate analysis result data including QC results, protein inference results, and hierarchical summary results. Additionally, the result generation module (260) may be configured to generate and store MuData in a single object format containing the generated analysis result data. This MuData has a standardized HDF5-based data structure containing multiple objects (corresponding to each modality), thereby facilitating data sharing and external verification.

[0047] The statistical analysis module (270) can be configured to perform various statistical analyses, including a permutation test on the analysis result data.

[0048] The visualization module (280) can be configured to provide various visualization functions for the analysis result data.

[0049] FIG. 3 is an exemplary flowchart illustrating an integrated data processing method for LC-MS / MS data analysis according to one embodiment of the present disclosure.

[0050] A computing device (100) can import search result data from a protein search engine (310). The computing device (100) can perform quality control (QC) on the imported search result data (320). The computing device (100) can perform protein inference based on the search result data (330). The computing device (100) can perform hierarchical summarization on the protein inference results (340). The computing device (100) can perform statistical analysis and visualization on the analysis results (350). The computing device (100) can generate analysis result data including these QC results, protein inference results, and hierarchical summary results, and store and manage them as a single object Mudata (360). A more detailed description of each step of this method will be provided later in relation to FIG. 4.

[0051] FIG. 4 is a diagram illustrating an exemplary workflow for integrated data processing for LC-MS / MS data analysis according to one embodiment of the present disclosure.

[0052] As illustrated in FIG. 4, experimental mass spectrometry raw data (402) can be generated for a sample to be analyzed by a mass spectrometry instrument (401). This raw data (402) can be converted (403) into a data format (e.g., mzml) for searching through a protein database (DB). A protein search engine (e.g., SAGE, DIA-NN, etc.) can perform identification work by searching the protein DB for this experimental mass spectrometry data and, as a result, generate search result data (404).

[0053] The data import module (210) can import search result data from a protein search engine, and the resulting search result data can be converted into a standardized data structure for subsequent data processing and input into the system (200). The data import module (210) can be configured to recognize the directory structure and filename pattern of the search result data (404) and apply import logic specific to the search engine. Additionally, the data import module (210) can generate imported search result data by combining the search result data (404) with experimental metadata (405). The experimental metadata (405) may include experimental parameters such as enzyme, modification (PTM) information, and LC / MS settings.

[0054] For example, when importing search result data from protein search engines SAGE and DIA-NN, the data import module (210) can automatically explore and parse their search result directories, and SAGE’s Protein.Group and DIA-NN’s Protein.Names can be mapped to protein.group. Additionally, SAGE’s *.tsv, DIA-NN’s report.tsv, and related output files can be automatically recognized. Furthermore, experimental metadata (405) can be matched with an annotation file (tsv / json) and combined with the search result data.

[0055] The data import module (210) can generate the imported search result data into a single object format (e.g., h5mu) and transmit it to subsequent data processing (QC, protein inference, hierarchical summary, etc.). This standardized data structure is extensible and can be extended to MaxQuant, pepXML, mxTab, etc.

[0056] The QC module (220) can perform quality control (QC) on the input search result data to derive values ​​for QC indicators. QC indicators can be defined as indicators related to proteomics. QC indicators may include, but are not limited to, at least q-values, precursor isolation purity, peptide length, charge, lost cleavage, intensity distribution, PCA, and correlation, and QC may be performed on additional QC indicators (e.g., retention time (RT) versus peptide, etc.). The QC indicators calculated through the QC module (220) can be visualized individually or comprehensively as graphs, charts, etc. (e.g., through the visualization module (280)).

[0057] The protein inference module (230) can perform protein inference based on search result data. The protein inference module (230) can classify protein groups based on the parsimony rule method in the search result data and can be configured based on unique peptide evidence excluding shared peptides. Additionally, the protein inference module (230) can reclassify protein groups into Identical, Subset, Subsumable, and Distinguishable relationships and can perform protein inference so that all protein groups are finally confirmed to be in a Distinguishable state.

[0058] The hierarchical summary module (240) can perform hierarchical summarization on the protein inference results. The hierarchical summary module (240) can perform quantitative summarization at the PSM (precursor) -> peptide -> protein stages, and this can be performed as an independent quantitative summarization separate from the protein inference. Additionally, the hierarchical summary module (240) can perform summarization based on top N features at each stage. The application history of this hierarchical summarization can be recorded as provenance so that it can be traced. Additionally, the hierarchical summary module (240) can be configured to also perform independent quantitative summarization on PTM sites.

[0059] The normalization / correction module (250) can perform normalization and / or batch correction on the protein inference results. To this end, the normalization / correction module (250) can support log2 transform, median centering, quantile normalization, and GIS batch correction. Additionally, the history of such normalization / correction application can be recorded as a source so that it can be traced. Also, in the example of FIG. 4, the normalization / correction module (250) is illustrated as being between the protein inference module (230) and the hierarchical summary module (240), but the normalization / correction module (250) may be placed after the hierarchical summary module (240). In this case, the normalization / correction module (250) may be configured to apply normalization / correction at the PSM, peptide, and protein stages.

[0060] The result generation module (260) can generate analysis result data that includes both QC results and data analysis results, and can generate and store MuData (406) in a single object format that includes such analysis result data.

[0061] The statistical analysis module (270) can perform statistical analysis on the analysis result data and can provide a permutation test based on the Welch t-test as a basic test. Additionally, the statistical analysis module (270) can be configured to selectively support other commonly used test techniques (Student's T-test, Wilcoxon rank-sum, median difference test, etc.). Additionally, the statistical analysis module (270) can provide cyclic-based q-value estimation as FDR (False Discovery Rate) correction and can selectively support additional techniques such as the Benjamin-Hochberg method.

[0062] The visualization module (270) can perform visualization of analysis result data and statistical analysis. To this end, the visualization module (270) can provide various visualization functions such as PCA, UMAP (Uniform Manifold Approximation & Projection), volcano plot, etc.

[0063] As described above, the integrated data processing system (200) according to the present disclosure can provide an integrated data processing solution that implements the entire process of LC-MS / MS data analysis into a unified workflow.

[0064] The differentiating features of the integrated data processing solution according to the present disclosure compared to existing technology are summarized as follows:

[0065]

[0066] It should be understood that any particular order or hierarchy of steps in any of the presented processes is an example of exemplary approaches. It should be understood that, based on design priorities, any particular order or hierarchy of steps in the processes may be rearranged within the scope of the invention. The appended method claims provide elements of various steps in an exemplary order, but do not imply limitation to the particular order or hierarchy presented.

[0067] As used herein, terms such as “component,” “unit (or part),” “module,” “system,” etc., may refer to computer-related entities, hardware, firmware, software, combinations of software and hardware, or executions of software. For example, a component may be a process, processor, object, execution thread, program, and / or computer executed on a processor, but is not limited thereto. For example, both an application executed on a computing device and the computing device itself may be a component. One or more components may reside within a processor and / or execution thread, and a component may be localized within a single computer or distributed among two or more computers. Additionally, these components may be executed from various computer-readable media having various data structures stored therein.

[0068] The description of the presented embodiments is provided so that any person skilled in the art may use or practice the present invention. Various modifications to these embodiments will be apparent to those skilled in the art, and the general principles defined herein may be applied to other embodiments without departing from the scope of the present invention. Thus, the present invention is not limited to the embodiments presented herein, but should be interpreted in the broadest possible scope consistent with the principles and novel features presented herein. Explanation of the symbols

[0069] 100: Computing device 110: Processor 120: Storage media 122: Operating System 124: Computer program 130: Memory 140: Network Interface 150: System bus 200: Integrated Data Processing System 210: Data Import Module 220: QC Module 230: Protein Inference Module 240: Hierarchical Summary Module 250: Normalization / Correction Module 260: Result generation module 270: Statistical Analysis Module 280: Visualization Module

Claims

Claim 1 An integrated data processing method for LC-MS / MS data analysis executable by a computing device, comprising: a step of importing identification data from a protein search engine; a step of performing quality control (QC) on the identification data; a step of performing protein inference based on the identification data; a step of performing hierarchical summarization on the protein inference results; and a step of generating analysis result data including the QC results, the protein inference results, and the hierarchical summarization results, wherein the step of generating analysis result data includes the step of generating and storing MuData in a single object format containing the analysis result data. Claim 2 An integrated data processing method for LC-MS / MS data analysis, wherein the step of importing data comprises combining the search result data with experimental metadata to generate data-imported search result data in a single object format. Claim 3 An integrated data processing method for LC-MS / MS data analysis, wherein the step of performing the QC comprises deriving a value for at least one item among QC indicators including q-value, precursor isolation purity, peptide length, charge, missed cleavage, intensity distribution, Principal Component Analysis (PCA), and correlation. Claim 4 An integrated data processing method for LC-MS / MS data analysis, wherein the step of performing protein inference in claim 1 includes the step of classifying protein groups based on the parsimony rule method in the search result data. Claim 5 An integrated data processing method for LC-MS / MS data analysis, wherein the step of performing the hierarchical summary according to claim 1 comprises the step of performing independent quantitative summaries for PSM (Peptide Spectrum Match), peptide, and protein, respectively, from the protein inference results. Claim 6 An integrated data processing method for LC-MS / MS data analysis, wherein the step of performing the hierarchical summary in claim 5 further includes the step of performing an independent quantitative summary for post-translational modification (PTM) sites. Claim 7 delete Claim 8 An integrated data processing method for LC-MS / MS data analysis, further comprising: a step of performing a statistical analysis including a permutation test on the analysis result data in claim 1; and a step of performing visualization on the analysis result data. Claim 9 An integrated data processing system for LC-MS / MS data analysis, comprising: a data import module configured to import search result data of a protein search engine; a QC module configured to perform quality control (QC) on the search result data; a protein inference module configured to perform protein inference based on the search result data; a hierarchical summary module configured to perform hierarchical summarization on the protein inference results; and a result generation module configured to generate analysis result data including the QC results, the protein inference results, and the hierarchical summary results, wherein the result generation module is additionally configured to generate and store MuData in a single object format containing the analysis result data. Claim 10 An integrated data processing system for LC-MS / MS data analysis, wherein the data import module is additionally configured to combine the search result data with experimental metadata to generate data-imported search result data in a single object format. Claim 11 An integrated data processing system for LC-MS / MS data analysis, wherein the QC module is additionally configured to derive values ​​for at least one item among QC indicators including q-value, precursor separation purity, peptide length, charge, lost cleavage, intensity distribution, PCA, and correlation. Claim 12 In claim 9, the protein inference module is additionally configured to classify protein groups based on the parsimony rule method in the search result data, an integrated data processing system for LC-MS / MS data analysis. Claim 13 An integrated data processing system for LC-MS / MS data analysis, wherein the hierarchical summary module is additionally configured to perform independent quantitative summaries for PSM, peptide, and protein, respectively, from the protein inference results. Claim 14 In claim 13, the above hierarchical summary module is additionally configured to perform independent quantitative summaries for PTM sites, an integrated data processing system for LC-MS / MS data analysis. Claim 15 delete Claim 16 An integrated data processing system for LC-MS / MS data analysis, further comprising: a statistical analysis module configured to perform statistical analysis including permutation testing on the analysis result data in claim 9; and a visualization module configured to perform visualization on the analysis result data. Claim 17 A computer program stored on a computer-readable medium, comprising computer-executable instructions for executing a method according to any one of claims 1 through 6 and 8.