A disease biomarker discovery method, system and medium

CN118588168BActive Publication Date: 2026-09-25SUZHOU UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410558349.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-05-08
Publication Date
2026-09-25
Estimated Expiration
2044-05-08

AI Technical Summary

Technical Problem

[0003]本发明的目的是提出一种疾病生物标志物发掘方法以解决无法在两种或以上不同条件下(病理/生理等)高通量分析异常变化糖基化糖蛋白,无法快速有效识别与疾病相关联糖基化的瓶颈

Benefits of technology

[0012]由于上述技术方案运用,本发明与现有技术相比具有下列优点:本申请通过应用搜索速度较快的MSFragger,缩小了糖蛋白数据库筛选的范围,再次使用GlycReSoft进行筛选则进一步缩小N-糖蛋白筛查范围,给Byonic提供靶向N-糖蛋白数据库,Byonic在蛋白质修饰糖基化这一方面可以把更多的具体信息输出,以此来得到潜在目的糖蛋白的修饰信息,输出的结果进行再一步分析与优化,并为下一步的程序分析提供重要基础参考,减轻了人工分析的负担,极大地提高临床实际应用的生物标志物的筛选流程,并通过综合应用不同的程序,增加了检测的精密度和准确度。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118588168B_ABST
    Figure CN118588168B_ABST
Patent Text Reader

Abstract

The application discloses a disease biomarker mining method, system and medium, and the disease biomarker mining method comprises the following steps: screening a large range of total protein library based on a plurality of bioinformatics tools to obtain a large range of screening results; analyzing the screening results based on a GDAS system program to screen a list of glycoproteins with significant changes; calculating the proportions of different glycopeptides of proteomes based on the list of glycoproteins, and narrowing the range of a protein database; analyzing and searching glycosylation type change information in diseases based on the protein database after the range is narrowed to obtain target glycoproteins; performing structure analysis on the screened target glycoproteins to obtain analysis results; and screening potential biomarkers in clinical and research through comprehensive use of a plurality of bioinformatics tools, so that the efficiency and specificity of screening are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of disease discovery technology, specifically to a method, system, and medium for discovering disease biomarkers. Background Technology

[0002] Screening for disease-related biomarkers through glycoprotein analysis is a crucial method for early disease diagnosis and prognosis. Protein glycosylation, the enzymatic attachment of sugar molecules (glycans) to proteins, has become a promising avenue for biomarker discovery. Glycan modifications of proteins affect their structure and physiological function; therefore, alterations in glycosylation patterns are often associated with disease occurrence and progression. Screening for glycoprotein biomarkers typically employs mass spectrometry, which involves analyzing and comparing the mass spectra of a large number of clinical samples with those of human proteomes and glycans using bioinformatics techniques to identify altered glycosylation patterns. However, the sheer size of human protein libraries and the limitations of current analytical software in accurately and rapidly analyzing high-throughput glycoprotein databases pose a significant challenge to biomarker discovery. Summary of the Invention

[0003] The purpose of this invention is to propose a method for discovering disease biomarkers to overcome the bottleneck of being unable to perform high-throughput analysis of abnormally changed glycosylated glycoproteins under two or more different conditions (pathological / physiological, etc.) and to quickly and effectively identify glycosylation associated with diseases.

[0004] To achieve the above-mentioned objectives, the technical solution adopted by this invention is: a method for discovering disease biomarkers, comprising the following steps: S1, based on a variety of bioinformatics tools, performs a large-scale screening of total protein libraries, resulting in a wide range of screening results; S2, based on the GDAS system program, is developed using Python. It sets different running modes and filters and analyzes the output results of other bioinformatics tools to obtain a list of glycoproteins with significant changes. S3 calculates the proportion of different glycopeptides in the proteome based on the glycoprotein list and narrows down the scope of the glycoprotein database; S4. Based on the narrowed-scope glycoprotein database analysis, information on changes in glycosylation types in the disease is found to obtain target glycoproteins. S5. Structural analysis was performed on the screened target glycoproteins to obtain the analysis results.

[0005] Furthermore, in step S1, various bioinformatics tools include MSFragger, GlycReSoft, O-Pair, or Byonic.

[0006] Furthermore, glycosylation types include N-glycosylation and O-glycosylation. GlycReSoft was used to screen N-glycosylated proteins, revealing disease-associated N-glycoproteins; O-Pair was used to screen O-glycosylated proteins, revealing disease-associated O-glycoproteins.

[0007] Furthermore, the algorithm analysis steps in step S2 are as follows: Read the output file, filter the output file, and group the output file according to the filtering results; Determine if the groups are duplicated; If the groups are not repeated, calculate the intensity ratio; If groups are duplicated, computer language is used to analyze and determine whether protein IDs are duplicated by setting up loops and conditional statements. Protein IDs are unique in the protein database; each identified protein corresponds to a single ID, establishing a mapping relationship. If protein IDs are duplicated, a unique data type is set for further division, and the total peak intensity of all ion peaks is calculated for a single mass spectrometry analysis result file. Then, the sum of the ion peaks of a single protein is calculated and divided by the total peak intensity of the file to obtain the intensity ratio. If the protein IDs are not duplicated, it is only necessary to calculate the overall peak intensity within the mass spectrometry file and obtain the intensity ratio through the corresponding calculation.

[0008] Furthermore, the algorithm steps for narrowing down the protein database in step S3 are as follows: Read the output file and the database, and determine whether the database ID in the output file meets the requirements; If the requirements are met, extract the rows from the database and create a new file with the same format and encoding as the original file; If the requirements are not met, skip the judgment.

[0009] Furthermore, step S5, "performing structural analysis on the screened target glycoproteins and obtaining the analysis results," specifically includes: Based on the screened target glycoproteins, the total glycopeptides and glycans were analyzed, and the site-specific glycoforms of biomarkers were analyzed by Byonic analysis. The site-specific glycoform analysis of biomarkers was used to analyze the structure of glycopeptides, glycosylation sites, and glycans, and the analysis results were obtained.

[0010] The present invention also provides a disease biomarker discovery system, including a processor, a memory, and at least one program, the program being stored in the memory and configured to be executed by the processor, the program including instructions for performing the disease biomarker discovery method as described in any of the preceding claims.

[0011] The present invention also provides a computer-readable storage medium storing a computer program that causes a computer to execute in order to implement the disease biomarker discovery method described in any of the preceding claims.

[0012] Due to the application of the above technical solutions, this invention has the following advantages compared with the prior art: This application narrows the scope of glycoprotein database screening by applying MSFragger, which has a faster search speed, and further narrows the scope of N-glycoprotein screening by using GlycReSoft again, providing Byonic with a targeted N-glycoprotein database. Byonic can output more specific information on protein modification glycosylation, thereby obtaining modification information of potential target glycoproteins. The output results are further analyzed and optimized, and provide an important basic reference for the next step of program analysis, reducing the burden of manual analysis, greatly improving the screening process of biomarkers in clinical applications, and increasing the precision and accuracy of detection by comprehensively applying different programs. Attached Figure Description

[0013] Figure 1 The diagram shows the overall algorithm block diagram of the disease biomarker discovery method provided in the embodiments of the present invention; Figure 2 The flowchart of the algorithm for shrinking the glycoprotein library provided in the embodiment of the present invention is shown; Figure 3 A flowchart of the analysis algorithm provided in an embodiment of the present invention is shown. Detailed Implementation

[0014] To enable those skilled in the art to better understand the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present application, and not all embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative effort should fall within the scope of protection of the present application.

[0015] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate for the embodiments of this application described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or system that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or systems.

[0016] It should be noted that, unless otherwise specified, the embodiments and features described in this application can be combined with each other. This application will now be described in detail with reference to the accompanying drawings and embodiments.

[0017] like Figures 1-3 As shown, this embodiment of the invention provides a method for discovering disease biomarkers, including the following steps: S1, based on a variety of bioinformatics tools, performs a large-scale screening of total protein libraries, resulting in a wide range of screening results; S2, based on the GDAS (Glycoproteomics Data Analysis Software) system program. This program is developed using the Python language, sets different running modes, and filters and analyzes the output results of other bioinformatics tools to obtain a list of glycoproteins with significant changes. S3 calculates the proportion of different glycopeptides in the proteome based on the glycoprotein list and narrows down the scope of the glycoprotein database; S4. Based on the narrowed-scope glycoprotein database analysis, information on changes in glycosylation types in the disease is found to obtain target glycoproteins. S5. Structural analysis was performed on the screened target glycoproteins to obtain the analysis results.

[0018] According to an embodiment of the present invention, the various bioinformatics tools used in step S1 include MSFragger, GlycReSoft, O-Pair, or Byonic.

[0019] Specifically, MSFragger represents ultra-fast and comprehensive peptide identification in shotgun proteomics; GlycReSoft represents a software package for automatically identifying glycans from LC / MS data; O-Pair is also known as MetaMorpheus; and Byonic is also known as Baiannik.

[0020] According to embodiments of the present invention, the glycosylation types include N-glycosylation and O-glycosylation. GlycReSoft was used to screen N-glycosylated proteins, revealing disease-associated N-glycoproteins; O-Pair was used to screen O-glycosylated proteins, revealing disease-associated O-glycoproteins.

[0021] According to an embodiment of the present invention, the analysis algorithm steps in step S2 are as follows: Read the output file, filter the output file, and group the output file according to the filtering results; Determine if the groups are duplicated; If the groups are not repeated, calculate the intensity ratio; If groups are duplicated, computer language is used to analyze and determine whether protein IDs are duplicated by setting up loops and conditional statements. Protein IDs are unique in the protein database; each identified protein corresponds to a single ID, establishing a mapping relationship. If protein IDs are duplicated, a unique data type is set for further division, and the total peak intensity of all ion peaks is calculated for a single mass spectrometry analysis result file. Then, the sum of the ion peaks of a single protein is calculated and divided by the total peak intensity of the file to obtain the intensity ratio. If the protein IDs are not duplicated, it is only necessary to calculate the overall peak intensity within the mass spectrometry file and obtain the intensity ratio through the corresponding calculation.

[0022] According to an embodiment of the present invention, the algorithm steps for narrowing down the glycoprotein database in step S3 are as follows: Read the output file and the database, and determine whether the database ID in the output file meets the requirements; If the requirements are met, extract the rows from the database and create a new file with the same original format and encoding. If the requirements are not met, skip the judgment.

[0023] According to an embodiment of the present invention, step S5, "performing structural analysis on the screened target glycoproteins and obtaining the analysis results," specifically includes: Based on the screened target glycoproteins, the total glycopeptides and glycans were analyzed, and the site-specific glycoforms of biomarkers were analyzed by Byonic analysis. The site-specific glycoform analysis of biomarkers was used to resolve glycopeptides, glycosylation sites, and glycan structures, and the analysis results were obtained.

[0024] In summary, this method first uses mass spectrometry to identify glycoproteins in the human proteome using software capable of rapid big data analysis, quickly screening out a glycoprotein library. Next, by comparing the entire mass spectrometry data, the glycoprotein library obtained in the first step is screened to identify glycoproteins that exhibit abnormal changes in disease, including N-linked and O-linked glycoproteins, thus obtaining a glycoprotein library closely related to the disease. Finally, this database is analyzed using software to obtain detailed glycosylation changes observed in the disease, thereby establishing potential glycoprotein biomarkers and providing a foundation for subsequent validation. This technology provides a solution for large-scale protein databases and overcomes the limitations of existing analytical tools. By combining database establishment, initial screening of N-glycoproteins and O-glycoproteins using different analytical tools, and annotation of disease-specific glycosylation, it achieves efficient and specific discovery of disease biomarkers.

[0025] This embodiment also provides a disease biomarker discovery system, including a processor, a memory, and at least one program. The program is stored in the memory and configured to be executed by the processor. The program includes instructions for performing any of the above-described disease biomarker discovery methods.

[0026] Those skilled in the art will understand that, for ease of explanation, the example is provided with one memory and one processor. In actual terminals or servers, multiple processors and memories may exist. Memory can also be referred to as storage medium or storage device, etc., and the embodiments of this application do not limit this.

[0027] It should be understood that in the embodiments of this application, the processor may be a central processing unit (CPU), or it may be other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The processor may also be a general-purpose microprocessor, graphics processing unit (GPU), or one or more integrated circuits to execute relevant programs to achieve the functions required by the embodiments of this application.

[0028] The processor can also be an integrated circuit chip with signal processing capabilities. In implementation, each step of this application can be completed through integrated logic circuits in the processor hardware or instructions in software form. The aforementioned processor can implement or execute the methods, steps, and logic block diagrams disclosed in the embodiments of this application. The steps of the methods disclosed in the embodiments of this application can be directly manifested as execution by a hardware decoding processor, or execution by a combination of hardware and software modules in the decoding processor. The software modules can reside in random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, registers, or other mature storage media in the art. This storage medium is located in memory; the processor reads information from the memory and, in conjunction with its hardware, completes the functions required by the units included in the methods, systems, and storage media of the embodiments of this application.

[0029] It should also be understood that the memory mentioned in the embodiments of this application can be volatile memory or non-volatile memory, or may include both volatile and non-volatile memory. The non-volatile memory can be read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), or flash memory. The volatile memory can be random access memory (RAM), which is used as an external cache.

[0030] By way of example, but not limitation, many forms of RAM are available, such as static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDR SDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronous link dynamic random access memory (SLDRAM), and direct rambus RAM (DR RAM).

[0031] The memory can also be a read-only optical disc (CD-ROM) or other optical disc storage, optical disc storage (including compressed optical discs, laser discs, optical discs, digital universal optical discs, Blu-ray discs, etc.), magnetic disk storage media, or other magnetic storage devices, or any other medium capable of carrying or storing desired program code in the form of instructions or data structures that can be accessed by a computer, but is not limited thereto. The memory can exist independently and be connected to the processor via a bus. The memory can also be integrated with the processor. The memory can store programs, and when the program stored in the memory is executed by the processor, the processor performs the various steps of the method determined in the above embodiments of this application.

[0032] It should be noted that when the processor is a general-purpose processor, DSP, ASIC, FPGA, or other programmable logic device, discrete gate or transistor logic device, or discrete hardware component, the memory (storage module) is integrated into the processor. It should be noted that the memory described herein is intended to include, but is not limited to, these and any other suitable types of memory.

[0033] It should be understood that the term "and / or" in this article is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A existing alone, A and B existing simultaneously, or B existing alone. Additionally, the character " / " in this article generally indicates that the preceding and following related objects have an "or" relationship.

[0034] In implementation, each step of the above method can be completed by integrated logic circuits in the processor's hardware or by instructions in software. The steps of the method disclosed in the embodiments of this application can be directly implemented by a hardware processor, or by a combination of hardware and software modules within the processor. The software modules can reside in mature storage media in the art, such as random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, or registers. Since this storage medium is located in memory, the processor reads information from the memory and, in conjunction with its hardware, completes the steps of the above method; to avoid repetition, these will not be described in detail here.

[0035] Those skilled in the art will recognize that the various illustrative logical blocks (ILBs) and steps described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementations should not be considered beyond the scope of this application.

[0036] In the above embodiments, implementation can be achieved, in whole or in part, through software, hardware, firmware, or any combination thereof. When implemented in software, it can be implemented, in whole or in part, as a computer-programmed program product. A computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a processor, all or part of the flow or function according to the embodiments of this application is generated. The computer can be a general-purpose computer, a computer network, or other programmable device.

[0037] This embodiment also provides a computer-readable storage medium storing a computer program that causes a computer to execute in order to implement the above-described method for discovering disease biomarkers.

[0038] It should be noted that computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, computer instructions can be transmitted from one website, computer, server, or data center to another via wired (e.g., coaxial cable, fiber optic) or wireless (e.g., infrared, wireless, microwave, etc.) means, or from one website, computer, server, or data center to a mobile phone processor via a wired means. A computer-readable storage medium can be any usable medium that a computer can access, or a data storage system such as a server or data center that integrates one or more usable media. Usable media can be magnetic media (e.g., floppy disks, hard disks), optical media (e.g., DVDs), or semiconductor media (e.g., solid-state drives), etc.

[0039] Finally, it should be noted that the above are merely preferred embodiments of the present invention and are not intended to limit the present invention. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art can still modify the technical solutions described in the foregoing embodiments or make equivalent substitutions for some of the technical features. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.

Claims

1. A method for discovering disease biomarkers, characterized in that, Includes the following steps: S1, based on a variety of bioinformatics tools, performs a large-scale total protein library screening to obtain a wide range of glycoprotein screening results; S2 uses the GDAS system program, which is developed based on the Python language. It sets different running modes and filters and analyzes the output results of other bioinformatics tools to obtain a list of glycoproteins with significant changes. S3, calculates the content ratio of different glycopeptides in the proteome based on the glycoprotein list, and narrows down the scope of the glycoprotein database; S4. Based on the narrowed-scope glycoprotein database analysis, information on changes in glycosylation types in the disease is found to obtain target glycoproteins. S5, perform structural analysis on the screened target glycoproteins to obtain the analysis results; Step S1 involves various bioinformatics tools, including MSFragger, GlycReSoft, O-Pair, or Byonic. Glycosylation types include N-glycosylation and O-glycosylation; GlycReSoft was used to screen N-glycosylated proteins to identify N-glycoproteins with significant changes that were associated with the disease; at the same time, O-Pair was used to screen O-glycosylated proteins to obtain O-glycoproteins associated with the disease. The algorithm analysis steps in step S2 are as follows: Read the output file, filter the output file, and group the output file according to the filtering results; Determine if the groups are duplicated; If the groups are not repeated, calculate the ratio of glycopeptide or polysaccharide content; If groups are duplicated, computer language is used to analyze and determine whether protein IDs are duplicated by setting up loop paragraphs and conditional statements; protein IDs are unique in the protein database, and each identified protein corresponds to one ID, with a mapping relationship; If protein IDs are duplicated, a unique data type is set for further division, and the total peak intensity of all ion peaks is calculated for a single mass spectrometry analysis result file. Then, the sum of the ion peaks of a single protein is calculated and divided by the total peak intensity of the file to obtain the intensity ratio. If the protein IDs are not duplicated, it is only necessary to calculate the overall peak intensity within the mass spectrometry file and obtain the intensity ratio through the corresponding calculation.

2. The method for discovering disease biomarkers as described in claim 1, characterized in that, The algorithm steps for narrowing down the protein database in step S3 are as follows: Read the output file and the database, and determine whether the other information corresponding to the protein ID in the output file meets the requirements. If the requirements are met, the row corresponding to the ID in the database is extracted, which contains the complete amino acid sequence information of the corresponding protein, and a new file with the same format as the original database file is created. If the requirements are not met, skip the judgment.

3. The method for discovering disease biomarkers as described in claim 2, characterized in that, Step S5, "performing structural analysis on the screened target glycoproteins and obtaining the analysis results," specifically includes: Based on the screened target glycoproteins, the total glycopeptides and glycans were analyzed, and the site-specific glycoforms of biomarkers were analyzed by Byonic analysis. The site-specific glycoform analysis of biomarkers was used to resolve glycopeptides, glycosylation sites, and glycan structures, and the analysis results were obtained.

4. A disease biomarker discovery system, characterized in that, The device includes a processor, a memory, and at least one program stored in the memory and configured to be executed by the processor, the program including instructions for performing the disease biomarker discovery method as described in any one of claims 1-3.

5. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that causes a computer to execute in order to implement the disease biomarker discovery method according to any one of claims 1-3.