A method for pathogen immunogram lineage evaluation based on high-throughput antibody epitope profile detection data
By identifying significantly enriched epitope peptides using a generalized Poisson distribution model, the problems of data analysis lag and fixed thresholds in high-throughput antibody detection technology were solved, enabling efficient and accurate assessment of pathogen immune lineages and improving the sensitivity and adaptability of detection.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-17
- Publication Date
- 2026-04-10
AI Technical Summary
Existing high-throughput antibody epitope detection technologies suffer from data analysis lag, making it difficult to comprehensively assess an individual's overall immune status. Furthermore, fixed threshold judgments lead to insufficient sensitivity and misjudgments, and they cannot adapt to changes in different samples and pathogens.
A pathogen immune lineage assessment method based on a generalized Poisson distribution model was adopted. Through data preprocessing, comparison and counting, combined with the generalized Poisson distribution model, significantly enriched epitope peptides were identified, and calibration coefficients were calculated to dynamically determine the immune response of pathogens.
It enables efficient processing of information from tens of thousands of viral sites, improves the sensitivity and specificity of pathogen immunity detection, is highly adaptable, and can support dynamic assessment of population pathogens and analysis of vaccine efficacy.
Smart Images

Figure CN120954506B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of immunology, and in particular to a pathogen immunological lineage evaluation method based on high-throughput antibody epitope spectrum detection data. BACKGROUND
[0002] Traditional pathogenic microorganism detection methods are usually limited to the current infection status of one or a small number of predefined pathogens, and it is difficult to achieve comprehensive characterization of unknown or multiple pathogen past infection histories. With the rise of high-throughput antibody epitope detection technology, it is possible to simultaneously screen hundreds of pathogen-related antigen epitopes.
[0003] Existing high-throughput antibody epitope detection technologies (such as PhIP-seq, PepSeq, VirScan, SLISA) can simultaneously screen thousands of pathogen-related antigen epitopes, providing new possibilities for large-scale pathogen immunity detection. However, the lagging nature of current data analysis methods limits their practical application value, such as the lack of efficient and systematic data analysis methods, making it difficult to statistically process large-scale data, resulting in difficulty in interpreting detection results. In practice, single-site antibody detection is still the main method, which cannot comprehensively evaluate the overall immune status of individuals. Moreover, existing methods rely too much on fixed threshold values to determine antibody reaction positivity or negativity, but fixed threshold values cannot balance the sensitivity and specificity of detection, resulting in insufficient sensitivity; especially for some pathogens, the proportion of peptide coverage in the library is low, and fixed threshold values are prone to missed detection; in addition, fixed threshold values cannot effectively correct the background noise generated by protein non-specificity, which may lead to misjudgment. The analysis results and the conclusion of whether the pathogen is infected or not are difficult to adapt to changes in different samples, different pathogens, and population immunity.
[0004] Therefore, there is an urgent need for a flexible and accurate data analysis method to analyze high-throughput pathogen immunological lineage, thereby providing important data support for disease diagnosis, vaccine development, and epidemiological research. SUMMARY
[0005] The purpose of the present application is to solve the problems in the background art and provide a pathogen immunological lineage evaluation method based on high-throughput antibody spectrum detection data.
[0006] To this end, the present application adopts the following technical solutions:
[0007] In a first aspect, the present application provides a pathogen immunological lineage evaluation method based on high-throughput antibody epitope spectrum detection data, comprising the following steps:
[0008] (1) Data preparation step: obtaining sequencing data of an antigen epitope library and sequencing data of an immunoprecipitation sample;
[0009] (2) Data preprocessing step: quality control of all sequencing data in step (1), including removing adapter sequences at both ends of read data and filtering read data with too short length;
[0010] (3) Data alignment and counting step: based on the sequencing data of all antigen epitope libraries, a reference library index is constructed; the filtered read data in step (2) is aligned with the reference library index, and the read count of each pathogen epitope in the antigen epitope library and the read count of each pathogen epitope in the immunoprecipitation sample are counted;
[0011] (4) Epitope peptide identification step: based on the generalized Poisson distribution model to identify significantly enriched epitope peptides in the immunoprecipitation sample;
[0012] (5) Calculate calibration coefficient step: for pathogen i, the corresponding calibration coefficient ; is the total number of peptide segments of pathogen i, N is the total number of positive epitope peptides detected for all pathogens, i.e. the number of significantly enriched epitopes, and M is the total number of antigen epitope peptide segments in the library;
[0013] (6) Judgment step: calculate the enrichment immune characterization feature of pathogen i , is the total number of positive peptides of pathogen i; if > 1, it is determined that the pathogen is a positive immune response.
[0014] Further, the sequencing data of the antigen epitope peptide library in step (1) is obtained as follows:
[0015] Proteomes of n pathogens are obtained and split into a series of equal-length epitope peptides; assuming that the i-th pathogen has epitopes, the total number of epitopes of all n pathogens is ; all epitope peptide sequences are converted into corresponding DNA sequences, oligo synthesis is performed, in vitro translation and transcription are performed using mRNA display technology, and a one-to-one nucleic acid-peptide complex, i.e. an antigen epitope peptide library, is formed; the DNA fragments of the antigen epitope peptide library are collected for PCR amplification, library construction, and high-throughput sequencing to obtain the sequencing data of the antigen epitope library.
[0016] Further, the antigen epitope library in step (1) is interacted with the antibodies in the biological sample, and the antigen epitope peptides that interact with the antibodies in the biological sample are captured by immunoprecipitation; the sequences of the captured peptides are eluted, PCR amplified, and library constructed, and then high-throughput sequencing is performed to obtain the sequencing data of the immunoprecipitation sample.
[0017] Further, the read count of each pathogen peptide in the antigen peptide library in step (3) is denoted as The read count of each pathogen peptide in the co-precipitated sample is recorded as follows: ;in This is the read count for the j-th peptide in the antigen peptide library. is the read count of the j-th peptide in the immunoprecipitation sample, where M is the total number of pathogen peptides.
[0018] Specifically, it includes the following:
[0019] (401) Based on the statistical principle of the generalized Poisson distribution, the peptides of the antigen epitope library are grouped according to the number of reads, and the total number of peptides in each group is: For each group For each peptide segment, the read count data of the corresponding peptide segment in the immunoprecipitation sample are obtained and the maximum likelihood estimation is performed to obtain the fitting parameters λ and θ of each group of generalized Poisson models.
[0020] (402) Perform linear fitting on each set of fitting parameters obtained in step (401) to predict the model parameters λ and θ at different signal levels, and calculate the expected number of reads E for each peptide in the immunoprecipitation sample; compare the actual number of reads X obtained by each peptide in the immunoprecipitation sample with the expected number of reads E, calculate the statistical p value, thereby confirming whether it is an enriched epitope peptide, and obtain the enriched peptide set vector. ;in For the first The peptide segment showed positive peptide values in the immunoprecipitation sample, and had... ;
[0021] Furthermore, the likelihood estimation formula in step (401) is as follows: ,in, ,and k is the total number of peptides in each group.
[0022] Furthermore, the probability mass function of the generalized Poisson distribution model in step (402) is: Therefore, the expected number of reads E in the immunoprecipitation sample at different peptide signal levels can be calculated.
[0023] Furthermore, if the p-value is less than 0.05, then the pathogen peptide is a significantly enriched epitope.
[0024] In a second aspect, the present invention provides a computer device including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the steps of the method described in the first aspect.
[0025] In a third aspect, the present application provides a computer readable storage medium having stored thereon a computer program which, when executed by a processor, implements the steps of the method of the first aspect.
[0026] Compared with the related art, the pathogen immunity lineage evaluation method based on high-throughput antibody epitope spectrum detection data provided by the present application has the following beneficial effects:
[0027] Based on the cross fields of bioinformatics, immunology and computer science, the present application provides an Enrichment-based Pathogen Immunity Characterization (EPIC) evaluation method, which can extract key information from high-throughput antibody spectrum detection data, distinguish effective signals from non-specific binding noise background, summarize the information of numerous epitopes of a pathogen to the overall infection score of the pathogen, efficiently process thousands of viral site information into an immune status evaluation spectrum of an individual against multiple pathogens, and intuitively present the detection results through image visualization. This method greatly improves the throughput, accuracy and clinical application value of pathogen immunity detection, and can be used for individual pathogen antibody detection and population pathogen dynamic monitoring. It is suitable for pathogen infection screening, immune response research and disease diagnosis.
[0028] Specifically,
[0029] (1) The present application can integrate multiple antigen epitope signals of a pathogen, summarize to the overall infection score of the pathogen, realize the overall evaluation of the immune status, break through the judgment limitation of single epitope detection, and improve the sensitivity and specificity of pathogen immunity detection.
[0030] (2) The present application corrects the design differences of different pathogens in the epitope library by calculating the normalized signal of the pathogen, so that the immunity detection of different pathogens is comparable.
[0031] (3) The present application verifies the signals among the epitopes simultaneously, thereby overcoming the limitation of a single fixed threshold and improving the accuracy of detection.
[0032] (4) The present application dynamically calculates the EPIC score in combination with the overall experimental signal background of the sample, normalizes the signal according to the background information of the individual and the library, so that the detection is more adaptive, and the final immunity judgment has better adaptability and tolerance to the volatility of each link of the experiment and the sample. Thus, it can support population pathogen dynamic evaluation, such as large-scale statistical analysis of the epidemic trend of the pathogen and the effect of vaccination.
[0033] (5) The application is compatible with PhIP-seq, PepSeq, SLISA, protein chip and other high-throughput antibody spectrum detection technologies, and can be widely applied to pathogen infection screening, vaccine evaluation, epidemiological research and immune monitoring. BRIEF DESCRIPTION OF DRAWINGS
[0034] Figure 1 Figure 1 is a schematic diagram of the correspondence between the pathogen and the epitope peptide segment of the application.
[0035] Figure 2 Figure 2 is a high-throughput sequencing sample read base quality distribution graph of the application.
[0036] Figure 3 Figure 3 is a schematic diagram of the "ATCG" base composition distribution of the high-throughput sequencing sample of the application.
[0037] Figure 4 Figure 4 is a schematic diagram of the sample sequence alignment result statistics of the application.
[0038] Figure 5 Figure 5 is a schematic diagram of the model lambda and theta parameter distribution of the application.
[0039] Figure 6 Figure 6 is a schematic diagram of the EPIC correction effect of the application.
[0040] Figure 7 Figure 7 is a schematic diagram of the stability of the EPIC score of the application in representing the individual viral infection state.
[0041] Figure 8 Figure 8 is a schematic diagram of the evaluation comparison of the EPIC score of the application on repeated samples.
[0042] Figure 9 Figure 9 is a schematic diagram of the population pathogen EPIC score immunity map of the application.
[0043] Figure 10 Figure 10 is a schematic diagram of the individual pathogen EPIC score immunity map of the application, wherein different colors of blocks represent immunity to different known viruses, the larger the block area, the greater the immunity, and the smaller the block area, the smaller the immunity. DETAILED DESCRIPTION
[0044] The embodiments of the application are described in detail below, wherein the same or similar reference numerals represent the same or similar elements or elements with similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary and are only used to explain the application and cannot be regarded as a limitation of the application.
[0045] As used herein, and unless otherwise indicated, all terms have their ordinary meanings. It will be further understood that, as used in the specification and the appended claims, the singular forms "a," "an" and "the" include plural or work equivalents, unless the context clearly dictates otherwise. All patents and publications mentioned herein are incorporated by reference.
[0046] The steps mentioned in the embodiments are only for the convenience of description, and are not related to the substantial sequence. The different steps in the embodiments can be combined in different sequences to achieve the purpose of the application. The modules and methods not described in detail in the application can be implemented by conventional technical means, so they will not be described in detail.
[0047] The application will be further described below in conjunction with the drawings and specific embodiments.
[0048] Please refer to Figures 1 to 9 The pathogen immunological lineage evaluation method provided by the application based on high-throughput antibody epitope spectrum detection data includes the following steps:
[0049] (1) Data preparation step: obtain the sequencing data of the antigen epitope library (Input library) and the sequencing data of the immunoprecipitation (IP) sample. The sequencing data of the antigen peptide library is obtained as follows:
[0050] First, obtain the proteome of n pathogens, and split it into a series of equal-length epitope peptides. For example, Figure 2 As shown in the figure, it is assumed that the ith pathogen has peptides, and the total number of peptides of all n pathogens is That is, the total number of peptides in the antigen peptide library is M.
[0051] All peptide sequences are converted into corresponding DNA sequences, and after oligo synthesis, in vitro translation and transcription are performed using mRNA display technology to form a one-to-one nucleic acid-peptide complex. These nucleic acid-peptide complexes constitute the antigen peptide library.
[0052] Collect the DNA fragments of the antigen peptide library for PCR amplification, library construction, and illumina high-throughput sequencing to obtain the sequencing data of the antigen epitope peptide library.
[0053] The prepared antigen epitope library is interacted with antibodies in a biological sample (such as serum), and the antigen epitope peptides interacting with the antibodies in the biological sample are captured by immunoprecipitation. The captured peptide sequences are eluted, PCR amplified, and library constructed, and then high-throughput sequencing is performed to obtain sequencing data of the immunoprecipitation sample.
[0054] (2) Data preprocessing step: quality control is performed on all sequencing data in step (1), including removing adapter sequences at both ends of read data and filtering read data with too short length.
[0055] Specifically, quality control is performed on fastq files obtained by sequencing all immunoprecipitation samples and antigen peptide library samples, as shown in FIG. 2, using the fastqc tool to detect read base quality distribution, "ATCG" base composition distribution, etc., and using the cutadapt tool to remove adapter sequences at both ends of the read. Figure 3
[0056] (3) Data alignment and counting step: based on the sequencing data of all antigen peptide libraries, a reference library index is constructed. (It is constructed by using the bowtie2-build command based on all antigen peptide library DNA sequences)
[0057] The read data preprocessed in step (2) is aligned with the reference library index, and the alignment results are counted, as shown in FIG. 3, including the percentage of unaligned reads, the percentage of uniquely aligned reads, etc. of each sample, wherein the abscissa represents the sample and the ordinate represents the percentage value. Figure 4 According to the alignment results, the read count of each pathogen epitope peptide in the antigen epitope library sample is counted using the samtools tool, denoted as an array Y, wherein Y i represents the read count of the i th peptide in the antigen peptide library, and there are a total of P peptides.
[0058] wherein Y i represents the read count of the i th peptide in the antigen peptide library, and there are a total of P peptides. wherein Y i represents the read count of the i th peptide in the antigen peptide library, and there are a total of P peptides. wherein Y i represents the read count of the i th peptide in the antigen peptide library, and there are a total of P peptides. wherein Y i represents the read count of the i th peptide in the antigen peptide library, and there are a total of P peptides.
[0059] (4) Peptide identification step: taking the obtained antigen epitope peptide library peptide count Y and the immunoprecipitation sample peptide read count R as input data, the significantly enriched epitope peptides in the IP sample are identified based on the generalized Poisson distribution (generalized Poisson) model. Specifically, it includes the following:
[0060] (401) Grouping the peptides in the antigenic peptide library according to the read counts. That is, grouping the peptides with the same read counts into the same group, and totally obtaining groups.
[0061] Then, for each group of peptides, the maximum likelihood estimation of the read count data of the corresponding peptides in the immunoprecipitation sample is obtained, and the fitting parameters λ and θ of the generalized Poisson model of each group are obtained. That is, the read count data of the corresponding peptides in the immunoprecipitation sample is obtained, denoted as , wherein is the read count of the i-th peptide in the immunoprecipitation sample in the group, and the group has a total of peptides. The maximum likelihood estimation (MLE) is used to fit , and two parameters and of the generalized Poisson model are obtained. The maximum likelihood estimation formula is: , wherein , and .
[0062] (402) The above processing is performed for all groups, and the parameter sets corresponding to the groups are obtained and , denoted as and . Linear fitting is performed on and , as shown in Figure 5 , so that the prediction models and corresponding to different antigenic peptide library signal levels are obtained. The different level prediction parameters and are brought into the probability mass function (pmf) of the generalized Poisson model: , that is, the expected read count of different peptide signal levels in the IP sample can be calculated. According to the actual observed read count of each peptide in the IP sample and the expected read count , the statistical p value is calculated. Finally, according to the p values obtained by all peptides, the multiple test correction is performed to obtain the adjusted p value, and according to , it is calculated whether each pathogen peptide segment in the IP sample is an enriched epitope peptide segment. For the i-th pathogen of the j-th antigenic peptide library, the read count of the corresponding peptide in the immunoprecipitation sample is obtained, denoted as , wherein is the read count of the i-th peptide in the immunoprecipitation sample in the group, and the group has a total of peptides.The results of epitope peptide enrichment were used to determine the outcome. Indications. A value of 1 indicates a significantly enriched peptide (positive), while a value of 0 indicates that the peptide was not significantly enriched in this sample.
[0063] (5) Steps for calculating calibration coefficients: For the first... There are pathogens, If there are epitope peptides, then the total number of positive peptides for this pathogen is [number missing]. .
[0064] Number of positive epitopes for each pathogen The total number of positive epitopes detected for all pathogens is . For pathogen i, its corresponding correction coefficient Let N be the total number of peptides for pathogen i, N be the total number of positive peptides detected for all pathogens, and M be the total number of peptides in the antigen peptide library. This dynamic normalization calculation accurately reflects the intensity of an individual's antibody response to a specific pathogen, and it does not depend on a fixed threshold but rather achieves adaptive judgment based on the relative signal within the sample and the library background.
[0065] (6) Judgment steps: Calculate the overall immunity EPIC score of pathogen i, i.e. , This represents the total positive peptide fragments of pathogen i. If If the value is greater than 1, it indicates that the epitope of the pathogen is significantly enriched in the IP sample, meaning that the IP sample has significant immunity to the pathogen, which largely reflects the sample's history of previous infection with the corresponding pathogen. It reflects the ratio of the degree of pathogen epitope enrichment to the background difference in the pathogen experiment.
[0066] In summary, by comprehensively assessing the immunogenicity of a pathogen in IP samples by aggregating signals from all epitopes of the pathogen, the biases inherent in traditional fixed threshold methods can be avoided, significantly reducing false positive / false negative rates. By incorporating a calibration coefficient B, the EPIC scoring formula introduces weighting factors such as the total number of pathogen epitope peptides and the total number of global positive epitopes, eliminating the influence of pathogen genome size differences on the results and making cross-pathogen immune response assessments more equitable. Furthermore, EPIC normalizes signals based on individual and library background information, making the detection more adaptive and allowing the final immunogenicity assessment to better accommodate and tolerate fluctuations in experimental and sample processes.
[0067] from Figure 6The comparison between the left and right graphs can clearly show that the number of viral enriched peptides obtained from the high-throughput antibody profiling data is significantly positively correlated with the size of the virus (left graph), and after the EPIC score calculation, the correlation between the viral EPIC score and the size of the virus is significantly reduced (right graph). Therefore, the EPIC score of the present application corrects the bias of different size distributions of viruses in the library, and can more fairly evaluate the infectivity status of different viruses. The English marked by the dots in the graph is the name of different viruses.
[0068] And, as shown in Figure 7 and Figure 8 , the correlation of the EPIC evaluation in repeated experimental samples is significantly higher than that between non-repeated samples; and the scatter plot also shows that the EPIC evaluation calculated from different repeated samples is highly consistent, proving that the EPIC evaluation can stably represent the viral infectivity status of individuals.
[0069] The present scheme can also obtain the pathogen immunity map of the population, as shown in Figure 9 . This result shows the immunity map of each person in a population cohort against different pathogens (hundreds of viruses shown in the graph) under different body conditions. Whether the virus is large or small, the EPIC score can clearly and accurately show the immunity status of the biological sample against hundreds of viruses.
[0070] As shown in Figure 10 , the EPIC score can dynamically monitor the changes in the immunity map of individuals against hundreds of viruses over time and under different body conditions. For example, the first row in the graph shows the changes in the viral immunity map of a subject 1 before and after SARS-CoV-2 infection for several weeks to several months. It can be seen that before and after the infection of SARS-CoV-2, the immunity status of the subject against different viruses shows obvious differences (among them, the immunity changes of human rhinovirus A (HRV8A), echovirus 16 (9ENTO), echovirus 12 (EC12T), and poliovirus 1 (POL1M) type are the most obvious, and the color blocks of other colors also represent other different known common viruses, which are not described in detail); the same phenomenon can also be observed in the viral immunity map of another patient (subject 2) in the second row. The changes in the immunity map of subject 2 are most obvious for human rhinovirus A (HRV8A), echovirus 30 (EC30B), and echovirus 6 (EC06C). In addition, by comparing the immunity maps of the two patients, it can also be found that the EPIC viral immunity map shows high individual specificity, and the composition mode and response intensity of the immune characteristic spectrum are significantly different in the population. Different color blocks in the graph represent the immunity of different viruses.
[0071] In one embodiment, the present application further provides a computer device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the steps in the segmentation method provided by the above-mentioned embodiment when executing the computer program.
[0072] In one embodiment, the present application further provides a computer readable storage medium, having a computer program stored thereon, wherein the computer program is executable by a processor to implement the steps in the segmentation method provided by the above-mentioned embodiment.
[0073] A person of ordinary skill in the art can understand that all or part of the processes in the above-mentioned embodiment methods can be completed by a computer program instructing related hardware, and the computer program can be stored in a non-volatile computer readable storage medium. When the computer program is executed, it can include the processes of the above-mentioned embodiments. Any reference to memory, storage, database or other medium used in the embodiments provided by the present application can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM) or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. As an illustration but not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct RAM bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM).
[0074] Each technical feature of the above-mentioned embodiments can be combined arbitrarily. In order to make the description simple, all possible combinations of each technical feature in the above-mentioned embodiments are not described, however, as long as the combination of these technical features does not exist contradictory, it should be considered as the scope of the present application.
[0075] The several embodiments of the application are described in more detail and specifically, but they should not be understood as limiting the scope of the patent. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, a number of modifications and improvements can be made, and these improvements and refinements should be considered as within the scope of the present application. Therefore, the scope of the patent of the present application should be subject to the appended claims.
Claims
1. A method for assessing the immune lineage of pathogens based on high-throughput antibody epitope profiling data, characterized in that, Includes the following steps: (1) Data preparation steps: Obtain sequencing data of antigen epitope library and sequencing data of biological sample after incubation with library for immunoprecipitation; The sequencing data of the antigen epitope library were obtained as follows: Obtain the proteomes of n pathogens and split them into a series of epitope peptides of equal length; define the i-th pathogen as having m i If there are n antigenic epitopes, then the total number of epitopes for all n pathogens is... ; All peptide sequences are converted into corresponding DNA sequences, synthesized via oligo, and then translated and transcribed in vitro using mRNA display technology to form one-to-one nucleic acid-peptide complexes, i.e., viral antigen epitope libraries. The DNA fragments of the antigen epitope library are collected for PCR amplification, library construction, and high-throughput sequencing to obtain the sequencing data of the antigen epitope library. (2) Data preprocessing steps: Quality control is performed on all sequencing data in step (1), including removing adapter sequences at both ends of reads and filtering out reads that are too short. (3) Data comparison and counting steps: Based on the sequencing data of all antigen epitope libraries, construct a reference library index; compare the read data filtered in step (2) with the reference library index, and count the read count of each pathogen epitope in the antigen epitope library and the read count of each pathogen epitope peptide in the immunoprecipitation sample. (4) Epitope peptide identification steps: Identify significantly enriched antigenic epitopes in immunoprecipitation samples based on the generalized Poisson distribution model; (5) Steps for calculating calibration coefficients: For pathogen i, the corresponding calibration coefficients are... ; m i denoted as the total number of antigenic epitopes of pathogen i, N as the total number of positive epitopes detected by all pathogens, and M as the total number of epitopes in the antigenic epitope library. (6) Judgment Step: Calculate the enrichment immunological characterization features of pathogen i. , n i The total positive antigenic epitopes of pathogen i; if If the value is greater than 1, the biological sample is considered to have a positive immune response against the pathogen.
2. The method for evaluating the immune lineage of pathogens based on high-throughput antibody epitope profiling data according to claim 1, characterized in that, The antigen epitope library in step (1) is incubated with a biological sample. The antibodies in the biological sample will interact with the antigen epitope peptides in the library. The antigen peptides that interact with the antibodies in the biological sample are captured by immunoprecipitation. The captured peptide sequences are eluted, amplified by PCR, and constructed into a library. Then, high-throughput sequencing is performed to obtain the sequencing data of the immunoprecipitated sample, which is the raw data of antibody epitope profiling.
3. The method for evaluating the immune lineage of pathogens based on high-throughput antibody epitope profiling data according to claim 1, characterized in that, The read count of each pathogen epitope peptide in the antigen peptide library in step (3) is recorded as follows: The read count of each pathogen epitope peptide in the co-precipitated sample is recorded as . ;in y j This is the read count for the j-th peptide in the antigen peptide library. x j is the read count of the j-th peptide in the immunoprecipitation sample, where M is the total number of pathogen peptides.
4. The method for evaluating the immune lineage of pathogens based on high-throughput antibody epitope profiling data according to claim 1, characterized in that, The specific steps for identifying whether the peptide segment is a significantly enriched positive immunoprecipitation result are as follows: Includes the following processes: (401) Based on the statistical principle of the generalized Poisson distribution, the peptides in the antigen epitope peptide library are grouped according to the number of reads, with the total number of peptides in each group being... k For each group k For each peptide segment, the read count data of the corresponding peptide segment in the immunoprecipitation sample are obtained and the maximum likelihood estimation is performed to obtain the fitting parameters λ and θ of each group of generalized Poisson models. (402) Perform linear fitting on each set of fitting parameters obtained in step (401) to predict the model parameters λ and θ at different signal levels, and calculate the expected number of reads E for each epitope peptide in the immunoprecipitation sample; compare the actual number of reads X obtained by each peptide in the immunoprecipitation sample with the expected number of reads E, calculate the statistical p value, thereby confirming whether it is an enriched peptide, and obtain the enriched peptide set vector. ;in h j For the first j The peptide segment showed positive epitope peptide values in the immunoprecipitation sample, and had .
5. The method for evaluating the immune lineage of pathogens based on high-throughput antibody epitope profiling data according to claim 4, characterized in that, The likelihood estimation formula in step (401) is: ,in, ,and .
6. The method for evaluating the immune lineage of pathogens based on high-throughput antibody epitope profiling data according to claim 4, characterized in that, The probability mass function of the generalized Poisson distribution model in step (402) is: Therefore, the expected number of reads E in the immunoprecipitation sample with different epitope peptide signal levels can be calculated.
7. The method for evaluating the immune lineage of pathogens based on high-throughput antibody epitope profiling data according to claim 5, characterized in that, If the p-value is less than 0.05, then the pathogen peptide is an enriched peptide.
8. A computer device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the steps of the method as described in any one of claims 1 to 7.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the steps of the method as described in any one of claims 1 to 7.
Citation Information
Patent Citations
Detection of an antibody against a pathogen
US20160320406A1
Method for detecting the presence of malignant cells using a multi-protein DNA replication complex
US6093543A