Analysis, prediction methods and equipment for VJ gene preference of neutralizing antibodies against new coronavirus

By conducting scientific statistical analysis on the VJ gene selection frequency and pairing frequency of neutralizing antibodies against the new coronavirus and combining the visualization results, the shortcomings of the existing technology in the analysis of the neutralizing antibody V(D)J gene rearrangement relationship were solved, and the prediction of neutralizing antibody preference and the early prediction of the probability of neutralizing antibody production were achieved.

CN115831235BActive Publication Date: 2025-10-14SOUTHERN MEDICAL UNIVERSITY
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210957595.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2021-11-18
Filing Date
2022-08-10
Publication Date
2025-10-14
Estimated Expiration
2042-08-10

AI Technical Summary

Technical Problem

In the existing technology, the analysis results of the relationship between the V(D)J gene rearrangement and neutralizing ability of neutralizing antibodies against the new coronavirus are presented in a single form, lack statistical support, and cannot scientifically predict the relationship between V(D)J gene rearrangement and neutralizing antibodies.

Method used

The VJ gene selection frequency of human anti-new coronavirus neutralizing antibodies and non-neutralizing antibodies was analyzed through the coronavirus database, and the OR value, P value and Padj value were calculated using statistical methods. The preference was displayed in visual forms such as horizontal grouped bar charts, volcano charts, heat maps and chord charts. The probability of neutralizing antibody production was predicted based on the VJ gene frequency of peripheral blood B cells.

Benefits of technology

A scientific analysis of the neutralizing antibody VJ gene selection and pairing preference has been achieved, which can intuitively display the preference and make early predictions on whether the probability of the subjects producing neutralizing antibodies is higher than the baseline level of healthy people.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115831235B_ABST
    Figure CN115831235B_ABST
Patent Text Reader

Abstract

The application belongs to the technical field of bioinformatics analysis, and discloses a kind of anti-new coronavirus neutralizing antibody VJ gene preference analysis method, including antibody database acquisition data, data screening, selection frequency statistics, selection preference analysis, selection frequency and preference visualization, pairing frequency statistics, pairing preference analysis, pairing frequency and preference visualization.The beneficial aspects of the present application are that the neutralizing antibody and VJ selection and pairing frequency are tested by scientific statistical methods, and are visualized to make the result interpretation more simple and intuitive.Therefore, according to the above method, a new method for early prediction of anti-new coronavirus neutralizing antibody production based on peripheral blood B cell VJ gene frequency and preference is established, which can predict whether the probability of producing neutralizing antibody by the subject is higher than the baseline level of healthy people.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of bioinformatics analysis technology, and specifically relates to a method for analyzing the VJ gene preference of anti-new coronavirus neutralizing antibodies, a method for predicting anti-new coronavirus neutralizing antibodies, and electronic equipment. Background Art

[0002] In higher mammals, germline V(D)J genes are present on immature B lymphocytes in the bone marrow. They consist of multiple consecutive V, D, and J gene segments located at corresponding positions on the chromosomes. During B cell development and maturation, germline V(D)J gene rearrangement occurs. Either randomly or through a potential mechanism, a set of V(D)J genes is selected to join together and form the antigen-binding portion of the B-cell antigen receptor (BCR) located on the cell surface. Due to this high repertoire diversity, the BCR can recognize nearly all foreign antigens. Upon activation by foreign antigens in peripheral immune organs, B cells undergo somatic hypermutation (SHM) and class switch recombination (CSR), ultimately differentiating into plasma cells capable of secreting the corresponding antibodies. Therefore, the neutralizing potency of the secreted antibodies is determined by the V(D)J gene rearrangements in the corresponding B cells.

[0003] Currently, several neutralizing antibodies against the novel coronavirus have been used in clinical practice, and many studies on neutralizing antibodies against the novel coronavirus have attempted to explain the relationship between neutralizing antibodies and V(D)J gene rearrangement. However, the main defects of these studies are: (1) the results are presented in a single format, which only shows the selection frequency of the V(D)J gene of neutralizing antibodies against the novel coronavirus; (2) the background of the selection frequency of the V(D)J gene itself is not considered, and there is a lack of statistical analysis results to support it. Therefore, it is impossible to scientifically predict the relationship between V(D)J gene rearrangement and the neutralizing ability of neutralizing antibodies. Summary of the Invention

[0004] The purpose of the present invention is to provide a method for analyzing the VJ gene preference of anti-new coronavirus neutralizing antibodies, a method for predicting anti-new coronavirus neutralizing antibodies, and an electronic device. Based on the data of human anti-new coronavirus neutralizing antibodies and non-neutralizing antibodies included and continuously updated in the coronavirus database (CoV-AbDab), the method analyzes whether the selection frequency of the VJ gene of anti-new coronavirus neutralizing antibodies is different from that of non-neutralizing antibodies, performs scientific statistical analysis, and visualizes the results in different forms.

[0005] The first aspect of the embodiments of the present invention discloses a method for analyzing the VJ gene preference of anti-new coronavirus neutralizing antibodies, comprising:

[0006] Obtain antibody data from the coronavirus antibody database, and screen the antibody data for neutralizing and non-neutralizing antibodies against the novel coronavirus;

[0007] The neutralizing antibodies and the non-neutralizing antibodies were grouped according to the six genes of IGHV, IGHJ, IGKV, IGKJ, IGLV, and IGLJ, and the selection frequency of each gene in the neutralizing antibodies and the non-neutralizing antibodies was counted;

[0008] Calculate the selection preference OR value, P value, and Padj value of each gene in the neutralizing antibody;

[0009] Sort by OR value from large to small, and display the selection frequency of each gene in the neutralizing antibody and the non-neutralizing antibody in the form of a horizontal grouped bar graph;

[0010] The OR value and Padj value of each gene are displayed in the form of a volcano plot, highlighting the preferred genes.

[0011] The second aspect of the present invention discloses another method for analyzing the VJ gene preference of neutralizing antibodies against the new coronavirus, comprising:

[0012] Obtain antibody data from a coronavirus antibody database, and screen the antibody data for neutralizing and non-neutralizing antibodies against the novel coronavirus;

[0013] Grouping by heavy chain and light chain, and counting the pairing frequencies of each V gene and each J gene in the neutralizing antibody and the non-neutralizing antibody;

[0014] Calculate the pairing preference P value, OR value, and Padj value of each VJ gene pair in the neutralizing antibody and the non-neutralizing antibody;

[0015] The pairing frequencies of the VJ genes in the neutralizing antibodies and the non-neutralizing antibodies are displayed in the form of a heat map, and the first preference analysis results are highlighted;

[0016] The specific VJ gene pairing of significant genes is displayed in the form of a chord diagram, and the second preference analysis results are highlighted.

[0017] The third aspect of an embodiment of the present invention discloses an electronic device, comprising a memory storing executable program code and a processor coupled to the memory; the processor calls the executable program code stored in the memory to execute the VJ gene preference analysis method for neutralizing antibodies against the new coronavirus disclosed in the first aspect or the second aspect.

[0018] A fourth aspect of the embodiments of the present invention discloses a method for predicting anti-novel coronavirus neutralizing antibodies, which is performed according to the anti-novel coronavirus neutralizing antibody VJ gene preference analysis method disclosed in the first or second aspect of the claim, and the prediction method includes:

[0019] The VJ gene sequencing results of healthy human peripheral blood B cells were obtained from a public dataset, and the read ratio of each gene was calculated as the baseline frequency, which was used to characterize the baseline level of healthy people.

[0020] Obtain sequencing data of the subject's peripheral blood, and calculate the read ratio of each VJ gene in the subject's peripheral blood as the actual frequency based on the sequencing data. The actual frequency is used to represent the actual level of the subject;

[0021] Based on the aforementioned VJ gene selection and pairing preference analysis results, gene segments that are significantly preferred in neutralizing antibodies are determined as favorable genes, and gene segments that are significantly preferred in non-neutralizing antibodies are determined as unfavorable genes; wherein Padj <= 0.05 indicates a significantly preferred gene segment;

[0022] Using the Fisher exact test and taking p = 0.05 as the threshold, the frequency difference between the actual frequency of each favorable gene and unfavorable gene measured in the peripheral blood of the subject is calculated as being higher / lower than the baseline frequency corresponding to the gene; based on the frequency difference, it is judged whether the probability of the subject producing neutralizing antibodies is higher than the baseline level of healthy people.

[0023] The fifth aspect of the embodiments of the present invention discloses a computer-readable storage medium, which stores a computer program, wherein the computer program enables a computer to execute the VJ gene preference analysis method for neutralizing antibodies against the new coronavirus disclosed in the first aspect or the second aspect.

[0024] The beneficial effects of the present invention lie in that the provided anti-new coronavirus neutralizing antibody VJ gene preference analysis method, anti-new coronavirus neutralizing antibody prediction method and electronic equipment, by conducting scientific statistical analysis on the VJ gene selection frequency and pairing frequency of human anti-new coronavirus neutralizing antibodies and non-neutralizing antibodies, and highlighting them accordingly in the visualization results, can more simply and intuitively display the VJ gene selection preference and pairing preference of neutralizing antibodies. More importantly, the above method can be used to establish a new method for early prediction of anti-new coronavirus neutralizing antibody production based on peripheral blood B cell VJ gene frequency and preference, which can predict whether the probability of a subject producing neutralizing antibodies is higher than the baseline level of healthy people. BRIEF DESCRIPTION OF THE DRAWINGS

[0025] The accompanying drawings herein illustrate specific examples of the technical solutions described in the present invention, and together with the specific implementation methods constitute a part of the specification, and are used to explain the technical solutions, principles and effects of the present invention.

[0026] Unless otherwise specified or defined, the same reference numerals in different drawings represent the same or similar technical features, and the same or similar technical features may also be represented by different reference numerals.

[0027] Figure 1 Schematic diagram of the process of the method for analyzing the VJ gene preference of neutralizing antibodies against the new coronavirus of the present invention;

[0028] Figure 2 The frequency statistics of the 10 genes with the highest neutralizing antibody frequencies in the IGHV group in the embodiments of the present invention are as follows;

[0029] Figure 3 In the embodiment of the present invention, the top 10 genes of the frequency preference analysis results of the IGHV group were selected;

[0030] Figure 4 A histogram of the IGHV group in the VJ selection frequency in an embodiment of the present invention;

[0031] Figure 5 A volcano plot of the VJ selection preference analysis in an embodiment of the present invention;

[0032] Figure 6 10 V genes of the heavy chain group in the VJ pairing frequency results in the embodiment of the present invention;

[0033] Figure 7 The results of V genes with significantly different composition ratios of heavy chain groups in the VJ pairing preference analysis in the examples of the present invention are as follows;

[0034] Figure 8 This is the heat map result of the heavy chain group in the VJ pairing frequency in the embodiment of the present invention;

[0035] Figure 9 This is the chord diagram result of the heavy chain group in the VJ pairing preference analysis in the examples of the present invention. DETAILED DESCRIPTION

[0036] To facilitate understanding of the present invention, specific embodiments of the present invention will be described in more detail below with reference to the accompanying drawings.

[0037] Unless otherwise specified or defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the art. In the context of combining the technical solution of the present invention with realistic scenarios, all technical and scientific terms used herein may also have meanings corresponding to the purpose of implementing the technical solution of the present invention. "First, second..." used herein is merely used to distinguish names and does not represent a specific quantity or order. The term "and / or" used herein includes any and all combinations of one or more related listed items.

[0038] The present invention is further described below by using an embodiment in conjunction with the accompanying drawings. The following embodiment is implemented by R software according to the technical solution of the present invention, and a specific implementation method and operation process are given. However, the protection scope of the present invention is not limited to the following embodiment, and the implementation tools, specific judgment thresholds, highlight mark colors and other display methods therein do not constitute limitations on the present invention. It should be pointed out that for ordinary technicians in this field, without departing from the concept of the present invention, several variations and improvements can be made, all of which fall within the scope of protection of the present invention.

[0039] The embodiment of the present invention discloses a method for analyzing the VJ gene preference of neutralizing antibodies against the new coronavirus, and the main process diagram is as follows: Figure 1 shown.

[0040] S1. Obtain antibody data from the database and perform data input, organization, and screening.

[0041] 1.1 Download all coronavirus antibody data from the coronavirus antibody database CoV-AbDab (http: / / opig.stats.ox.ac.uk / webapps / covabdab / ), a total of 4711 strains, and input them into R software;

[0042] 1.2 Among them, 1953 strains with neutralizing antibodies and 1403 strains with non-neutralizing antibodies against SARS-CoV-2 were screened;

[0043] 1.3 Antibodies with the "Human" tag on the VJ gene were screened, and 1,679 human anti-COVID-19 neutralizing antibodies and 1,349 non-neutralizing antibodies were obtained.

[0044] The VJ gene data of the above antibodies serve as input data for the next step of analysis. The next step of analysis includes step S2 or S3. Steps S2 and S3 can be analyzed independently of each other.

[0045] S2. VJ gene selection frequency preference analysis and visualization

[0046] 2.1. VJ gene selection frequency statistics, specifically including the following steps:

[0047] 2.1.1 Based on the gene information of the immunoglobulin heavy variable cluster (IGHV), immunoglobulin heavy joining cluster (IGHJ), immunoglobulin kappa variable cluster (IGKV), immunoglobulin kappa joining cluster (IGKJ), immunoglobulin lambda variable cluster (IGLV), and immunoglobulin lambda joining cluster (IGLJ) contained in each antibody, all antibodies were further divided into six groups for analysis. Duplicates were removed and erroneous gene information was screened out.

[0048] Specifically, based on the annotation information of all immunoglobulin V(D)J genes in the IMGT database, genes not included in the IMGT annotation were screened out. Among the total 3028 antibodies (1679 neutralizing antibodies and 1349 non-neutralizing antibodies), 3007 contained IGHV genes (1667 neutralizing antibodies and 1340 non-neutralizing antibodies), 2691 contained IGHJ genes (1489 neutralizing antibodies and 1202 non-neutralizing antibodies), 1856 contained IGKV genes (1011 neutralizing antibodies and 845 non-neutralizing antibodies), 1650 contained IGKJ genes (900 neutralizing antibodies and 750 non-neutralizing antibodies), 1151 contained IGLV genes (647 neutralizing antibodies and 504 non-neutralizing antibodies), and 1027 contained IGLJ genes (575 neutralizing antibodies and 452 non-neutralizing antibodies).

[0049] 2.1.2 Taking the IGHV group as an example, the table function of the base package was used to perform selection frequency statistics on each IGHV gene of neutralizing antibodies and non-neutralizing antibodies, and then integrated. The IGHV group genes finally included 50 IGHV genes and the selection frequency of each IGHV gene in neutralizing antibodies and non-neutralizing antibodies. The selection frequency statistics of the 10 genes with the highest frequency in neutralizing antibodies were as follows: Figure 2 shown.

[0050] 2.2 Statistical analysis of VJ gene selection preference, specifically including the following steps:

[0051] 2.2.1. In the frequency statistics results described in 2.1.2, the fisher.test function of the stats package was used to calculate the OR value and P value of each gene. Specifically, in R software, run "fisher.test(matrix(c(n,m,Nn,Mm),ncol=2,nrow=2))[[3]]" to output the OR value, and run "fisher.test(matrix(c(n,m,N,M),ncol=2,nrow=2))[[1]]" to output the P value, where n is the frequency of the gene counted in the neutralizing antibody group, m is the frequency of the gene counted in the non-neutralizing antibody group, N is the frequency of all neutralizing antibodies in the group, and M is the frequency of all non-neutralizing antibodies in the group.

[0052] Taking the IGHV3-53 gene of the IGHV group as an example, the frequency of the IGHV3-53 gene in neutralizing antibodies is 201, and the frequency in non-neutralizing antibodies is 33. The total number of neutralizing antibodies in the IGHV group is 1667, and the total number of non-neutralizing antibodies is 1340. Then OR = fisher.test(matrix(c(201,33,1667-201,1340-33),ncol=2,nrow=2))[[3]]=5.427569, P = fisher.test(matrix(c(201,33,1667-201,1340-33),ncol=2,nrow=2))[[3]]=8.569462e-25, that is, the calculated OR value of the IGHV3-53 gene in neutralizing antibodies is 5.428, and the P value is 8.569E-25.

[0053] 2.2.2 In the calculation results described in 2.2.1, use the order function of the base package to sort the genes by P value from small to large, and then use the p.adjust function of the stats package to perform multiple correction of the P value using the "fdr" method to obtain the Padj value. The results are sorted from large to small by OR value to obtain the VJ gene selection frequency preference analysis results. The results of the first 10 genes in the IGHV group are as follows Figure 3 shown.

[0054] 2.3 Visualization of VJ gene selection frequency, specifically including the following steps:

[0055] 2.3.1 Group by IGHV, IGHJ, IGKV, IGKJ, IGLV, and IGLJ. Based on the VJ gene selection frequency preference analysis results described in 2.2.2, use the ggplot2 package's ggplot function and geom_bar function to plot the frequencies of neutralizing and non-neutralizing antibodies as horizontal grouped bar charts, with the neutralizing antibody group on the left and the non-neutralizing antibody group on the right.

[0056] 2.3.2 Based on the OR value, set a significance mark for each gene according to the preset rules, and add the significance mark of each gene to the histogram.

[0057] The following rules were used to set significance marks for each gene: "*" indicates 0.05 >= Padj > 0.01, "**" indicates 0.01 >= Padj > 0.001, "***" indicates 0.001 >= Padj > 0.0001, and "****" indicates Padj <= 0.0001. In other words, Padj <= 0.05 indicates a significantly preferred gene segment.

[0058] Using the for function of the base package and the annotate function of the ggplot2 package, significant markers were added cyclically in the histogram. Genes with OR values ​​> 1 were added to the histogram of neutralizing antibodies, and genes with OR values ​​< 1 were added to the histogram of non-neutralizing antibodies.

[0059] The bar graph results of the IGHV group are as follows: Figure 4 As shown, in particular, IGHV1-58, IGHV3-66, IGHV3-53, IGHV1-2, and IGHV1-69 had significant selection preferences among neutralizing antibodies, while IGHV3-30, IGHV1-46, IGHV3-48, IGHV3-13, IGHV3-30-3, IGHV3-64D, and IGHV2-26 had significant selection preferences among non-neutralizing antibodies.

[0060] 2.4 Visualization of VJ gene selection preference analysis, specifically including the following steps:

[0061] 2.4.1 Combine the V and J gene selection frequency and preference analysis results of IGHV, IGHJ, IGKV, IGKJ, IGLV, and IGLJ as described in 2.2.2.

[0062] 2.4.2 Convert the OR values ​​of all genes to Log2OR and the Padj values ​​to -Log 10 Padj.

[0063] 2.4.3 Using the ggplot2 package's ggplot function and geom_point function, with Log2OR as the horizontal axis, -Log 10 Padj is the vertical axis, and a scatter plot is drawn, and genes with significant preferences are highlighted in different colors. Among them, genes with OR values ​​> 1 are marked in red, indicating that they are significantly enriched in neutralizing antibodies, and genes with OR < 1 are marked in blue, indicating that they are significantly enriched in non-neutralizing antibodies. At the same time, genes are grouped according to heavy chain, light chain, V gene, and J gene. The volcano plot results are as follows Figure 5 shown.

[0064] S3. VJ gene pairing frequency preference analysis and visualization

[0065] 3.1. VJ gene pairing frequency statistics: Group by heavy chain and light chain, and count the pairing frequencies of each V gene and each J gene in neutralizing and non-neutralizing antibodies. The specific steps include:

[0066] 3.1.1 Grouped by heavy chain, kappa chain (light chain), and lambda chain (light chain), in the VJ pairing data of 1679 human anti-novel coronavirus neutralizing antibodies and 1349 non-neutralizing antibodies as described in 1.3, further, based on the complete immunoglobulin V(D)J gene annotation information in the IMGT database, genes not included in the IMGT annotation were screened out.

[0067] 3.1.2 Group by heavy chain, kappa chain (light chain), and lambda chain (light chain), and count the J genes and their frequencies of each V gene paired with each other in neutralizing antibodies and non-neutralizing antibodies. Taking the heavy chain group as an example, the heavy chain has 49 V genes and the frequency of pairing with 6 J genes is counted. The results of 10 V genes are as follows Figure 6 shown.

[0068] 3.2. Statistical analysis of VJ gene pairing preferences: Calculate the pairing preference P value, OR value, and Padj value for each VJ gene pair in neutralizing and non-neutralizing antibodies. This specifically includes the following steps:

[0069] 3.2.1 Group by heavy chain, kappa chain (light chain), and lambda chain (light chain). Based on the paired frequency statistics in 3.1.2, use the fisher.test function in the stats package to perform a Fisher exact test on the frequency ratios of the same V gene pairing with different J genes in neutralizing and non-neutralizing antibodies, and calculate the P value for differences in the two ratios.

[0070] Specifically, in R software, run "fisher.test(matrix(c(A,B,C…N,a,b,c…n),ncol=2,nrow=x))[[1]]" to output P value, wherein A, B, C…N are the frequency of each J gene paired with the same V gene in neutralizing antibodies, respectively, a, b, c…n are the frequency of the corresponding J gene paired with the V gene in non-neutralizing antibodies, respectively, and x is the total number of paired J genes paired with the V. Taking IGHV3-30 as an example, the J genes paired with IGHV3-30 in neutralizing antibodies are IGHJ1, IGHJ2, IGHJ3, IGHJ4, IGHJ5 and IGHJ6, with frequencies of 2, 1, 20, 90, 14 and 50, respectively, and the frequencies of IGHV3-30 paired with the corresponding six J genes in non-neutralizing antibodies are 1, 2, 17, 125, 10 and 32, respectively, thus P=fisher.test(matrix(c(2,1,20,90,14,50,1,2,17,125,10,32),ncol=2,nrow=6))[[1]]=0.0328581<0.05, that is, the composition ratio of J genes paired with IGHV3-30 is different between neutralizing antibodies and non-neutralizing antibodies.

[0071] 3.2.2 Further, screen out P<0.05, that is, the V genes with different two composition ratios, use the fisher.test function of the stats package to compare the frequency of each V gene with each J gene pairwise, and obtain the OR value and P value of pairwise comparison.

[0072] Specifically, in R software, run "fisher.test(matrix(c(n,m,N,M),ncol=2,nrow=2))[[3]]" to output OR value, and run "fisher.test(matrix(c(n,m,N-n,M-m),ncol=2,nrow=2))[[1]]" to output P value, wherein n is the frequency of a V gene paired with a J gene in neutralizing antibodies, m is the frequency of the V gene paired with the J gene in non-neutralizing antibodies, N is the frequency of the V gene paired with any J gene in neutralizing antibodies, and M is the frequency of the V gene paired with any J gene in non-neutralizing antibodies.

[0073] Taking IGHV3-30 as an example, the J genes paired with IGHV3-30 in neutralizing antibodies are IGHJ1, IGHJ2, IGHJ3, IGHJ4, IGHJ5, and IGHJ6, totaling 6 J genes, with frequencies of 2, 1, 20, 90, 14, and 50, respectively, totaling 177, that is, there are 177 strains of neutralizing antibodies in which IGHV3-30 is paired with any J gene; while in non-neutralizing antibodies, the frequencies of IGHV3-30 pairing with the corresponding 6 J genes are 1, 2, 17, 125, 10, and 32, respectively, totaling 187, that is, there are 187 strains of non-neutralizing antibodies in which IGHV3-30 is paired with any J gene. Among them, the OR of the pairing of IGHV3-30 and IGHJ4 in neutralizing antibodies = fisher.test(matrix(c(90,125,177-90,187-125),ncol=2,nrow=2))[[3]]=0.5140575, P = fisher.test(matrix(c(90,125,177-90,187-125),ncol=2,nrow=2))[[1]]=0.002037995. All P values ​​were sorted from small to large using the order function of the base package and then corrected using the p.adjust function of the stats package using the "fdr" method to obtain the Padj value.

[0074] Taking the heavy chain group as an example, the statistical analysis results and the corresponding pairing frequencies are as follows: Figure 7 As shown, the composition ratio test showed that the five V genes had different pairing with the J gene in neutralizing antibodies and non-neutralizing antibodies. Further pairwise comparison results showed that the IGHV3-30 gene significantly favored pairing with the IGHJ6 gene in neutralizing antibodies, but significantly favored pairing with IGHJ4 in non-neutralizing antibodies; the IGHV3-48 gene significantly favored pairing with IGHJ4 in neutralizing antibodies, but significantly favored pairing with IGHJ6 in non-neutralizing antibodies; and the IGHV4-61 gene significantly favored pairing with the IGHJ6 gene in non-neutralizing antibodies.

[0075] 3.3 Visualization of VJ gene pairing frequencies: Display the pairing frequencies of VJ gene pairs in neutralizing and non-neutralizing antibodies in the form of a heat map, and highlight the results of the first preference analysis. The specific steps include:

[0076] 3.3.1 Based on the frequency statistics results described in 3.1.2, use the Heatmap function and Legend function of the Complexheatmap package to display the frequency statistics results in the form of a heat map.

[0077] 3.3.2 On the basis of the heat map as described in 3.3.1, the results of the constituent ratio test as described in 3.2.1 are displayed at the row name of each V gene according to the following rules: "*" represents 0.05 >= Padj > 0.01, "**" represents 0.01 >= Padj > 0.001, "***" represents 0.001 >= Padj > 0.0001, and "****" represents Padj <= 0.0001.

[0078] 3.3.3 On the basis of the heat map as described in 3.3.2, the results of the pairwise comparison as described in 3.2.2 are displayed in the grid of each VJ pair according to the following rules: if the OR of the pairwise comparison is greater than 1, it is displayed in the neutralizing antibody group, and if the OR is less than 1, it is displayed in the non-neutralizing antibody group. The heat map results of the heavy chain VJ pair frequency are shown as follows. Figure 8

[0079] 3.4, Visualization of VJ gene pair frequency significant bias genes, display the specific VJ gene pair of significant genes in the form of a chord diagram, and highlight the second preference analysis results, the specific steps are as follows:

[0080] 3.4.1 The pair frequency data of the V gene with a difference in constituent ratio as described in 3.2.2 is displayed in the form of a chord diagram by using the chordDiagram function of the cirlize package, and the frequency data is standardized by using the scale parameter.

[0081] 3.4.2 In the chord diagram as described in 3.4.1, the VJ gene pairs with a difference in the pairwise comparison results as described in 3.2.2 are highlighted, specifically, on the basis of all links being transparent, the link of the VJ gene pair with OR > 1 and Padj < 0.05 is reduced in transparency in the neutralizing antibody group, and the link of the VJ gene pair with OR < 1 and Padj < 0.05 is reduced in transparency in the non-neutralizing antibody group, thereby being highlighted on the chord diagram.

[0082] For example, as shown in Figure 9 , in the neutralizing antibody, the IGHV3-30 gene is significantly paired with IGHJ6, and the IGHV3-48 gene is significantly paired with IGHJ4; and in the non-neutralizing antibody, the IGHV3-30 gene is significantly paired with IGHJ4, the IGHV3-48 gene is significantly paired with IGHJ6, and the IGHV4-61 gene is significantly paired with IGHJ6.

[0083] The embodiment of the present application also discloses a new method for early prediction of anti-new coronavirus neutralizing antibody production based on peripheral blood B cell VJ gene frequency and preference, and the prediction method of the anti-new coronavirus neutralizing antibody comprises the following steps S401-S404: ​

[0084] S401, obtaining the sequencing results of the B cell VJ genes in the peripheral blood of healthy people from a public dataset, and calculating the read proportion of each gene as a baseline frequency, which is used to represent the baseline level of healthy people.

[0085] In the formula, the sequencing results of the B cell VJ genes in the peripheral blood of healthy people are obtained from a public dataset (Briney B, Inderbitzin A, Joyce C, Burton DR. Commonality despite exceptional diversity in the baseline human antibody repertoire. Nature. 2019; 566(7744): 393-397.), and the read proportion of each gene is calculated as a baseline frequency, which is used to represent the baseline level of healthy people. For example, the read of IGHV3-53 is 6344873, the read of all V genes is 363506788, and the baseline frequency of IGHV3-53 is 1.7455%.

[0086] S402, obtaining the sequencing data of the peripheral blood of the subject, and calculating the read proportion of each VJ gene in the peripheral blood of the subject as an actual frequency according to the sequencing data, which is used to represent the actual level of the subject.

[0087] In the formula, the sequencing data of the peripheral blood of the subject can be obtained by performing total RNA extraction, antibody gene sequence amplification and second-generation sequencing on the collected peripheral blood of the subject according to a standard method. Specifically, the peripheral blood of the subject is collected, total RNA is extracted, antibody gene sequences are amplified, and second-generation sequencing is performed to obtain sequencing data according to a standard method (Briney B, Inderbitzin A, Joyce C, Burton DR. Commonality despite exceptional diversity in the baseline human antibody repertoire. Nature. 2019; 566(7744): 393-397.), and the read proportion of each VJ gene in the peripheral blood of the subject is calculated as an actual frequency according to the sequencing data, which is used to represent the actual level of the subject.

[0088] S403, determining the gene fragments significantly preferred to be selected in neutralizing antibodies as favorable genes and the gene fragments significantly preferred to be selected in non-neutralizing antibodies as unfavorable genes according to the results of the VJ gene selection and pairing preference analysis; wherein Padj<=0.05 represents the gene fragments significantly preferred to be selected.

[0089] According to the aforementioned VJ gene selection and pairing preference analysis results, the gene segments that are significantly preferred in neutralizing antibodies are favorable genes, such as IGHV1-58, IGHV3-66, IGHV3-53, IGHV1-2, IGHV1-69 and other genes; the gene segments that are significantly preferred in non-neutralizing antibodies are unfavorable genes, such as IGHV3-30, IGHV1-46, IGHV3-48, IGHV3-13, IGHV3-30-3, IGHV3-64D, IGHV2-26 and other genes.

[0090] S404: Using the Fisher exact test and taking p = 0.05 as the threshold, calculate the frequency difference between the actual frequency of each favorable gene and unfavorable gene measured in the peripheral blood of the subject that is higher / lower than the baseline frequency corresponding to the gene; based on the frequency difference, determine whether the probability of the subject producing neutralizing antibodies is higher than the baseline level of healthy people.

[0091] Among them, it is calculated whether the actual frequency of each favorable gene and unfavorable gene measured in the peripheral blood of the subject is significantly increased or decreased compared with its corresponding baseline frequency. If the actual frequency of the favorable gene is significantly higher than the baseline frequency, or the actual frequency of the unfavorable gene is significantly lower than the baseline frequency, the probability of the subject producing neutralizing antibodies is higher; if the actual frequency of the favorable gene is significantly lower than the baseline frequency, or the actual frequency of the unfavorable gene is significantly higher than the baseline frequency, the probability of the subject producing neutralizing antibodies is lower.

[0092] That is to say, if the frequency difference between the actual frequency of the favorable gene and the baseline frequency corresponding to the favorable gene reaches a specified threshold, or the frequency difference between the actual frequency of the unfavorable gene and the baseline frequency corresponding to the unfavorable gene reaches a specified threshold, it is determined that the probability of the subject producing neutralizing antibodies is higher than the baseline level of healthy people; if the frequency difference between the actual frequency of the favorable gene and the baseline frequency corresponding to the favorable gene reaches a specified threshold, or the frequency difference between the actual frequency of the unfavorable gene and the baseline frequency corresponding to the unfavorable gene reaches a specified threshold, it is determined that the probability of the subject producing neutralizing antibodies is lower than the baseline level of healthy people.

[0093] For example, if a subject has 3623 reads of IGHV3-53 and 125633 reads of all V genes, the actual frequency of IGHV3-53 is 932 / 149754 = 2.8838%, which is 1.1383% higher than the baseline frequency (1.7455%). The Fisher exact test calculates p < 2.2e-16, which is statistically significant. Therefore, it can be predicted that the probability of this subject producing neutralizing antibodies is higher than the baseline level of healthy people.

[0094] An embodiment of the present invention discloses an electronic device, comprising a memory storing executable program code and a processor coupled to the memory; wherein the processor calls the executable program code stored in the memory to execute the VJ gene preference analysis method for neutralizing antibodies against the new coronavirus described in the above embodiments.

[0095] An embodiment of the present invention also discloses a computer-readable storage medium that stores a computer program, wherein the computer program enables a computer to execute the VJ gene preference analysis method for neutralizing antibodies against the new coronavirus described in the above embodiments.

[0096] An embodiment of the present invention discloses an electronic device, comprising a memory storing executable program code and a processor coupled to the memory; wherein the processor calls the executable program code stored in the memory to execute the method for predicting anti-new coronavirus neutralizing antibodies described in the above embodiments.

[0097] An embodiment of the present invention also discloses a computer-readable storage medium storing a computer program, wherein the computer program enables a computer to execute the method for predicting anti-new coronavirus neutralizing antibodies described in the above embodiments.

[0098] The purpose of the above embodiments is to exemplify and deduce the technical solution of the present invention, and to fully describe the technical solution, purpose and effect of the present invention. Its purpose is to enable the public to have a more thorough and comprehensive understanding of the disclosed content of the present invention, and it does not limit the scope of protection of the present invention.

[0099] The above embodiments are not exhaustive and may include many other embodiments not listed above. Any replacements and improvements made without violating the concept of the present invention are within the scope of protection of the present invention.

Claims

1. A method for analyzing the VJ gene preference of anti-new coronavirus neutralizing antibodies, characterized in that: include: Obtain antibody data from a coronavirus antibody database, and screen the antibody data for neutralizing and non-neutralizing antibodies against the novel coronavirus; The neutralizing antibodies and the non-neutralizing antibodies were grouped according to the six genes of IGHV, IGHJ, IGKV, IGKJ, IGLV, and IGLJ, and the selection frequency of each gene in the neutralizing antibodies and the non-neutralizing antibodies was counted; Calculate the selection preference OR value, P value, and Padj value of each gene in the neutralizing antibody; Sort by OR value from large to small, and display the selection frequency of each gene in the neutralizing antibody and the non-neutralizing antibody in the form of a horizontal grouped bar graph; The OR value and Padj value of each gene are displayed in the form of a volcano plot, highlighting the preferred genes; The calculation of the selection preference OR value, P value, and Padj value of each gene in the neutralizing antibody comprises: The OR value and P value of each gene were calculated using the selection frequency of each gene in the neutralizing antibodies and the non-neutralizing antibodies and the total number of the neutralizing antibodies and the non-neutralizing antibodies; After sorting the P values ​​from small to large, multiple corrections were performed to obtain the Padj value; The selection frequencies of the genes in the neutralizing antibodies and the non-neutralizing antibodies are displayed in the form of a horizontal grouped bar graph by sorting the genes from large to small according to the OR value, including: The frequencies of the neutralizing antibodies and the non-neutralizing antibodies are plotted on both sides of a bar graph using the ggplot2 function and the geom_bar function; According to the OR value, a significance mark is set for each gene according to a preset rule, and the significance mark of each gene is added to the histogram; The OR value and Padj value of each gene are displayed in the form of a volcano plot, highlighting the preferred genes, including: The OR values ​​of all genes were converted to Log2OR values, and the Padj values ​​were converted to -Log 10 Padj value; Using ggplot2 function and geom_bar function, the horizontal axis is Log2OR value and the horizontal axis is -Log 10 The Padj value is used as the vertical axis to draw a scatter plot, and the genes with significant preferences are highlighted in different colors.

2. A method for analyzing the VJ gene preference of neutralizing antibodies against the novel coronavirus, characterized in that: include: Obtain antibody data from a coronavirus antibody database, and screen the antibody data for neutralizing and non-neutralizing antibodies against the novel coronavirus; Grouping by heavy chain and light chain, and counting the pairing frequencies of each V gene and each J gene in the neutralizing antibody and the non-neutralizing antibody; Calculate the pairing preference P value, OR value, and Padj value of each VJ gene pair in the neutralizing antibody and the non-neutralizing antibody; The pairing frequencies of the VJ gene pairs in the neutralizing antibodies and the non-neutralizing antibodies are displayed in the form of a heat map, and the first preference analysis results are highlighted; Use chord diagrams to display the specific VJ gene pairings of significant genes and highlight the results of the second preference analysis; The calculating of the pairing preference P value, OR value, and Padj value of each VJ gene pair in the neutralizing antibody and the non-neutralizing antibody comprises: The frequency of pairing of each V gene with each J gene in the neutralizing antibody and the non-neutralizing antibody, as well as the total number of V genes, was used to analyze the difference in the constituent ratios of the two groups using the Fisher's exact test to obtain the P value. V genes with P values ​​less than the specified threshold were screened, and each significantly preferred V gene was compared with each J gene pairwise using the Fisher exact test to obtain the pairing preference OR value and P value in the neutralizing antibody. The P values ​​were sorted from small to large, and multiple corrections were performed to obtain the Padj value.

3. The method for analyzing the VJ gene preference of anti-new coronavirus neutralizing antibodies according to claim 2, wherein: The highlighting of the first preference analysis result includes: Display the sudden change of the composition ratio test result in the heat map row name in the specified symbol form; The results of pairwise comparisons are displayed in the corresponding gene pairs in the neutralizing antibody and non-neutralizing antibody groups according to the OR value results.

4. The method for analyzing the VJ gene preference of anti-new coronavirus neutralizing antibodies according to claim 2, wherein: The highlighting of the second preference analysis result includes: In the pairwise comparison results, the links of the VJ gene pairs whose Padj values ​​are greater than or equal to the specified threshold are displayed transparently. In the chord diagrams of neutralizing antibodies and non-neutralizing antibodies, the transparency of the links of the VJ gene pairs whose Padj values ​​are less than the specified threshold is reduced according to the OR value results.

5. An electronic device, characterized in that It includes a memory storing executable program code and a processor coupled to the memory; the processor calls the executable program code stored in the memory to execute the anti-new coronavirus neutralizing antibody VJ gene preference analysis method according to any one of claims 1 to 4.

6. A method for predicting neutralizing antibodies against the novel coronavirus, characterized in that: The method is performed according to the anti-new coronavirus neutralizing antibody VJ gene preference analysis method according to any one of claims 1 to 4, wherein the prediction method comprises: The VJ gene sequencing results of healthy human peripheral blood B cells were obtained from a public dataset, and the read ratio of each gene was calculated as the baseline frequency, which was used to characterize the baseline level of healthy people. Obtain sequencing data of the subject's peripheral blood, and calculate the read ratio of each VJ gene in the subject's peripheral blood as the actual frequency based on the sequencing data. The actual frequency is used to represent the actual level of the subject; Based on the aforementioned VJ gene selection and pairing preference analysis results, gene segments that are significantly preferred in neutralizing antibodies are determined as favorable genes, and gene segments that are significantly preferred in non-neutralizing antibodies are determined as unfavorable genes; wherein Padj <= 0.05 indicates a significantly preferred gene segment; Using the Fisher exact test and taking p = 0.05 as the threshold, the frequency difference between the actual frequency of each favorable gene and unfavorable gene measured in the peripheral blood of the subject is calculated as being higher / lower than the baseline frequency corresponding to the gene; based on the frequency difference, it is judged whether the probability of the subject producing neutralizing antibodies is higher than the baseline level of healthy people.

Citation Information

Patent Citations

  • Preparation method of pig-source single-chain genetic engineering antibody against foot-and-mouth disease viruses

    CN110372791A

  • Identifying The Pairing Of Variable Domains Of Light And Heavy Chains Of Antibodies

    US20210166787A1