A non-invasive diagnosis method for infectious diseases based on deep learning

By extracting the library-level characteristics and sequence-level characteristics of the antibody library and building a deep learning model, the problems of insufficient feature extraction and low prediction accuracy in the prior art are solved, and higher disease diagnosis accuracy is achieved.

CN114512244BActive Publication Date: 2025-05-06GUANGDONG GENERAL HOSPITAL
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111496240.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-12-09
Publication Date
2025-05-06
Estimated Expiration
2041-12-09

AI Technical Summary

Technical Problem

The existing disease diagnosis model based on the antibody library has shortcomings in the problems of insufficient feature extraction and low prediction accuracy.

Method used

Using a deep learning-based method, the library-level characteristics and sequence-level characteristics of the antibody group library are extracted, and the library-level model and sequence-level model are constructed, and the diagnostic accuracy is improved through feature screening and model integration.

Benefits of technology

It significantly improves the prediction accuracy of the disease diagnosis model and can more effectively mine disease association information in the high-diversity antibody library.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114512244B_ABST
    Figure CN114512244B_ABST
Patent Text Reader

Abstract

The present invention provides a non-invasive diagnosis method and system for infectious diseases based on deep learning: including the following features: obtaining sample antibody repertoire data, determining a training set and a test set; extracting repertoire level features and sequence level features for the antibody repertoire data; using the extracted repertoire level features and the extracted sequence level features to respectively construct an initial prediction model; using the training set to train the initial prediction model, screening out the repertoire level features and sequence level features that need to be retained; by inputting the screened repertoire level features and sequence level features into the initial prediction model, respectively, an optimized prediction model is obtained; and the optimized prediction model is used to evaluate the performance of the test set. This method can effectively mine disease association information hidden in a high-diversity antibody repertoire, and effectively improve the diagnostic accuracy of the prediction model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of artificial intelligence platform construction, and in particular to a non-invasive diagnosis method for infectious diseases based on deep learning. Background Art

[0002] Early diagnosis of various infectious diseases is a key means to effectively control sudden disease pandemics. For example, the symptoms caused by the new coronavirus infection are very similar to those of other respiratory diseases, making its diagnosis extremely complicated. Due to its strong infectivity and mortality rate, rapid and accurate diagnostic methods are of great significance to controlling the pandemic and treating patients. Other causes such as meningococcal pathogens that can cause sepsis have early symptoms similar to those caused by more self-limiting viral diseases, and delays in diagnosis and treatment can lead to death or serious complications. Therefore, screening biomarkers that can accurately distinguish the source of disease infection is the key to developing high-precision non-invasive diagnostic models.

[0003] The antibody repertoire, that is, the complete set of antibodies in a tissue or individual, is a key component of the adaptive immune system. The antibody repertoire of an adult contains about 10 13 A unique antibody or B cell receptor (B cell receptor, BCR). Since each antibody specifically binds to a specific epitope of a single antigen, the high diversity of the antibody repertoire can ensure that the antibodies in it can recognize a wide enough range of pathogens, thereby providing protection to the human body. Due to the memory effect of the immune repertoire, the antibody molecules in it can reflect both the individual's immune history and the current immune status. Therefore, the antibody repertoire has great potential for diagnosing various immune-mediated diseases.

[0004] The huge diversity of antibody repertoires provides potential for disease diagnosis, but it also brings challenges. With the development of high-throughput sequencing technology, tens of thousands of antibody molecules can be obtained in a single sequencing, of which only 0.01% or even less may be related to the individual's current disease state. Extracting associated features that can reflect the disease state from the highly diverse antibody repertoire and constructing a high-precision prediction model that can reflect the relationship between the features and the disease state is the key to developing a disease diagnosis model based on the antibody repertoire.

[0005] Currently, there are a few studies devoted to constructing disease prediction methods based on antibody repertoires. For example, Ostmeyer et al. first divided the complementarity determining region 3 (CDR3) amino acid sequence of antibody molecules into small fragments of equal length (k-mer), and characterized these small amino acid fragments with 5 amino acid physicochemical properties, and finally predicted the disease status of individuals based on the maximum fragment score (Ostmeyer J et al. Statistical classifiers for diagnosing disease from immune repertoires: a case study using multiple sclerosis. BMC bioinformatics, 2017, 18 (1): 1-10). The accuracy (correct judgment rate) of the leave-one-out cross-test of the training set (23 samples) was 87%, and the accuracy of the independent test (102 samples) was 72%. Eliyahu et al. used the similarity clusters of antibody molecules to characterize each antibody repertoire and constructed a logistic regression model to distinguish between individuals with spontaneous hepatitis C virus clearance and chronic infection (20 samples), with a correct judgment rate of 91% (Eliyahu S et al. Antibody repertoire analysis of hepatitis C virus infections identifies immune signatures associated with spontaneous clearance. Frontiers in immunology, 2018, 9: 3004). Shemesh et al. used CDR3 similarity clusters to characterize the antibody repertoire and constructed a logistic regression model to predict celiac disease patients (100 samples), with an F1 score of 0.85 (Shemesh O et al. Machine learning analysis of B-cell receptor repertoires stratifies celiac disease patients and controls. Frontiers in immunology, 2021, 12: 633). Huang et al. used public antibody molecules (public clones) to characterize the antibody repertoire and constructed a logistic regression model to predict immunoglobulin A nephropathy (112 samples), with an area under the receiver operating characteristic curve (AUC) score of only 0.7384 (Huang C et al. The landscape and diagnostic potential of T and B cell repertoire in Immunoglobulin A Nephropathy. Journal of autoimmunity, 2019, 97: 100-107). Shoukat et al. characterized the antibody repertoire using the frequency of occurrence of antibody molecule k-mers, and used a hierarchical clustering method to predict an individual's SARS-CoV-2 infection status (55 samples), with a correct judgment rate of only 47% (Shoukat MS et al. Use of machine learning to identify a T cell response to SARS-CoV-2. Cell Reports Medicine, 2021, 2(2): 100192).

[0006] In general, these methods use less training data to build prediction models. In addition, they only use a single feature to characterize the antibody repertoire, resulting in insufficient information extraction. Finally, these methods mostly use traditional machine learning algorithms, and their learning performance is not enough to adapt to the ultra-high diversity of the antibody repertoire. These problems have led to low prediction accuracy of current disease diagnosis models based on antibody repertoires. Summary of the invention

[0007] In view of the shortcomings of insufficient feature extraction of antibody repertoire characterization and low prediction accuracy of disease diagnosis model in the prior art, the present invention provides a non-invasive diagnosis method for infectious diseases based on deep learning, comprising the following steps:

[0008] (1) Collect a certain number of pathogen-infected samples and a certain number of healthy control samples, and randomly assign all the collected pathogen-infected samples and healthy control samples into training sets and test sets according to a certain ratio;

[0009] (2) obtaining antibody repertoire data of the sample in step (1);

[0010] (3) extracting repertoire-level features and sequence-level features for the antibody repertoire data;

[0011] (4) using the extracted repertoire-level features and the extracted sequence-level features to construct an initial repertoire-level model and an initial sequence-level model, respectively;

[0012] (5) using the training set in step (1) to train the initial repertoire level model and the initial sequence level model respectively, and screen out the repertoire level features and sequence level features that need to be retained;

[0013] (6) obtaining an optimized repertoire level model by inputting the screened repertoire level features into the initial repertoire level feature model, and obtaining an optimized sequence level model by inputting the screened sequence level features into the initial sequence level feature model;

[0014] (7) Use the optimized repertoire-level model and the optimized sequence-level model to evaluate the performance of the test set in step (1).

[0015] Furthermore, in step (3), the extracted repertoire-level features include at least: 1) V gene segment usage frequency, 2) VJ gene segment combined usage frequency, 3) α Evenness index, 4) diversity index; The extracted sequence-level features include: 1) at least two physicochemical characteristics of the CDR3 amino acid sequence: KF8 and F5, 2) sequence frequency, and 3) read complexity.

[0016] Furthermore, the physicochemical characteristics of the CDR3 amino acid sequence can also be selected from at least one of Mole%_Tiny, Mole%_Aliphatic, Mole%_Aromatic, Mole%_NonPolar, Mole%_Basic, F1, F6, KF8, F5, KF1, KF2 and KF3. In addition, the physicochemical characteristics of the CDR3 amino acid sequence can also be other physicochemical characteristics obtained by statistically analyzing the physicochemical characteristics of each amino acid constituting the CDR3 amino acid sequence, such as polarity, non-polarity, pH value, whether it is charged and hydrophobicity.

[0017] Further, in step (5), based on all the extracted repertoire-level features, the repertoire-level features retained after screening include V gene segment usage frequency, VJ gene segment combined usage frequency, α Evenness index and diversity index.

[0018] Furthermore, the screening of all the extracted repertoire-level features is carried out in the following manner: based on the constructed initial repertoire-level feature model and based on the correct judgment rate of 10 cross-tests of the training set, the importance of the subsets under all possible combinations of all the extracted repertoire features is evaluated, and the repertoire-level features that need to be retained are determined based on the principle of the highest training accuracy.

[0019] Further, in step (5), based on all the extracted sequence-level features, the sequence-level features retained after screening include read complexity, sequence frequency, KF8 and F5.

[0020] Furthermore, the screening of all extracted sequence-level features is performed in the following manner: The selection of sequence-level features is performed in the following manner: Based on the constructed initial sequence-level feature model, and based on the correct judgment rate of 10 cross-tests in the training set, the importance of each sequence-level feature is evaluated, and then each sequence-level feature is arranged in descending order of importance, and the sequence-level features are introduced one by one. If the introduced sequence-level feature cannot increase the model training accuracy, the sequence-level feature is eliminated, and the next sequence-level feature is introduced until all the extracted sequence-level features are traversed, and finally the sequence-level features that need to be retained are determined.

[0021] Furthermore, the initial sequence-level model includes an input layer, a first convolutional layer, a second convolutional layer, a third convolutional layer, a fully connected layer, an average pooling layer, and a result output layer; the input layer is a 160×160 two-dimensional matrix composed of all extracted sequence-level features, each row represents an antibody molecule, and each column represents a sequence-level feature; the first convolutional layer consists of 512 1×160 filters, the step size is set to 160, the padding parameter is set to 0, the activation function is ReLU, and then a 1×2 average pooling layer is connected; the second convolutional layer consists of 256 3×1 filters, the step size and padding are The parameters are all set to 1, the activation function is ReLU, and it is followed by a 1×2 average pooling layer; the third convolutional layer consists of 128 3×1 filters, the step size and padding parameters are all set to 1, the activation function is ReLU, and it is followed by a 1×2 average pooling layer; the third convolutional layer is followed by three fully connected layers based on the ReLU activation function, with the number of nodes being 2560, 1000, and 64 respectively; the activation function of the resulting output layer is Sigmoid; the SLM model parameter optimization uses the AdamOptimizer optimizer with a learning rate of 0.01, and the training epoch and batch size are 250 and 50 respectively.

[0022] Furthermore, in step (7), during the evaluation process, an integrated probability is output, which is the average of the output probability of the optimized RLM model and the output probability of the optimized SLM model. If the integrated probability is greater than 0.5, the test sample is a pathogen-infected sample, otherwise it is a healthy sample.

[0023] Furthermore, the number of collected pathogen-infected samples and healthy control samples is no less than 100, and the ratio is 2:1 to 5:1.

[0024] The present invention also provides a non-invasive diagnosis system for infectious diseases based on deep learning, the system comprising: (1) a training set and a test set selection module; (2) an antibody repertoire data acquisition module; (3) a repertoire level feature extraction module; (4) a sequence level feature extraction module; (5) an initial prediction model construction module; (6) a repertoire level feature screening module; (7) a sequence level feature screening module; (8) a prediction model performance evaluation module, wherein:

[0025] The training set and test set selection module collects a certain number of pathogen-infected samples and a certain number of healthy control samples, and randomly allocates all the collected pathogen-infected samples and healthy control samples as training sets and test sets according to a certain ratio;

[0026] The antibody repertoire data acquisition module acquires the antibody repertoire data of the samples collected by the training set and test set selection modules;

[0027] The repertoire level feature extraction module and the sequence level feature extraction module respectively extract the repertoire level features and sequence level features of the antibody repertoire data acquired by the antibody repertoire data acquisition module;

[0028] The initial prediction model building module builds an initial repertoire-level feature model and an initial sequence-level feature model by inputting all extracted repertoire-level features and sequence-level features;

[0029] The repertoire-level feature screening module and the sequence-level feature screening module respectively train the initial repertoire-level feature model and the initial sequence-level feature model using the training set to screen out the repertoire-level features and sequence-level features that need to be retained, and then obtain an optimized repertoire-level model by inputting the screened repertoire-level features into the initial repertoire-level feature model, and obtain an optimized sequence-level model by inputting the screened sequence-level features into the initial sequence-level feature model;

[0030] The prediction model performance evaluation module uses the optimized repertoire level model and the optimized sequence level model to perform performance evaluation on the test set respectively.

[0031] In the present invention, the "Index" in the primer sequence is a base sequence of a certain length, which is an essential component of high-throughput sequencing. Although the length of the Index of each high-throughput measurement platform is different, the ultimate purpose is the same, which is to split the sequencing data and identify different samples. For example, the length of the Index can be a fixed base sequence of 6bp, a fixed base sequence of 8bp, or a fixed base sequence of 10 to 12bp.

[0032] Technical effects of the present invention: Compared with the traditional diagnostic model based on antibody repertoire, the method proposed in the present invention has the following advantages: (1) The data used to train the diagnostic model includes 602 samples, which is the largest amount of data in the current existing models; (2) Compared with the traditional method that only uses a single feature to characterize the antibody repertoire, the present invention extracts more comprehensive features to characterize the antibody repertoire, especially extracts two different levels of features, namely, repertoire level features and sequence level features, to characterize the antibody repertoire, which can effectively mine the disease association information hidden in the high-diversity antibody repertoire; (3) The present invention finds that not all features can promote the improvement of the prediction accuracy of the diagnostic model, so feature selection is performed on the extracted multiple features, and the optimal feature subsets are selected for the repertoire level features and sequence level features for the construction of the diagnostic model; (4) Because the repertoire level features are one-dimensional features and the sequence level features are two-dimensional matrix features, the present invention designs two different deep neural network models for these two different types of features, and effectively improves the model diagnosis accuracy by integrating the results of the two models. BRIEF DESCRIPTION OF THE DRAWINGS

[0033] Figure 1 An artificial intelligence-based disease diagnosis model constructed for the present invention. DETAILED DESCRIPTION

[0034] The following description of the embodiments is only used to help understand the method of the present invention and its core idea. It should be noted that for those of ordinary skill in the art, without departing from the principles of the present invention, the present invention may also be improved and modified, and these improvements and modifications also fall within the scope of protection of the present invention. The following description of the disclosed embodiments enables professionals and technicians in this field to implement or use the present invention. It will be apparent to professionals and technicians in this field that various modifications to these embodiments, and the general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention will not be limited to these embodiments shown herein, but can be applied to a wider range consistent with the principles and novel features disclosed herein.

[0035] If no specific techniques or conditions are specified in the examples, the techniques or conditions described in the literature in the art or the product instructions are followed.

[0036] Example 1: Training set and test set selection

[0037] A certain number of pathogen-infected samples and a certain number of healthy control samples are collected, and all the collected pathogen-infected samples and healthy control samples are randomly allocated as training sets and test sets according to a certain ratio. The number of samples in each category is not less than 100, preferably not less than 200. The ratio is 2:1 to 5:1, preferably 2.5:1 to 4.5:1.

[0038] Example 2: Antibody repertoire data acquisition

[0039] (1) Sample collection

[0040] 1 ml of human peripheral blood samples (PBMCs) were collected from the pathogen-infected subjects and healthy controls in Example 1, collected in anticoagulant tubes containing EDTA, and stored at room temperature for no more than 4 hours. PBMCs were separated by density gradient centrifugation using lymphocyte separation fluid (Axis-Shield, 1114547), and the separated cells were dissolved in RLT buffer (Qiagen), 1% β-mercaptoethanol (Sigma) was added, and then stored at -80°C for short-term storage.

[0041] (2) RNA extraction and reverse transcription

[0042] RNA was extracted using the RNeasy Mini Kit (Qiagen, 74106) according to the instructions, and RNA concentration was measured using NanoDrop 2000c. 500 ng of each sample was taken for reverse transcription.

[0043] Reverse transcription was performed using Thermo SuperScript TM II Reverse Transcriptase and Takara RACE 5' / 3'Kit, follow the instructions for specific methods. Prepare the following reaction system 1:

[0044]

[0045] Place the reaction in a PCR instrument at 72°C for 1 min, then transfer to ice and place for 2 min;

[0046] Then prepare the following reaction system 2:

[0047]

[0048] Reaction conditions: 42°C for 90 min, 70°C for 15 min, 4°C.

[0049] (3) Dual-end barcode dual-end UMI amplification

[0050]

[0051] Primer F1 / R1 is a mixture of 8 V region specific primers in equal amounts, and its sequence is SEQ ID No1-8.

[0052] PCR program: pre-denaturation at 95°C for 3 min, (98°C for 20 s, 60°C for 15 s, 72°C for 15 s) 2 cycles, 72°C for 3 min.

[0053] After the first round of PCR, the PCR product was purified using Agencourt AMPure XP magnetic beads, and the operation was carried out according to the instructions, and the supernatant was aspirated for the second round of PCR.

[0054]

[0055] Primer F2 / R2 is SEQ ID No 9-10.

[0056] PCR program: pre-denaturation at 95°C for 3 min, (98°C for 20 s, 60°C for 15 s, 72°C for 15 s) for 30 cycles, 72°C for 3 min.

[0057] (4) Agarose gel electrophoresis, gel excision, purification, and next generation sequencing

[0058] Run agarose gel electrophoresis at 100V for 45min, and cut out the target band of about 500bp under blue light. Then use MN DNA purification and recovery kit for gel purification, and operate according to the instructions of the kit. The purified product is sequenced by Miseq2*300 to obtain the initial sequencing data. The sequencing primer sequences used in sequencing are SEQ ID No 11-12.

[0059] Example 3: Sequencing data processing

[0060] The initial sequencing data (FASTQ file) obtained in Example 2 was subjected to downstream analysis using MiXCR (version 3.0.7) software. The analysis with the same V / J gene and CDR3 nucleic acid sequence was defined as the same antibody molecule (clone), and only IgG antibody molecules were retained. The analysis parameters are as follows:

[0061] Align:mixcr align--library my_library-t 8-r align_log.txt R1 R2alignments.vdjca-s hs

[0062] Assemble:mixcr assemble-r assemble_log.txt-OseparateByV=true-OseparateByJ=true-OseparateByC=true alignments.vdjca clones.clna

[0063] Export clones:mixcr exportClones–c IGH clones.clna clones.txt

[0064] Export Alignments:mixcr exportAlignments-f-readIds-cloneId-vHit-vAlignment-jHit-jAlignment-cHit-cAlignment-nFeature FR1-nFeature CDR1-nFeature FR2-nFeature CDR2-nFeature FR3-nFeature CDR3-nFeature FR4-aaFeatureFR1-aaFeature CDR1-aaFeature FR2-aaFeature CDR2-aaFeature FR3-aaFeature CDR3-aaFeature FR4-defaultAnchorPointsclones.clna alignments.txt

[0065] In this embodiment, the antibody repertoire data used for sequencing data processing can also be obtained from network databases such as the IMGT (International Immunogenetics Database) database.

[0066] Example 4: Repertoire-level feature extraction

[0067] For the antibody repertoire data in Example 2, repertoire level features were extracted, and the repertoire level features are as follows:

[0068] ① The usage frequency of the three types of gene segments, V, D, and J, and the combined usage frequency of the VJ gene segments;

[0069] ② The frequency of occurrence of CDR3 amino acid fragments of different lengths;

[0070] ③ α Diversity index and α Evenness indicator calculation is as follows:

[0071]

[0072] In the above formula, f i is the proportion of the i-th antibody molecule, and N is the number of antibody types. The value range of a is [0,10], with a step size of 0.2. 0 The value of E is fixed to 1, so it is not calculated here. 0 E, but introduce ∞ E, its value is ln(1 / p), where p is the proportion of the most abundant antibody.

[0073] ④ Diversity index: D50, vertex gini index and cluster gini index; D50 is defined as the proportion of antibodies that cover half of the reads in the library. The gini index is defined as follows:

[0074]

[0075] For the vertex gini index, all antibodies are sorted in descending order according to the number of reads corresponding to them. i and N are the cumulative probability of reads of the first i antibodies and the total number of antibodies in the library, respectively. The cluster gini index calculation is similar to the vertex gini index calculation, but the antibody is replaced by the antibody cluster. Taking each antibody as a vertex, if the edit distance between the nucleic acid sequences of two antibodies is 1, they are connected to each other. For this antibody network, a group of antibodies that are connected to each other is defined as a cluster.

[0076] ⑤k-mer feature: The antibody amino acid sequence is divided into small segments consisting of k amino acids in an overlapping manner, and the value of k is 3-6. For the antibody repertoire of each sample, only the 1000 k-mers with the highest frequency are taken, and the antibody repertoire is characterized by their frequency.

[0077] ⑥ Antibody cluster characteristics: Antibody molecules of equal length with the same V and J gene segments and a similarity greater than 85% are defined as an antibody cluster. Antibody clusters that appear in at least 20 samples are selected as the final selected antibody clusters, and their proportion in each sample is used as the characteristic.

[0078] Example 5: Sequence-level feature extraction

[0079] For the antibody repertoire data in Example 2, sequence-level features were extracted, and the sequence-level features are as follows:

[0080] ① Various physical and chemical characteristics of the CDR3 amino acid sequence, such as Mole%_Tiny (the proportion of alanine A, cysteine ​​C, glycine G, serine S and threonine T in the amino acid sequence); Mole%_Aliphatic (the proportion of alanine A, isoleucine I, leucine L and valine V in the amino acid sequence); Mole%_Aromatic (the proportion of phenylalanine F, histidine H, tryptophan W and tyrosine Y in the amino acid sequence), Mole%_NonPolar (the proportion of alanine A, cysteine ​​C, phenylalanine F, glycine G, isoleucine I, leucine L , the proportion of methionine M, proline P, valine V, tryptophan W and tyrosine Y in the amino acid sequence), Mole%_Basic (the proportion of histidine H, lysine K and arginine R in the amino acid sequence), F1 (amino acid hydrophobicity index), F6 (protein charge), KF8 (normalized frequency of amino acid physicochemical properties in the α-helical region), F5 (protein sequence flexibility), KF1 (amino acid helix or bend preference), KF2 (amino acid side chain size index) and KF3 (amino acid extended structure preference). These physicochemical characteristics were calculated based on the R language toolkit Peptides.

[0081] ②Sequence frequency: the ratio of the number of reads corresponding to the antibody molecule to the total number of reads;

[0082] ③Read complexity, calculated as follows:

[0083]

[0084] In the above formula, k represents the number of reads with completely identical antibody molecule sequences, n i represents the number of the i-th read, and N represents the total number of reads.

[0085] Example 6: Building an initial prediction model

[0086] The disease diagnosis model involved in the present invention consists of two sub-models (such as repertoire-level-based model, RLM) and sequence-level model (sequence-level-based model, SLM) Figure 1As shown). The RLM model is a three-layer fully connected deep neural network model. The input layer is a one-dimensional feature vector composed of the library-level features extracted in Example 4; the second hidden layer consists of 256 nodes, and its output introduces the rectified linear unit (Rectified Linear Unit, ReLU) activation function and 30% of the node discard regularization processing; the third hidden layer consists of 64 nodes, and its output introduces the rectified linear unit (Rectified Linear Unit, ReLU) activation function and 30% of the node discard regularization processing; the RLM model parameter optimization adopts the AdamOptimizer optimizer with a learning rate of 0.01, and the training epoch and batch size are both set to 100; the activation function of the resulting output layer is Softmax.

[0087] The SLM model is a three-layer convolutional neural network model. The input layer is a 160×160 two-dimensional matrix composed of the sequence level features extracted in Example 5, each row represents an antibody molecule, and each column represents a sequence level feature. The first convolution layer consists of 512 1×160 filters, the stride is set to 160, the padding parameter is set to 0, the activation function is ReLU, and then a 1×2 average pooling layer is connected; the second convolution layer consists of 256 3×1 filters, the stride and padding parameters are set to 1, the activation function is ReLU, and then a 1×2 average pooling layer is connected; the third convolution layer consists of 128 3×1 filters, the stride and padding parameters are set to 1, the activation function is ReLU, and then a 1×2 average pooling layer is connected; the third convolution layer is connected to three fully connected layers based on the ReLU activation function, and the number of nodes is 2560, 1000 and 64 respectively. The activation function of the output layer is Sigmoid. The AdamOptimizer optimizer with a learning rate of 0.01 is used for SLM model parameter optimization. The training epoch and batch size are 250 and 50 respectively.

[0088] Example 7: Repertoire-level feature screening

[0089] The frequency of V gene segment usage, D gene segment usage, J gene segment usage, VJ gene segment combined usage, and the frequency of occurrence of CDR3 amino acid segments of different lengths. α Diversity indicators, αEvenness index, diversity index, k-mer feature and antibody cluster feature, a total of 10 types of features, the initial RLM model in Example 6 was trained using the training set in Example 1, and the importance of the subsets under all possible combinations of the 10 features was evaluated based on the correct judgment rate of 10 cross-tests (10-Fold Cross-validation, 10-CV) of the training set. V gene segment usage frequency, VJ gene segment joint usage frequency, α The combination of the Evenness index and the diversity index, a total of four types of features, has the highest training accuracy and is defined as the retained library-level features.

[0090] Example 8: Sequence-level feature screening

[0091] For the 160 sequence-level features in Example 3: 158 physicochemical properties, sequence frequency and read complexity of the CDR3 amino acid sequence, the initial SLM model in Example 6 was trained using the training set in Example 1, and the importance of each feature was evaluated based on the 10-CV correct rate of the training set, and then the features were arranged in descending order of importance, and introduced one by one. If the introduction of a feature does not increase the accuracy of model training, the feature is removed and the next feature is introduced until all features are traversed. Finally, four sequence-level features, read complexity, sequence frequency, KF8 and F5, are retained.

[0092] The four repertoire-level features retained after screening in Example 7 and the four sequence-level features retained after screening in Example 8 are respectively input into the input layer of the initial RLM model constructed in Example 6 and the input layer of the initial SLM model, thereby obtaining optimized RLM models and SLM models, so as to subsequently perform performance evaluation using the test set in Example 1.

[0093] The test set is tested using the constructed optimized prediction model, and the integrated probability of the optimized prediction model is finally output. The integrated probability is the average of the output probability of the optimized RLM model and the output probability of the optimized SLM model. If the integrated probability is greater than 0.5, the test sample is a pathogen-infected sample, otherwise it is a healthy sample.

[0094] Example 9: Prediction Performance Evaluation

[0095] To verify the predictive effect of the prediction model optimized in Example 8, a sample containing 276 healthy control samples and 326 pathogen infection samples (including HBV infection, influenza virus infection, meningococcal infection, HCV infection and unknown pathogen infection) was collected. All samples were randomly divided into a training set (482 samples) and a test set (120 samples). Table 1 lists the new model and three reference models: support vector machine model (SVM), decision tree model (DT) and random forest model (RF). The new model independent test accuracy (ACC), Matthews correlation coefficient (MCC), F1 and AUC scores are the highest among all models, and each score is higher than the scores of the reported diagnostic model based on antibody repertoire.

[0096] Table 1. Independent test accuracy of each model

[0097] Evaluation indicators SVM DT RF Model of the present invention ACC 0.8750 0.8417 0.9000 0.9500 SN 0.9077 0.9077 0.9231 0.9538 SP 0.8364 0.7636 0.8727 0.9455 MCC 0.7481 0.6828 0.7985 0.8993 F1 0.8872 0.8613 0.9091 0.9538 AUC 0.9429 0.9178 0.9726 0.9883

[0098] The definitions of each evaluation index are as follows:

[0099]

[0100] In the above formula, TP, TN, FP and FN are true positive, true negative, false positive and false negative respectively. SN and SP are positive judgment rate (infection judgment rate) and negative judgment rate (healthy judgment rate) respectively. The higher the evaluation index score in the above formula, the better the model prediction performance.

[0101] sequence:

[0102] SEQ ID No 1-8

[0103] ctacacgacgctcttccgatctactgnnnnnnnnnnnnncaggtgcagctggtggagtctgg

[0104] ctacacgacgctcttccgatctactgnnnnnnnnnnnnncaggtccagctggtgcagtctgg

[0105] ctacacgacgctcttccgatctactgnnnnnnnnnnnnncaggtcaccttgaggggag

[0106] ctacacgacgctcttccgatctactgnnnnnnnnnnnnngaggtgcagctggtggagtcc

[0107] ctacacgacgctcttccgatctactgnnnnnnnnnnnnngaggtgcagctggtggagtct

[0108] ctacacgacgctcttccgatctactgnnnnnnnnnnnnncaggtgcagctacagcagtggg

[0109] ctacacgacgctcttccgatctactgnnnnnnnnnnnnncaggtgcagctgcaggagtcgg

[0110] SEQ ID No 9-10

[0111] ctacacgacgctcttccgatct

[0112] cagacgtgtgctcttccgatct

[0113] SEQ ID No 11-12

[0114] aatgatacggcgaccaccgagatctacac-Index-acactctttccctacacgacgctcttccgatct

[0115] caagcagaagacggcatacgagat-Index-gtgactggagttcagacgtgtgctcttccgatct <110> Guangdong Provincial People's Hospital <120> A non-invasive diagnosis method for infectious diseases based on deep learning <160> 12 <170> Electronic sequence table creation and verification tool V1.0 <210> 1 <211> 61 <212> DNA <213> Artificial sequence <400> 1 ctacacgacg ctcttccgat ctactgnnnn nnnnnnnnca ggtgcagctg gtggagtctg 60 g 61 <210> 2 <211> 61 <212> DNA <213> Artificial sequence <400> 2 ctacacgacg ctcttccgat ctactgnnnn nnnnnnnnca ggtccagctg gtgcagtctg 60 g 61 <210> 3 <211> 56 <212> DNA <213> Artificial sequence <400> 3 ctacacgacg ctcttccgat ctactgnnnn nnnnnnnnca ggtcaccttg aggggag 56 <210> 4 <211> 59 <212> DNA <213> Artificial sequence <400> 4 ctacacgacg ctcttccgat ctactgnnnn nnnnnnnnga ggtgcagctg gtggagtcc 59 <210> 5 <211> 59 <212> DNA <213> Artificial sequence <400> 5 ctacacgacg ctcttccgat ctactgnnnn nnnnnnnnga ggtgcagctg gtggagtct 59 <210> 6 <211> 60 <212> DNA <213> Artificial sequence <400> 6 ctacacgacg ctcttccgat ctactgnnnn nnnnnnnnca ggtgcagcta cagcagtggg 60 <210> 7 <211> 60 <212> DNA <213> Artificial sequence <400> 7 ctacacgacg ctcttccgat ctactgnnnn nnnnnnnnca ggtgcagctg caggagtcgg 60 <210> 8 <211> twenty two <212> DNA <213> Artificial sequence <400> 8 ctacacgacg ctcttccgat ct 22 <210> 9 <211> twenty two <212> DNA <213> Artificial sequence <400> 9 cagacgtgtg ctcttccgat ct 22 <210> 10 <211> 58 <212> DNA <213> Artificial sequence <223> The Index sequence is inserted between the 29th base C and the 30th base A. The Index sequence varies depending on the sequencing platform. The number of bases here counts the number of bases that are not included in the Index sequence. <400> 10 aatgatacgg cgaccaccga gatctacaca cactctttcc ctacacgacg ctcttccgat 60 ct 62 <210> 11 <211> 58 <212> DNA <213> Artificial sequence <223> The Index sequence is inserted between the 24th base T and the 25th base G. The Index sequence varies depending on the sequencing platform. The number of bases here counts the number of bases that are not included in the Index sequence. <400> 11 caagcagaag acggcatacg agatgtgact ggagttcaga cgtgtgctct tccgatct 58

Claims

1. A non-invasive diagnosis method for infectious diseases based on deep learning, characterized in that: The method shown comprises the following steps: (1) Collect a certain number of pathogen-infected samples and a certain number of healthy control samples, and randomly assign all the collected pathogen-infected samples and healthy control samples into training sets and test sets according to a certain ratio; (2) obtaining antibody repertoire data of the sample in step (1); (3) for the antibody repertoire data, extracting repertoire-level features and sequence-level features respectively, wherein the extracted repertoire-level features include at least: 1) V gene segment usage frequency, 2) VJ gene segment joint usage frequency, 3) α Evenness index, 4) diversity index; the extracted sequence-level features include: 1) at least two physicochemical characteristics of CDR3 amino acid sequence: KF8 and F5, 2) sequence frequency, 3) read complexity; (4) using the extracted repertoire-level features and the extracted sequence-level features to construct an initial repertoire-level model and an initial sequence-level model, respectively; The initial repertoire level model is a three-layer fully connected deep neural network model, including an input layer, a second hidden layer, a third hidden layer and a result output layer; the input layer is a one-dimensional feature vector composed of all extracted repertoire level features; the second and third hidden layers contain 256 and 64 nodes respectively, and their outputs are both introduced with ReLU activation function and 30% node drop regularization processing; the parameter optimization of the repertoire level model adopts AdamOptimizer optimizer with a learning rate of 0.01, and the training epoch and batch size are both set to 100; the activation function of the result output layer is Softmax; The initial sequence-level model includes an input layer, a first convolutional layer, a second convolutional layer, a third convolutional layer, a fully connected layer, an average pooling layer, and a result output layer; the input layer is a 160×160 two-dimensional matrix composed of all extracted sequence-level features, each row represents an antibody molecule, and each column represents a sequence-level feature; the first convolutional layer consists of 512 1×160 filters, the step size is set to 160, the padding parameter is set to 0, the activation function is ReLU, and then a 1×2 average pooling layer is connected; the second convolutional layer consists of 256 3×1 filters, the step size and padding parameters are The activation function is set to 1, and a 1×2 average pooling layer is connected afterwards. The third convolutional layer consists of 128 3×1 filters, and the step size and padding parameters are both set to 1. The activation function is ReLU, and a 1×2 average pooling layer is connected afterwards. The third convolutional layer is connected to three fully connected layers based on the ReLU activation function, with 2560, 1000 and 64 nodes respectively. The activation function of the output layer is Sigmoid. The AdamOptimizer optimizer with a learning rate of 0.01 is used for sequence-level model parameter optimization, and the training epoch and batch size are 250 and 50 respectively. (5) Based on the constructed initial repertoire-level feature model and the correct judgment rate of 10 cross-tests of the training set in step (1), the importance of the subsets under all possible combinations of all extracted repertoire-level features is evaluated. Based on the principle of the highest training accuracy, the repertoire-level features to be retained are screened out. The repertoire-level features retained after screening include the usage frequency of V gene segments, the combined usage frequency of VJ gene segments, the α Evenness index, and the diversity index; Based on the constructed initial sequence-level feature model, the importance of each sequence-level feature is evaluated based on the 10 cross-test correctness ratios of the training set in step (1), and then each sequence-level feature is sorted in descending order of importance. The sequence-level features are introduced one by one. If the introduced sequence-level feature cannot increase the model training accuracy, the sequence-level feature is removed and the next sequence-level feature is introduced until all the extracted sequence-level features are traversed, and finally the sequence-level features to be retained are screened out. The sequence-level features retained after screening include read complexity, sequence frequency, KF8 and F5; (6) obtaining an optimized repertoire level model by inputting the screened repertoire level features into the initial repertoire level feature model, and obtaining an optimized sequence level model by inputting the screened sequence level features into the initial sequence level feature model; (7) Using the optimized repertoire level model and the optimized sequence level model to perform performance evaluation on the test set in step (1); during the evaluation process, an integrated probability is output, which is the average of the output probability of the optimized repertoire level model and the output probability of the optimized sequence level model. If the integrated probability is greater than 0.5, the test sample is a pathogen-infected sample, otherwise it is a healthy control sample.

2. The method according to claim 1, characterized in that The number of pathogen-infected samples and healthy control samples collected was no less than 100.

3. The method according to claim 2, characterized in that The ratio of the collected pathogen-infected samples to healthy control samples was 2:1 to 5:

1.

4. The method according to claim 3, characterized in that The ratio of the collected pathogen-infected samples to healthy control samples was 2.5:1 to 4.5:

1.

5. A non-invasive diagnosis system for infectious diseases based on deep learning, characterized in that: The system comprises: (1) a training set and a test set selection module; (2) an antibody repertoire data acquisition module; (3) a repertoire level feature extraction module; (4) a sequence level feature extraction module; (5) an initial prediction model construction module; (6) a repertoire level feature screening module; (7) a sequence level feature screening module; (8) a prediction model performance evaluation module, wherein: The training set and test set selection module collects a certain number of pathogen-infected samples and a certain number of healthy control samples, and randomly allocates all the collected pathogen-infected samples and healthy control samples as training sets and test sets according to a certain ratio; The antibody repertoire data acquisition module acquires the antibody repertoire data of the samples collected by the training set and test set selection modules; The repertoire level feature extraction module and the sequence level feature extraction module respectively extract the repertoire level features and sequence level features of the antibody repertoire data acquired by the antibody repertoire data acquisition module; the extracted repertoire level features at least include: 1) V gene segment usage frequency, 2) VJ gene segment joint usage frequency, 3) αEvenness index, 4) diversity index; the extracted sequence level features include: 1) at least two physical and chemical characteristics of CDR3 amino acid sequence: KF8 and F5, 2) sequence frequency, 3) read complexity; The prediction model building module builds an initial repertoire-level feature model and an initial sequence-level feature model by inputting all extracted repertoire-level features and sequence-level features; The initial repertoire level model is a three-layer fully connected deep neural network model, including an input layer, a second hidden layer, a third hidden layer and a result output layer; the input layer is a one-dimensional feature vector composed of all extracted repertoire level features; the second and third hidden layers contain 256 and 64 nodes respectively, and their outputs are both introduced with ReLU activation function and 30% node drop regularization processing; the parameter optimization of the repertoire level model adopts AdamOptimizer optimizer with a learning rate of 0.01, and the training epoch and batch size are both set to 100; the activation function of the result output layer is Softmax; The initial sequence-level model includes an input layer, a first convolutional layer, a second convolutional layer, a third convolutional layer, a fully connected layer, an average pooling layer, and a result output layer; the input layer is a 160×160 two-dimensional matrix composed of all extracted sequence-level features, each row represents an antibody molecule, and each column represents a sequence-level feature; the first convolutional layer consists of 512 1×160 filters, the step size is set to 160, the padding parameter is set to 0, the activation function is ReLU, and then a 1×2 average pooling layer is connected; the second convolutional layer consists of 256 3×1 filters, the step size and padding parameters are The activation function is set to 1, and a 1×2 average pooling layer is connected afterwards. The third convolutional layer consists of 128 3×1 filters, and the step size and padding parameters are both set to 1. The activation function is ReLU, and a 1×2 average pooling layer is connected afterwards. The third convolutional layer is connected to three fully connected layers based on the ReLU activation function, with 2560, 1000 and 64 nodes respectively. The activation function of the output layer is Sigmoid. The AdamOptimizer optimizer with a learning rate of 0.01 is used for sequence-level model parameter optimization, and the training epoch and batch size are 250 and 50 respectively. The repertoire-level feature screening module and the sequence-level feature screening module are based on the constructed initial repertoire-level feature model, and the importance of the subsets of all possible combinations of all extracted repertoire features is evaluated based on the correct judgment rate of 10 cross tests of the training set. The repertoire-level features that need to be retained are screened out based on the principle of the highest training accuracy. The repertoire-level features retained after screening include the usage frequency of V gene segments, the joint usage frequency of VJ gene segments, the αEvenness index, and the diversity index; the initial sequence-level feature model constructed is also based on the correct judgment rate of 10 cross tests of the training set to evaluate the importance of each sequence-level feature, and then each sequence-level feature is sorted in descending order of importance, and sequence-level features are introduced one by one. If the introduced sequence-level feature cannot increase the model training accuracy, the sequence-level feature is removed, and the next sequence-level feature is introduced until all the extracted sequence-level features are traversed, and finally the sequence-level features that need to be retained are screened out. The sequence-level features retained after screening include read complexity, sequence frequency, KF8, and F5; Then, the optimized repertoire level model is obtained by inputting the screened repertoire level features into the initial repertoire level feature model, and the optimized sequence level model is obtained by inputting the screened sequence level features into the initial sequence level feature model; the prediction model performance evaluation module uses the optimized repertoire level model and the optimized sequence level model to perform performance evaluation on the test set respectively. During the evaluation process, an integrated probability will be output, which is the average of the output probability of the optimized repertoire level model and the output probability of the optimized sequence level model. If the integrated probability is greater than 0.5, the test sample is a pathogen-infected sample, otherwise it is a healthy control sample.

6. The non-invasive diagnosis system for infectious diseases according to claim 5, characterized in that: The number of pathogen-infected samples and healthy control samples collected was no less than 100.

7. The non-invasive diagnosis system for infectious diseases according to claim 6, characterized in that: The ratio of the collected pathogen-infected samples to healthy control samples was 2:1 to 5:

1.

8. The non-invasive diagnosis system for infectious diseases according to claim 7, characterized in that: The ratio of the collected pathogen-infected samples to healthy control samples was 2.5:1 to 4.5:1.

Citation Information

Patent Citations

  • Dual specificity antibodies and methods of making and using

    AU2007202323A1

  • Immunosignaturing: path to early diagnosis and health monitoring

    CN104781668A