Peptide chips and their application in the preparation of AIDS diagnostic products

By developing polypeptide chips and machine learning prediction models containing characteristic polypeptides, the cost and complexity of existing AIDS detection methods are solved, and low-cost and efficient AIDS detection is achieved.

CN115963260BActive Publication Date: 2025-05-23ZHUHAI CARBON CLOUD DIAGNOSIS TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210737190.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-06-27
Publication Date
2025-05-23
Estimated Expiration
2042-06-27

AI Technical Summary

Technical Problem

The existing AIDS testing methods are harsh, costly, and expensive data analysis and testing costs, making it difficult to achieve low-cost AIDS testing.

Method used

Develop a polypeptide chip that contains characteristic polypeptides. Through polypeptide chip detection, imaging, fluorescence intensity detection, combined with machine learning training methods, a prediction model is built to achieve low-cost and efficient detection of AIDS.

Benefits of technology

It reduces the cost and complexity of AIDS testing, improves the accuracy and reliability of testing, simplifies the experimental process, and reduces the impact of the experimental environment on the results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115963260B_ABST
    Figure CN115963260B_ABST
Patent Text Reader

Abstract

The present invention provides a polypeptide chip and its application in the preparation of AIDS diagnostic products. The polypeptide chip includes characteristic polypeptides, and the characteristic polypeptides include polypeptides shown in SEQ ID NOs: 1 to 123; or polypeptides shown in SEQ ID NOs: 1 to 129; or polypeptides shown in SEQ ID NOs: 1 to 163; or polypeptides shown in SEQ ID NOs: 1 to 172; or polypeptides shown in SEQ ID NOs: 1 to 207. The problem of harsh AIDS detection conditions in the prior art can be solved, and the method is applicable to the field of AIDS diagnosis.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of AIDS diagnosis, and in particular to a polypeptide chip and its application in the preparation of AIDS diagnosis products. Background Art

[0002] AIDS (HIV) is an infectious disease caused by human immunodeficiency virus infection, which attacks CD4T cells to make the human body lose its immune function. AIDS has a long incubation period and a high mortality rate, which is extremely harmful. The average incubation period of HIV in the human body is 8-9 years, during which there may be no symptoms. There is no effective vaccine to prevent AIDS or effective medicine to cure HIV infection. HIV has very poor survival ability outside the body and is mainly transmitted through body fluids. HIV testing is the key to preventing and controlling AIDS.

[0003] HIV testing mainly includes anti-HIV antibody testing, virus culture, nucleic acid testing and antigen testing. Generally, the initial screening test uses HIV antibody detection methods, such as enzyme-linked immunosorbent assay (ELISA), chemiluminescence method and colloidal gold rapid method. Enzyme-linked immunosorbent assay (ELISA) is currently the most commonly used method for HIV testing. Enzyme-linked immunosorbent assay, referred to as enzyme-linked immunosorbent assay, or ELISA method. The basis of ELISA is the solid phase of antigen or antibody and the enzyme labeling of antigen or antibody, the enzyme activity in ELISA, and the amount of product is directly related to the amount of the test substance in the specimen, and qualitative or quantitative analysis is performed based on the depth of color. ELISA includes several steps such as sample addition, incubation, washing, color development, and microplate reader reading. ELISA is based on enzyme activity, and there are many influencing factors in the determination. Sample collection and preservation, reagent storage and use, and experimental temperature requirements are high. In addition, laboratory biosafety is also very important, and the operating technology requirements for inspection personnel are also high.

[0004] However, the existing peptide chips contain a large number of peptides, and the manufacturing cost, testing cost and data analysis cost are high, making it difficult to conduct low-cost testing for AIDS. In addition, a large amount of data is irrelevant to AIDS testing, which affects the accuracy of the test results. Summary of the invention

[0005] The main purpose of the present invention is to provide a polypeptide chip and its application in the preparation of AIDS diagnostic products to solve the problem of harsh AIDS detection conditions in the prior art.

[0006] In order to achieve the above-mentioned purpose, according to the first aspect of the present invention, a polypeptide chip is provided, which includes characteristic polypeptides, and the characteristic polypeptides include polypeptides shown in SEQ ID NOs: 1 to 123; or polypeptides shown in SEQ ID NOs: 1 to 129; or polypeptides shown in SEQ ID NOs: 1 to 163; or polypeptides shown in SEQ ID NOs: 1 to 172; or polypeptides shown in SEQ ID NOs: 1 to 207.

[0007] In order to achieve the above-mentioned purpose, according to the second aspect of the present invention, a method for screening characteristic polypeptides is provided, which screening method comprises: obtaining the AIDS status and biological samples of the sample population, using the above-mentioned polypeptide chip to detect, image and detect the fluorescence intensity of the biological samples, and obtaining the fluorescence signal intensity information of the polypeptides on the polypeptide chip, wherein the sample population comprises a healthy population and an AIDS patient population; based on the fluorescence signal intensity information, screening for significantly different polypeptides with signal differences between AIDS patients and healthy people by statistical methods; using the fluorescence signal intensity information of the significantly different polypeptides of the sample population and the AIDS status, constructing a prediction model by a machine learning training method; the polypeptides appearing in the prediction model are the characteristic polypeptides.

[0008] Further, the biological sample comprises serum or plasma; preferably, the statistical method comprises variance threshold screening.

[0009] Furthermore, the method of machine learning training includes: dividing the sample population into a training set and a test set, establishing a mapping relationship between the fluorescence signal intensity information of the significantly different polypeptides in the training set or the test set and the AIDS status, and obtaining a training data set or a test data set respectively; using a machine learning method to construct a model for the training data set, and performing a five-fold cross-validation, and by adjusting the penalty factor and the variance threshold, constructing multiple preliminary models under different penalty factors and variance threshold conditions; using the test data set, respectively verifying the multiple preliminary models, and calculating the accuracy and recall rate of each preliminary model; the model whose accuracy and recall rate are both greater than the screening threshold is the prediction model; preferably, the machine learning method includes any one or more of the following: a logistic regression model, a support vector machine model or a naive Bayes model; preferably, the screening threshold for accuracy is 0.7, and the screening threshold for recall is 0.85.

[0010] Furthermore, when constructing multiple preliminary models, the accuracy of each preliminary model is calculated. Under the condition of ensuring the accuracy, the recall rate of each preliminary model is gradually improved by adjusting the penalty factor and variance threshold.

[0011] In order to achieve the above-mentioned purpose, according to the third aspect of the present invention, a method for constructing an AIDS prediction model is provided, and the construction method comprises: a) obtaining the AIDS status and biological samples of a sample population, using the above-mentioned polypeptide chip to detect, image, and detect the fluorescence signal intensity of the biological samples, extracting the fluorescence signal intensity information of the characteristic polypeptides, and establishing a data set of the biological samples, wherein the data set comprises a mapping relationship between the AIDS status of the sample population and the fluorescence signal intensity information of the characteristic polypeptides in the biological samples; b) constructing a prediction model by using the data set through a machine learning training method; the sample population comprises a healthy population and an AIDS patient group.

[0012] Furthermore, before using machine learning training, the data set is divided into a training set and a test set, the classification model is constructed using the training set, and the model effect is verified using the test set; preferably, the machine learning training method includes any one or more of the following: a logistic regression model, a support vector machine model, or a naive Bayes model; preferably, the characteristic polypeptide includes a polypeptide shown in SEQ ID NO: 1 to 123, or includes a polypeptide shown in SEQ ID NO: 1 to 129; or includes a polypeptide shown in SEQ ID NO: 1 to 163; or includes a polypeptide shown in SEQ ID NO: 1 to 172; or includes a polypeptide shown in SEQ ID NO: 1 to 207; preferably, the biological sample includes serum or plasma.

[0013] To achieve the above-mentioned purpose, according to a fourth aspect of the present invention, an electronic device for AIDS diagnosis is provided, the electronic device comprising a diagnostic model, the diagnostic model is constructed using the above-mentioned construction method, the diagnostic model uses the detection result of the polypeptide chip on the test sample as the model input, and outputs the diagnostic result, the detection result includes the fluorescence signal intensity information of the characteristic polypeptide extracted by detecting the test sample using the polypeptide chip; preferably, the biological sample includes serum or plasma; preferably, the characteristic polypeptide includes the polypeptide shown in SEQ ID NO: 1 to 123, or includes the polypeptide shown in SEQ ID NO: 1 to 129; or includes the polypeptide shown in SEQ ID NO: 1 to 163; or includes the polypeptide shown in SEQ ID NO: 1 to 172; or includes the polypeptide shown in SEQ ID NO: 1 to 207.

[0014] In order to achieve the above-mentioned purpose, according to the fifth aspect of the present invention, a computer-readable storage medium is provided, which includes a stored program, wherein when the program is executed, the device where the storage medium is located is controlled to execute the above-mentioned screening method or construction method.

[0015] In order to achieve the above-mentioned purpose, according to a sixth aspect of the present invention, a processor is provided, the processor being used to run a program, wherein the above-mentioned screening method or construction method is executed when the program is run.

[0016] In order to achieve the above-mentioned purpose, according to the seventh aspect of the present invention, there is provided an application of a characteristic polypeptide in the preparation of a product for diagnosing and / or treating AIDS, wherein the characteristic polypeptide includes a polypeptide shown in SEQ ID NOs: 1 to 123; or a polypeptide shown in SEQ ID NOs: 1 to 129; or a polypeptide shown in SEQ ID NOs: 1 to 163; or a polypeptide shown in SEQ ID NOs: 1 to 172; or a polypeptide shown in SEQ ID NOs: 1 to 207.

[0017] Furthermore, the application includes application in the preparation of AIDS diagnosis products, the preparation of AIDS treatment effect prediction products, or the preparation of AIDS targeted drugs; preferably, the above-mentioned AIDS diagnosis products include polypeptide chips.

[0018] By applying the technical solution of the present invention, the above-mentioned polypeptide chip utilizes characteristic polypeptides, and the number of polypeptides is greatly reduced compared with the existing polypeptide chips. The polypeptide chip can accurately detect characteristic polypeptides related to AIDS, and the detection cost is low. Compared with the HIV antibody detection method in the prior art, the requirements for experimental conditions and detection personnel are lower. The above-mentioned polypeptide chip can greatly expand the application scenarios of AIDS detection and realize the effective detection of HIV infection in samples using polypeptide detection technology. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] The drawings constituting a part of the present application are used to provide a further understanding of the present invention. The exemplary embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute an improper limitation of the present invention. In the drawings:

[0020] Figure 1 A cluster analysis diagram of sample data according to Example 2 of the present invention is shown, wherein the gray dots represent patients and the white dots represent normal persons.

[0021] Figure 2 A schematic diagram of a confusion matrix for training set fitting according to Example 2 of the present invention is shown; wherein 0 represents a healthy person, 1 represents a patient, TN represents a true negative, FN represents a false negative, TP represents a true positive, and FP represents a false positive.

[0022] Figure 3 A schematic diagram of the ROC curve of the training set model according to Example 2 of the present invention is shown.

[0023] Figure 4 A schematic diagram of the test set fitting confusion matrix according to Example 2 of the present invention is shown; wherein 0 represents a healthy person, 1 represents a patient, TN represents a true negative, FN represents a false negative, TP represents a true positive, and FP represents a false positive.

[0024] Figure 5A volcano plot is shown based on the p-values ​​obtained by T-test of all 126,501 polypeptides according to Example 2 of the present invention, wherein the discontinuous white dots are polypeptides of SEQ ID NOs: 1 to 123. DETAILED DESCRIPTION

[0025] It should be noted that, in the absence of conflict, the embodiments and features in the embodiments of the present application can be combined with each other. The present invention will be described in detail below in conjunction with the embodiments.

[0026] As mentioned in the background technology, there are many influencing factors in the existing AIDS detection methods, and the requirements for sample collection and preservation, reagent storage and use, experimental temperature, biosafety, and operating skills of detection personnel are all high. Therefore, in this application, the inventors tried to explore what kind of polypeptide combination can be used to prepare products for diagnosing AIDS, and obtained characteristic polypeptides that can be used to prepare products for diagnosing AIDS by analyzing and screening from a large amount of experimental data. Therefore, a series of protection schemes for this application are proposed.

[0027] In a first typical embodiment of the present application, a polypeptide chip is provided, which includes characteristic polypeptides, and the characteristic polypeptides include polypeptides shown in SEQ ID NOs: 1 to 123; or polypeptides shown in SEQ ID NOs: 1 to 129; or polypeptides shown in SEQ ID NOs: 1 to 163; or polypeptides shown in SEQ ID NOs: 1 to 172; or polypeptides shown in SEQ ID NOs: 1 to 207.

[0028] Peptide chip is a detection technology that fixes a series of peptide fragments to a carrier to detect paired unknown proteins. This technology can detect serum immune responses caused by viral infection. It has been used for antibody identification and verification, autoimmune disease research, tumor marker research, allergen research, and infectious disease research. Compared with other antibody detection technologies, this technology has high stability and design flexibility. Purely chemically synthesized peptide molecules have protective groups on their surfaces, which greatly extends the shelf life of peptide chips. Peptide chips stored at room temperature for more than two years still have complete biological activity. The design of peptide chips is flexible, high-density, and high-throughput. The probe proteins on the chip can be designed according to actual needs, and tens of thousands of protein-peptide biochemical reactions can be measured simultaneously.

[0029] The polypeptides composed of the above-mentioned characteristic polypeptides are set on this polypeptide chip. Compared with the existing polypeptide chip, the number of polypeptides set is greatly reduced (the polypeptide chip here can be composed of only the above-mentioned polypeptide combination, or it can contain a small amount of other polypeptides on the basis of the above-mentioned polypeptide combination, but the number of polypeptides on the whole is much less than that on the existing polypeptide chip for screening these characteristic polypeptides, such as 5,000 less, 6,000 less, 7,000 less, 8,000 less, 9,000 less, 10,000 less, or even more than 11,000 less), which greatly reduces the immune information unrelated to HIV infection, reduces the cost of data analysis, reduces the preparation and detection cost of polypeptide chips, and at the same time has the advantages of high stability and high throughput of polypeptide chips in the prior art. The combination of samples and polypeptide chips is based on protein-protein interactions, and the fluorescence signal detection process is not sensitive to temperature and has lower requirements for the experimental environment. The chip can be placed at room temperature for a long time without refrigeration, and the polypeptide chip experimental process can be fully automated. Not only is the experimental process simple, but it also reduces the risk of infection of experimenters during operation and minimizes the impact of the environment on the results. By using this polypeptide chip, it is possible to detect whether the subject has AIDS at a low cost and high efficiency.

[0030] In a second typical embodiment of the present application, a method for screening characteristic polypeptides is provided, the screening method comprising: obtaining the AIDS status and biological samples of the sample population, using the above-mentioned polypeptide chip to detect, image, and detect the fluorescence intensity of the biological samples to obtain the fluorescence signal intensity information of the polypeptides on the polypeptide chip, the sample population comprising a healthy population and an AIDS patient population; based on the fluorescence signal intensity information, screening for significantly different polypeptides with signal differences between AIDS patients and healthy people by statistical methods; using the fluorescence signal intensity information of the significantly different polypeptides of the sample population and the AIDS status to construct a prediction model by a machine learning training method; the polypeptides appearing in the prediction model are the characteristic polypeptides.

[0031] In a preferred embodiment, the biological sample comprises serum or plasma; preferably, the statistical method comprises variance threshold screening.

[0032] In a preferred embodiment, the method for machine learning training includes: dividing the sample population into a training set and a test set, establishing a mapping relationship between the fluorescence signal intensity information of significantly different polypeptides in the training set or the test set and the AIDS status, and obtaining a training data set or a test data set respectively; using a machine learning method to construct a model for the training data set and performing five-fold cross-validation, and constructing multiple preliminary models under different penalty factors and variance thresholds by adjusting the penalty factor and the variance threshold; using the test data set to verify the multiple preliminary models respectively, and calculating the accuracy and recall rate of each preliminary model; the model with both the accuracy and the recall rate greater than the screening threshold is the prediction model; preferably, the machine learning method includes any one or more of the following: logistic regression model, support vector machine model or naive Bayes model; preferably, the screening threshold for accuracy is 0.7, and the screening threshold for recall rate is 0.85.

[0033] For the thresholds of accuracy and recall rate, they can be flexibly selected according to the actual situation and requirements. In this application, the screening threshold is preferably that the accuracy is greater than 0.7 and the recall rate is greater than 0.85.

[0034] In a preferred embodiment, when constructing multiple preliminary models, calculate the accuracy of each preliminary model, and under the condition of ensuring the accuracy, gradually increase the recall rate of each preliminary model by adjusting the penalty factor and the variance threshold.

[0035] In the above method, a polypeptide chip is used to detect biological samples of the sample population (including healthy people and AIDS patients), so as to obtain the fluorescence signal intensity information of the polypeptides on the polypeptide chip. Using the data generated by the polypeptide chip, it is known that a large number of polypeptides on the chip can bind to antibodies in biological samples, so the antibody spectrum in the individual serum can be unbiasedly obtained using the immune characteristics of the test samples.

[0036] Based on the obvious differences in the immune systems between HIV-positive individuals after HIV infection and normal uninfected individuals, according to the fluorescence signal intensity information, significantly different polypeptides with differences between AIDS patients and healthy people are screened by statistical methods. These significantly different polypeptides are used for modeling, and corresponding features are identified by machine learning methods. Using the data of the training data set, by continuously training and optimizing the corresponding parameters of the model and the required significantly different polypeptides, a suitable model for predicting the HIV infection status of the sample is constructed. Finally, through another batch of independent test data sets, the HIV infection status is predicted, which can further verify the results of the model. In the above screening method, using a machine learning algorithm to construct a model to judge the result can ensure the accuracy and reliability of the detection result.

[0037] In a third typical embodiment of the present application, a method for constructing an AIDS prediction model is provided, the construction method comprising: a) obtaining the AIDS status of a sample population and a biological sample, using the above-mentioned polypeptide chip to detect, image, and detect the fluorescence signal intensity of the biological sample, extracting the fluorescence signal intensity information of the characteristic polypeptides, and establishing a data set of the biological sample, the data set comprising a mapping relationship between the AIDS status of the sample population and the fluorescence signal intensity information of the characteristic polypeptides in the biological sample; b) constructing a prediction model using the data set through a machine learning training method; the sample population comprises a healthy population and an AIDS patient group.

[0038] In a preferred embodiment, before using machine learning training, the data set is divided into a training set and a test set, the classification model is constructed using the training set, and the model effect is verified using the test set; preferably, the machine learning training method includes any one or more of the following: a logistic regression model, a support vector machine model, or a naive Bayes model; preferably, the characteristic polypeptide includes a polypeptide shown in SEQ ID NO: 1 to 123, or includes a polypeptide shown in SEQ ID NO: 1 to 129; or includes a polypeptide shown in SEQ ID NO: 1 to 163; or includes a polypeptide shown in SEQ ID NO: 1 to 172; or includes a polypeptide shown in SEQ ID NO: 1 to 207; preferably, the biological sample includes serum or plasma.

[0039] In a fourth typical embodiment of the present application, an electronic device for AIDS diagnosis is provided, the electronic device comprising a diagnostic model, the diagnostic model is constructed using the above-mentioned construction method, the diagnostic model uses the detection result of the polypeptide chip on the test sample as the model input, and outputs the diagnostic result, the detection result includes the fluorescence signal intensity information of the characteristic polypeptide extracted by detecting the test sample using the polypeptide chip; preferably, the biological sample includes serum or plasma; preferably, the characteristic polypeptide includes the polypeptide shown in SEQ ID NO: 1 to 123, or includes the polypeptide shown in SEQ ID NO: 1 to 129; or includes the polypeptide shown in SEQ ID NO: 1 to 163; or includes the polypeptide shown in SEQ ID NO: 1 to 172; or includes the polypeptide shown in SEQ ID NO: 1 to 207.

[0040] The above prediction model is a prediction model obtained by training a computer using the corresponding relationship between the fluorescence signal intensity information of the characteristic polypeptides of healthy people and AIDS patients and the disease conditions. There are differential antibodies in the biological samples of healthy people and AIDS patients. The differential antibodies are combined with the characteristic polypeptides on the polypeptide chip to display different fluorescence signal intensity information. The model can be used to process the fluorescence signal intensity information of the characteristic polypeptides of the subject to output whether the subject has AIDS. The above electronic device has a built-in prediction model. When the fluorescence signal intensity information of the characteristic polypeptides of the subject is input into the device, the electronic device can output the diagnosis result.

[0041] The device for detecting the fluorescence signal intensity information of the biological sample is a polypeptide chip detection device, which may include a fluorescence imager (such as Melecular Device Image Xpress Micro-4), a chip centrifuge (such as Labnet C1303T-230V), a plate washer (such as BioTek Instruments, 405TSUVS), an oscillator (such as a 96-well plate orbital oscillator, Thermo scientific, 88880026), and a mixer (such as a constant temperature mixer, Eppendorf Thermomixer C). In actual use, the detection device can be flexibly increased or decreased according to the specific needs of use.

[0042] In a fifth typical embodiment of the present application, a computer-readable storage medium is provided, which includes a stored program, wherein when the program is executed, the device where the storage medium is located is controlled to execute the above-mentioned screening method or construction method.

[0043] In a sixth typical implementation of the present application, a processor is provided, the processor being used to run a program, wherein the above-mentioned screening method or construction method is executed when the program is run.

[0044] It should be noted that, for the above-mentioned method embodiments, for the sake of simplicity, they are all described as a series of action combinations, but those skilled in the art should know that the present invention is not limited by the order of the actions described, because according to the present invention, certain steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should also know that the embodiments described in the specification are all preferred embodiments, and the actions involved are not necessarily required by the present invention.

[0045] It can be known from the description of the above implementation mode that a person skilled in the art can clearly understand that the present application can be implemented by means of software plus hardware devices such as a detection device. Based on such an understanding, the data processing part of the technical solution of the present application can be embodied in the form of a software product, and the computer software product can be stored in a storage medium such as ROM / RAM, a magnetic disk, an optical disk, etc., including a number of instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods of each embodiment of the present application or some parts of the embodiments.

[0046] The present application can be used in many general or special computing system environments or configurations, such as personal computers, server computers, handheld or portable devices, tablet devices, multiprocessor systems, microprocessor-based systems, set-top boxes, programmable consumer electronic devices, network PCs, minicomputers, mainframe computers, distributed computing environments including any of the above systems or devices, etc.

[0047] Obviously, those skilled in the art should understand that some modules or steps of the present application described above can be implemented in a general computing device, they can be concentrated on a single computing device, or distributed on a network composed of multiple computing devices, and optionally, they can be implemented with program codes executable by the computing device, so that they can be stored in a storage device and executed by the computing device, or they can be made into individual integrated circuit modules, or multiple modules or steps therein can be made into a single integrated circuit module for implementation. Thus, the present application is not limited to any specific combination of hardware and software.

[0048] In the seventh typical embodiment of the present application, a characteristic polypeptide is provided for use in the preparation of a product for diagnosing and / or treating AIDS, wherein the characteristic polypeptide includes a polypeptide shown in SEQ ID NOs: 1 to 123; or a polypeptide shown in SEQ ID NOs: 1 to 129; or a polypeptide shown in SEQ ID NOs: 1 to 163; or a polypeptide shown in SEQ ID NOs: 1 to 172; or a polypeptide shown in SEQ ID NOs: 1 to 207.

[0049] The product for diagnosing AIDS prepared using the above-mentioned characteristic polypeptide can be combined with the sample to be tested, and through the immune changes between the polypeptide and the sample to be tested, the individual HIV infection status can be judged, thereby realizing the effective detection of HIV infection in the sample using polypeptide detection technology. The above-mentioned characteristic polypeptide is a characteristic polypeptide obtained by screening from the existing polypeptide chip, and the above-mentioned characteristic polypeptide can be used to detect whether the test subject has AIDS. Compared with the tens of thousands of polypeptides in the existing polypeptide chip, the above-mentioned characteristic polypeptide contains a small number of polypeptides, has a low preparation cost, and can reduce the amount of data and workload generated by the test.

[0050] In a preferred embodiment, the application includes application in the preparation of a diagnostic AIDS product, a product for predicting the effect of AIDS treatment, or a product for preparing AIDS targeted drugs; preferably, the diagnostic AIDS product includes a polypeptide chip.

[0051] The theoretical basis for the ability of peptides to identify HIV infection is that they have the ability to stably bind to HIV antibodies. The increase in the number of CD4 T cells in AIDS patients with good treatment effects can stimulate B cells to produce higher HIV antibody titers, and these antibodies can also bind to peptides and provide stronger signals. Therefore, the above-mentioned characteristic peptides have the application of predicting the treatment effect of AIDS.

[0052] When HIV-infected and normal samples are tested on peptide chips, the essence is that the characteristic peptides will react immune-responsively with HIV-infected samples, thereby showing differences in fluorescence signal intensity between HIV-infected samples and normal samples. These characteristic peptides showing signal differences can therefore be used to assist in the development of vaccines or related drugs.

[0053] The beneficial effects of the present application will be further explained in detail below in conjunction with specific embodiments.

[0054] Example 1 Peptide chip experiment and data generation process

[0055] 1. Sample Preparation

[0056] The sample source of the training set of this application is 219 available blood / serum samples, of which 198 are HIV patient samples and 21 are healthy person samples. The 24 samples of the test set are from 20 HIV positive and 4 HIV negative samples provided by Spectrum Biotechnology. All the above HIV patients and healthy controls were judged by the gold standard ELISA / Western blot experiment (Western blot, WB) for AIDS testing.

[0057] 1) Sample loading

[0058] Serum or plasma samples were diluted 25 times twice with 1% D-mannitol solution in a 96-well deep-well plate to obtain a 625-fold diluted sample plate for use.

[0059] 2) Hydration and assembly of the chip

[0060] Place the chip in a chip hydration device, add ultrapure water to cover the chip, and hydrate on an orbital shaker at 55±5rpm / min for 20 minutes. Then spray the chip surface with isopropanol and place the chip in a centrifuge for centrifugal drying. Assemble the dried chip into an assay cassette according to the experimental design.

[0061] The polypeptide chip used in the examples of the present application is the V13 chip of Health Tell, which has 24 repeated polypeptide arrays, each array having more than 130,000 polypeptides. These 130,000 polypeptides are polypeptide sequences formed by combinations of 5-13 unbiased random amino acids.

[0062] 3) Sample and chip incubation

[0063] The diluted sample was added to the assembled chip at 90 μL / well and incubated on a constant temperature shaker for 1 hour.

[0064] 4) Sample cleaning

[0065] Place the assay cassette in a plate washer for washing.

[0066] 5) Fluorescent secondary antibody incubation

[0067] A 2 nM fluorescent secondary antibody solution was prepared with 0.75% casein solution, added to the assay cassette at a rate of 40 μL / well, and placed on a constant temperature shaker for shaking and incubation for 1 hour.

[0068] 6) Secondary antibody washing

[0069] Same as step 3).

[0070] 7) Imaging

[0071] The chips in the assay cassette were disassembled, cleaned, dried, assembled into an imaging cassette, and placed into the ImageXpress micro4 imager from Molecular Devices for scanning and imaging. Finally, a TIFF image file was obtained for each sample tested, which was the original data.

[0072] 2. Data Extraction

[0073] 1) Use the MIAMI pipeline written by myself to grid the TIFF images generated by each sample (it should be noted that the grid processing software in the prior art

[0074] The TIFF images generated by each sample can be gridded, such as the gridding software that comes with HealthTell), and then the fluorescence intensity values ​​of the features (the features here refer to the differential peptides) can be extracted, and a GPR5 data file and a corner images file can be output. Among them, the GPR5 file contains all the information of a sample and the fluorescence signal intensity information of all features.

[0075] 2) Extract the characteristic fluorescence signal intensity information from the GPR5 data files of all samples to generate the original fluorescence intensity (FG, foreground) data matrix. Then logarithmically transform the data of each sample to obtain the LFG (log-transferred foreground) data matrix, and perform z-score standardization to obtain the NLFG (normalized and log-transferred foreground) data matrix. This step will also generate a sample chip information file, which includes information such as the sample array position and the chip number used.

[0076] 3. Quality Control

[0077] 3.1 Single sample quality control

[0078] 1) Supersaturation analysis

[0079] Quality control indicator: the proportion of features (i.e., polypeptide fragments) whose original fluorescence intensity values ​​(FG) of a single sample exceed the upper limit of detection of the imager.

[0080] Quality control standard: The above ratio is qualified if ≤ 1%.

[0081] What to do when quality control fails: adjust the exposure time and rescan the chip until the ratio is ≤1%.

[0082] 2) Fluorescence signal distribution analysis

[0083] Quality control indicators: Draw a frequency density diagram for the logarithmically transformed fluorescence intensity signal (LFG) of a single sample to determine whether the distribution of blank control, standard sample and test sample is normal.

[0084] Quality control standards: For the blank control, confirm that the signal is distributed with narrow pulse width and high peak value, and the LFG distribution peak is within 3; for the standard samples and the serum samples to be tested, confirm that the signal is generally distributed with positive skewness, and the peak value is much larger than the blank control.

[0085] What to do when quality control fails: If the blank control and standard are abnormal, the quality control of the corresponding sample in the chip has failed; if the sample to be tested is abnormal, the sample quality needs to be confirmed and re-sampling should be considered.

[0086] 3) Grid positioning accuracy quality control

[0087] Quality control indicators: Adjusted RSquare and Root Mean Square (RMS) obtained by Mask Analysis of each sample except the negative control.

[0088] Quality control standard: Adjusted R Square ≥ 0.3 and RMS ≥ 0.3 are qualified.

[0089] What to do when quality control fails: Manually check the gridding of corner images. If it is manually confirmed that the gridding is incorrect, the sample fails quality control and needs to be retested.

[0090] 4) Outlier Analysis

[0091] Quality control index: the number of samples on a chip (24 samples, excluding negative controls) whose features represented by outliers account for ≤2%.

[0092] Quality control standard: If the number of the above samples does not exceed 2, the chip passes the quality control.

[0093] What to do when quality control fails: All samples on the chip need to be retested.

[0094] 5) Quality control peptide mean (coefficient of variation, CV) analysis

[0095] Quality control indicators: CV mean of the quality control peptide signal intensity within a single sample.

[0096] Quality control standard: The above CV mean value ≤ 1% is qualified.

[0097] What to do when quality control fails: The sample needs to be retested.

[0098] 3.2 System stability quality control

[0099] The signal intensities of all the standards tested in this batch (one standard per chip, i.e., 1 standard / 24 samples) were analyzed for correlation and CV values ​​to perform quality control of system stability.

[0100] 1) Standard product correlation

[0101] Quality control index: correlation coefficient of signal intensity of all standard samples tested in this batch.

[0102] Quality control standard: The above correlation coefficient is ≥0.8 and is qualified.

[0103] What to do when quality control fails: This batch of samples needs to be retested.

[0104] 2) Standard sample CV mean

[0105] Quality control indicators: CV mean of the signal intensity of all standard samples tested in this batch.

[0106] Quality control standard: The above CV mean is ≤4%.

[0107] What to do when quality control fails: This batch of samples needs to be retested.

[0108] 4. Data preprocessing:

[0109] 1) Get the original data FG:

[0110] The peptide chip technology V13 chip was used to perform sample detection according to the above standard process, and the signal values ​​of 126,501 peptide segments of the V13 chip were obtained. Each peptide segment signal value is called a feature, and its value range is 0 to 65535. The raw data is called FG (foreground) and stored in a GPR5 format file.

[0111] 2) Perform data correction:

[0112] The raw data FG of each peptide was extracted from the data matrix in GPR5 format. Since the raw fluorescence signal data was Log-Norm distributed, the FG was added with a constant of 100 and then logarithmically transformed to obtain LFG (Log-FG) to improve homoscedasticity. The measurement accuracy of the data was roughly proportional to the intensity. The peptide signal LFG of each array was subtracted from the median of all the peptides in the array to obtain the standardized data NLFG.

[0113] Example 2

[0114] 1. Characteristic peptide screening

[0115] This application uses NLFG data corrected by V13 peptide chip (see Example 1 for the experimental and data preprocessing process), and each sample contains 126051 features (ie, signal values ​​of 126501 peptide segments).

[0116] First, the data of the 219 training sets are reduced in dimension using the Sklearn PCA module (n_components = 2). The results of all training set samples are as follows: Figure 1(Gray dots are patients, white dots are normal people). Most of the normal samples can be clustered together, and the normal samples and HIV-infected samples have good distinction. It is considered that this batch of data can be used for model prediction. Since there are too many features in the original data, it is necessary to perform feature engineering on the training set, that is, only select important features for modeling and prediction. Using 219 samples in the training set, the variance threshold (VarianceThreshold) screening method of feature selection in Python's sklearn library is used for difference analysis. Variance threshold is a filter. Variance threshold screening in feature engineering is a filter method. After setting a threshold, features whose variance in the entire data set is less than the threshold will be removed. Based on biological considerations, when some feature values ​​fall into a basically consistent range, that is, when the upper and lower limits do not change much, the feature can be considered to be removed, so the feature engineering method of variance threshold is selected. Using the threshold VarianceThreshold (threshold = 0.9*(1-0.9)), all features greater than this threshold in the difference analysis results of AIDS patients and normal samples are screened out. This step initially screened a total of 832 peptides. In the subsequent steps, further screening will be performed among these 832 peptides.

[0117] The development environment of the model and feature peptide screening program for this application is Jupyter Notebook Python 3.8.5, and the main packages used are: sklearn, numpy, pandas, seaborn and matplotlib.

[0118] 2. Training the model

[0119] After screening the characteristic peptides, in order to obtain the best model to distinguish HIV-positive and normal samples, it is necessary to use machine learning methods to build a model, and adjust the corresponding parameters to train the model based on the results of model evaluation. The test data set used for modeling is 823 peptide signals obtained from the initial screening of 219 test samples.

[0120] The modeling method of the present invention uses the SVC model of the SVM module of the sklearn library, that is, the linear kernel in the SVC function of the support vector machine (SVC parameters: (kernel = "linear", probability = True, random_state = 100, class_weight = "balanced")) to model and predict the characteristic polypeptide. In order to prevent overfitting, we choose to train the model on the training set through five-fold cross validation. Cross validation is a commonly used method in modeling, among which K-fold cross validation is an effective way to reduce model overfitting. This method can better verify whether the model is true and effective, and at the same time adjust the model parameters so that the model achieves a better effect in the training set before predicting and evaluating on the test set. The training set data is divided into 5 groups, one subset of which is used as the test group, and the remaining 4 subsets of data are used as the training group. Repeat 5 times, and the test group and the training group are randomly divided each time. Finally, the average of the model test accuracy calculated 5 times is used as the accuracy of the model under this group (C). For the SVM model, C is the penalty factor and is the most important parameter. The larger the value, the smaller the tolerance for errors, and overfitting may occur. If C is too small, the tolerance rate will be too high, and such a model will be meaningless. At the same time, the number of characteristic peptides input to the model is adjusted by adjusting the variance threshold T, and the accuracy of the model under different characteristic conditions is calculated. The adjusted C value: 0.7-1 and the threshold of feature selection (represented by T, T = a*(1-a), represented by T(a) below, and the number in brackets represents the value of a). The results of parameter adjustment within the range of T(0.8) to T(0.85) are shown in Tables 1 and 2. Under the condition of ensuring accurate judgment, by adjusting C and T, the prediction index values ​​of the model on the training set, such as accuracy and recall value, are gradually improved. Recall and accuracy are the main indicators used to evaluate the model. The changes in accuracy are shown in Tables 1 and 2, and the changes in recall are shown in Tables 3 and 4.

[0121] Precision: TP / (TP+FP), the ratio of correctly predicted positive numbers to all positive numbers.

[0122] Recall: TP / (TP+FN), the ratio of correctly predicted positive samples to all positive samples.

[0123] Table 1 Results of accuracy changes after adjusting training set parameters T and C

[0124]

[0125]

[0126] Table 2 Results of accuracy changes after adjusting training set parameters T and C

[0127]

[0128] Table 3 Results of recall rate changes after adjusting training set parameters T and C

[0129]

[0130] Table 4 Adjustment of training set parameters T and C, the results of recall rate change

[0131]

[0132] 3. Model Validation

[0133] Using the same model and parameters in the training set, we analyzed and verified the actual performance of the model on validation data sets from different sources. Using the variance thresholds T(0.8) to T(0.9) of the features selected in the training set above and the SVC model with the same parameters (C: 1-0.7) after fitting the training set (SVC parameters: kernel = "linear", probability = True, random_state = 100, class_weight = "balanced"), the C value of the SVC function and the feature threshold T, the corresponding precision and recall are shown in Tables 5, 6, and 7 below.

[0134] Table 5 Results of adjusting the parameter T of the test set - accuracy

[0135]

[0136] Table 6 Results of adjusting the parameter T of the test set - recall rate

[0137]

[0138] Table 7 Results of parameter T adjustment for the test set - threshold determination

[0139]

[0140] Table 7 shows that when the value of T is selected to be greater than 0.85, up to 0.9, the model prediction results are not ideal.

[0141] From the model prediction results, we can see that the model will overfit after the variance threshold T (a>0.85). That is, the accuracy index of the training set is lower than that of the test set under the same conditions. It is believed that the generalization effect of the model is not good after fitting the training set under this parameter, that is, the overfitting problem occurs. Therefore, the variance threshold range: T (a) should be a value in the range of 0.82-0.85, and the C value of SVC should be 0.7-0.9 to obtain better recall and accuracy. When the values ​​of a in T(a) are 0.82, 0.824, 0.838, 0.84, and 0.85, respectively, the peptides used in the model are shown in Table 10, Table 10+Table 11, Table 10+Table 11+Table 12, Table 10+Table 11+Table 12+Table 13, and Table 10+Table 11+Table 12+Table 13+Table 14, respectively (Note: "Table 10+Table 11" represents all the peptides in Table 10 and Table 11, and other similar expressions have similar meanings). Based on the training results of the training set and the prediction results of the test set, when the threshold T(a=0.82), that is, Threshold=0.82×(1-0.82), and the C value is 0.8 (SVC parameters=(C=0.8, kernel="linear", probability=True, random_state=100, class_weight="balanced")), the best effect can be obtained on the data. Under this parameter, the accuracy of the model test set reached 83%, and the recall rate reached 94%, that is, 94% of the positive samples could be detected. At this time, the confusion matrix obtained after fitting the training set model is as follows Figure 2 As shown, ROC Figure 3 As shown, the AUC value is 0.97. The confusion matrix obtained after the model predicts the test set is as follows Figure 4 The F1, recall rate, sensitivity, etc. of the model on the training set are shown in Table 8; the confusion matrix on the test set is shown in Figure 4 , F1, sensitivity, specificity and other values ​​are shown in Table 9.

[0142] F1 Score is an indicator used in statistics to measure the accuracy of a binary classification model. The value range is generally between 0 and 1. It is believed that the closer the F1 value is to 1, the better the model effect.

[0143] Table 8

[0144]

[0145] Table 9

[0146]

[0147]

[0148] When the threshold is Threshold = (0.82 (1-0.82)), the numbers and sequences of the 123 characteristic peptides selected in the model are shown in Table 10. These 123 characteristic peptides are marked with white dots on the volcano plot of the training set, as shown in Table 10. Figure 5 As shown in Figure 2. P.adj is the adjusted P value and FDR is the false positive rate.

[0149] Table 10

[0150]

[0151]

[0152]

[0153] Changing the value of T will change the number of peptides included in the model. When T = (0.824 × (1-0.824)), based on the aforementioned 123 peptides, the model will automatically add 6 characteristic peptides (a total of 129) as shown in Table 11. The model performance is shown in Table 3.

[0154] Table 11

[0155]

[0156] When T = (0.838 × (1-0.838)), based on the aforementioned 129 peptides (Table 10 + Table 11), the model automatically adds 34 characteristic peptides (a total of 163) as shown in Table 12. The model performance is shown in Table 3.

[0157] Table 12

[0158]

[0159] When T=(0.84×(1-0.84)), on the basis of the 163 polypeptides in Table 10+Table 11+Table 12, 9 additional characteristic polypeptides (a total of 172) are shown in Table 13.

[0160] Table 13

[0161]

[0162] When T=(0.85×(1-0.85)), on the basis of the 172 polypeptides in Table 10+Table 11+Table 12+Table 13, 35 additional characteristic polypeptides (a total of 207) are added as shown in Table 14.

[0163] Table 14

[0164]

[0165]

[0166] From the above description, it can be seen that the above embodiments of the present invention achieve the following technical effects: The present invention utilizes the polypeptide chip technology and uses a known training data set to screen out significantly different polypeptides that can be used for diagnosing HIV. The above-mentioned characteristic polypeptides are screened out from the significantly different polypeptides by machine learning methods, and a prediction model is established. After performing polypeptide chip detection on the sample to be tested / test set sample, the signal values of the corresponding characteristic polypeptides can be obtained. The signal values are input into the trained model, and after model calculation, the probability of each sample being HIV positive / negative is obtained. According to the corresponding probability and the set threshold in the model, the model will output the final judgment result (positive / negative) of the sample. At the same time, the characteristic peptide segments discovered by the present invention that can be used for diagnosing HIV can also be used to design / customize a dedicated polypeptide chip.

[0167] The present invention can detect the HIV virus infection of a sample based on the polypeptide chip experimental technology. Compared with the currently commonly used AIDS diagnosis method, enzyme-linked immunosorbent assay (ELISA), the combination of the sample and the polypeptide chip is based on protein-protein interactions, and the fluorescence signal detection process is not sensitive to temperature and has lower requirements for the experimental environment. The chip can be stored at room temperature for a long time without refrigeration, and the polypeptide chip experimental process can be fully automated. Not only is the experimental process simple, reducing the risk of infection for the experimenter during the operation process, but also minimizing the impact of the environment on the results. Moreover, when judging the HIV infection result, it is not directly judged by the intensity of the color reaction, but after reading the fluorescence signal of the characteristic polypeptide (i.e., the fluorescence signal intensity information), a model is constructed using machine learning algorithms to judge the result, ensuring the accuracy and reliability of the detection result. The polypeptide chip detection technology can achieve full automation from experiment to result output, ensuring the stability of the detection result. Each chip obtains a TIFF image file as the original data. The original data is scanned, pre-processed, quality controlled, analyzed, and the results are interpreted through a laser scanning system and software. The entire experiment and data processing process has the advantages of high throughput, automation, high sensitivity, and multivariate analysis.

[0168] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. For those skilled in the art, the present invention can have various changes and modifications. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention. Sequence Listing <110> Zhuhai Carboncloud Intelligent Technology Co., Ltd. <120> Polypeptide Chip and Its Application in the Preparation of AIDS Diagnostic Products <130> PN175523SZTY <160> 207 <170> SIPOSequenceListing 1.0 <210> 1 <211> 6 <212> PRT <213> Artificial Sequence <220> <221> VARIANT <222> (1)..(6) <223> Peptides <400> 1 Gly Ala Pro Pro Asp Gly 1 5 <210> 2 <211> 9 <212> PRT <213> Artificial Sequence <220> <221> VARIANT <222> (1)..(9) <223> Peptides <400> 2 Ser Ser Ala Lys Lys Val Phe Ser Asp 1 5 <210> 3 <211> 10 <212> PRT <213> Artificial Sequence <220> <221> VARIANT <222> (1)..(10) <223> Peptides <400> 3 Asp Asp Phe Ser Asp Gly Ala Asn Lys Glu 1 5 10 <210> 4 <211> 10 <212> PRT <213> Artificial Sequence <220> <221> VARIANT <222> (1)..(10) <223> Peptides <400> 4 Lys Arg Arg Pro Trp Phe Ser His Ser Asp 1 5 10 <210> 5 <211> 9 <212> PRT <213> Artificial Sequence <220> <221> VARIANT <222> (1)..(9) <223> Peptides <400> 5 Ser Ala Pro Gly Lys Val Ser Asp Gly 1 5 <210> 6 <211> 8 <212> PRT <213> Artificial Sequence <220> <221> VARIANT <222> (1)..(8) <223> Peptides <400> 6 Gly Ser Pro Val Arg Val Leu Glu 1 5 <210> 7 <211> 8 <212> PRT <213> Artificial Sequence <220> <221> VARIANT <222> (1)..(8) <223> Peptides <400> 7 Ser Ser Arg Arg Phe Glu Asp Gly 1 5 <210> 8 <211> 9 <212> PRT <213> Artificial Sequence <220> <221> VARIANT <222> (1)..(9) <223> Peptides <400> 8 Ser Ser Pro Lys Arg Val Gly Phe Asp 1 5 <210> 9 <211> 9 <212> PRT <213> Artificial Sequence <220> <221> VARIANT <222> (1)..(9) <223> Peptides <400> 9 Pro Lys Lys Gln Val Pro Arg Glu Asp 1 5 <210> 10 <211> 7 <212> PRT <213> Artificial Sequence <220> <221> VARIANT <222> (1)..(7) <223> Peptides <400> 10 Leu Gly Ala Gly Gly Tyr Gly 1 5 <210> 11 <211> 9 <212> PRT <213> Artificial Sequence <220> <221> VARIANT <222> (1)..(9) <223> Peptides <400> 11 Ser Ser Arg Val Ser Ala Lys His Glu 1 5 <210> 12 <211> 7 <212> PRT <213> Artificial Sequence <220> <221> VARIANT <222> (1)..(7) <223> Peptides <400> 12 Asn Asn Ser Ser Tyr Asp Gly 1 5 <210> 13 <211> 9 <212> PRT <213> Artificial Sequence <220> <221> VARIANT <222> (1)..(9) <223> Peptides <400> 13 Ser Ser Lys Asp Trp Asn His Asp Gly 1 5 <210> 14 <211> 9 <212> PRT <213> Artificial Sequence <220> <221> VARIANT <222> (1)..(9) <223> Peptides <400> 14 Asn Gly Lys Leu Val Ser Leu Ser Gly 1 5 <210> 15 <211> 8 <212> PRT <213> Artificial Sequence <220> <221> VARIANT <222> (1)..(8) <223> Peptides <400> 15 Pro Pro Ala Trp Gly Ser Arg Gly 1 5 <210> 16 <211> 11 <212> PRT <213> Artificial Sequence <220> <221> VARIANT <222> (1)..(11) <223> Peptides <400> 16 Trp Arg His Asn Gln Lys Tyr Val Pro Val Glu 1 5 10 <210> 17 <211> 5 <212> PRT <213> Artificial Sequence <220> <221> VARIANT <222> (1)..(5) <223> Peptides <400> 17 Leu Leu Gln Pro Gly 1 5 <210> 18 <211> 8 <212> PRT <213> Artificial Sequence <220> <221> VARIANT <222> (1)..(8) <223> Peptides <400> 18 Leu Val Gln Pro Gln Phe Asp Gly 1 5 <210> 19 <211> 11 <212> PRT <213> Artificial Sequence <220> <221> VARIANT <222> (1)..(11) <223> Peptides <400> 19 Pro Pro Tyr Trp Asn Lys Asn Lys Pro Asp Gly 1 5 10 <210> 20 <211> 10 <212> PRT <213> Artificial Sequence <220> <221> VARIANT <222> (1)..(10) <223> Peptides <400> 20 Gly Ala Pro Asn Leu Glu Gly Arg Leu Asp 1 5 10 <210> twenty one <211> 8 <212> PRT <213> Artificial Sequence <220> <221> VARIANT <222> (1)..(8) <223> Peptides <400> twenty one Ser Ser Lys Gly Ala Glu Asp Gly 1 5 <210> twenty two <211> 7 <212> PRT <213> Artificial Sequence <220> <221> VARIANT <222> (1)..(7) <223> Peptides <400> twenty two Ser Ser Arg Val Arg His Glu 1 5 <210> twenty three <211> 7 <212> PRT <213> Artificial Sequence <220> <221> VARIANT <222> (1)..(7) <223> Peptides <400> twenty three Ser Ser Arg Ala Ser Asp Gly 1 5 <210> twenty four <211> 8 <212> PRT <213> Artificial Sequence <220> <221> VARIANT <222> (1)..(8) <223> Peptides <400> twenty four Arg Arg Val Phe Phe Arg Glu Asp 1 5 <210> 25 <211> 7 <212> PRT <213> Artificial Sequence <220> <221> VARIANT <222> (1)..(7) <223> Peptides <400> 25 Ser Ser Lys Trp Val Glu Asp 1 5 <210> 26 <211> 7 <212> PRT <213> Artificial Sequence <220> <221> VARIANT <222> (1)..(7) <223> Peptides <400> 26 Leu Arg Gln Pro Glu Asp Gly 1 5 <210> 27 <211> 10 <212> PRT <213> Artificial Sequence <220> <221> VARIANT <222> (1)..(10) <223> Peptides <400> 27 Asp Gly Val Lys Phe Ser Tyr Gln His Ser 1 5 10 <210> 28 <211> 11 <212> PRT <213> Artificial Sequence <220> <221> VARIANT <222> (1)..(11) <223> Peptides <400> 28 Lys Leu Leu Gly Ala Phe Asp Pro Arg His Asp 1 5 10 <210> 29 <211> 9 <212> PRT <213> Artificial Sequence <220> <221> VARIANT <222> (1)..(9) <223> Peptides <400> 29 Ser Ser Lys Ser Glu Asp Pro Lys Gly 1 5 <210> 30 <211> 10 <212> PRT <213> Artificial Sequence <220> <221> VARIANT <222> (1)..(10) <223> Peptides <400> 30 Ser Ala Arg Pro Arg His Tyr Ala Val Glu 1 5 10 <210> 31 <211> 10 <212> PRT <213> Artificial Sequence <220> <221> VARIANT <222> (1)..(10) <223> Peptides <400> 31 Tyr Asp Gly Glu Pro Leu Ser Tyr Asp Gly 1 5 10 <210> 32 <211> 8 <212> PRT <213> Artificial Sequence <220> <221> VARIANT <222> (1)..(8) <223> Peptides <400> 32 Ser Ser Arg His Ser Ser Glu Asp 1 5 <210> 33 <211> 7 <212> PRT <213> Artificial Sequence <220> <221> VARIANT <222> (1)..(7) <223> Peptides <400> 33 Gly Ser Pro Leu Leu Glu Asp 1 5 <210> 34 <211> 7 <212> PRT <213> Artificial Sequence <220> <221> VARIANT <222> (1)..(7) <223> Peptides <400> 34 Gly Ala Pro Ser His Glu Gly 1 5 <210> 35 <211> 9 <212> PRT <213> Artificial Sequence <220> <221> VARIANT <222> (1)..(9) <223> Peptides <400> 35 Ser Ser Lys Val Phe Glu Pro Tyr Gly 1 5 <210> 36 <211> 9 <212> PRT <213> Artificial Sequence <220> <221> VARIANT <222> (1)..(9) <223> Peptides <400> 36 Lys Leu Leu Asn Ser Gln His Ser Gly 1 5 <210> 37 <211> 8 <212> PRT <213> Artificial Sequence <220> <221> VARIANT <222> (1)..(8) <223> Peptides <400> 37 Lys Leu Leu Asn Ala Asn Gln Gly 1 5 <210> 38 <211> 11 <212> PRT <213> Artificial Sequence <220> <221> VARIANT <222> (1)..(11) <223> Peptides <400> 38 Asn Pro Val Glu Gly Ala Val Ser Ala Arg Asp 1 5 10 <210> 39 <211> 9 <212> PRT <213> Artificial Sequence <220> <221> VARIANT <222> (1)..(9) <223> Peptides <400> 39 Val Arg Ser Asp Gly Glu His Leu Ser 1 5 <210> 40 <211> 8 <212> PRT <213> Artificial Sequence <220> <221> VARIANT <222> (1)..(8) <223> Peptides <400> 40 Ser Ser Arg Asn Lys Leu Ser Asp 1 5 <210> 41 <211> 10 <212> PRT <213> Artificial Sequence <220> <221> VARIANT <222> (1)..(10) <223> Peptides <400> 41 Asn Val Asn Gln Tyr Asn Ser Asn Asp Gly 1 5 10 <210> 42 <211> 11 <212> PRT <213> Artificial Sequence <220> <221> VARIANT <222> (1)..(11) <223> Peptides <400> 42 Asn Tyr Val Asp His Val Leu Asn His Glu Gly 1 5 10 <210> 43 <211> 6 <212> PRT <213> Artificial Sequence <220> <221> VARIANT <222> (1)..(6) <223> Peptides <400> 43 Ser Ser Arg Arg Glu Asp 1 5 <210> 44 <211> 9 <212> PRT <213> Artificial Sequence <220> <221> VARIANT <222> (1)..(9) <223> Peptides <400> 44 Gly Ala Pro Arg Leu Ser Asp Ser Gly 1 5 <210> 45 <211> 9 <212> PRT <213> Artificial Sequence <220> <221> VARIANT <222> (1)..(9) <223> Peptides <400> 45 Ser Ala Lys Pro Phe Asp Pro Leu Ser 1 5 <210> 46 <211> 9 <212> PRT <213> Artificial Sequence <220> <221> VARIANT <222> (1)..(9) <223> Peptides <400> 46 Ser Ser Lys Phe Lys His Val Phe Asp 1 5 <210> 47 <211> 7 <212> PRT <213> Artificial Sequence <220> <221> VARIANT <222> (1)..(7) <223> Peptides <400> 47 Ser Ala Pro Val Arg Val Gly 1 5 <210> 48 <211> 8 <212> PRT <213> Artificial Sequence <220> <221> VARIANT <222> (1)..(8) <223> Peptides <400> 48 Gln Asp Val Lys Tyr Val Ser Glu 1 5 <210> 49 <211> 9 <212> PRT <213> Artificial Sequence <220> <221> VARIANT <222> (1)..(9) <223> Peptides <400> 49 Asn Lys Leu Val Ala Tyr Val Asp Gly 1 5 <210> 50 <211> 5 <212> PRT <213> Artificial Sequence <220> <221> VARIANT <222> (1)..(5) <223> Peptides <400> 50 Ser Ser Lys Ala Gly 1 5 <210> 51 <211> 8 <212> PRT <213> Artificial Sequence <220> <221> VARIANT <222> (1)..(8) <223> Peptides <400> 51 Ser Ser Lys Glu Asn Lys Leu Glu 1 5 <210> 52 <211> 8 <212> PRT <213> Artificial Sequence <220> <221> VARIANT <222> (1)..(8) <223> Peptides <400> 52 Ser Ser Lys Glu Gln Arg Ser Gly 1 5 <210> 53 <211> 11 <212> PRT <213> Artificial Sequence <220> <221> VARIANT <222> (1)..(11) <223> Peptides <400> 53 Ser Ala Lys Glu Asp Ala Phe Asp Pro Asp Gly 1 5 10 <210> 54 <211> 10 <212> PRT <213> Artificial Sequence <220> <221> VARIANT <222> (1)..(10) <223> Peptides <400> 54 Val Ala Gln Lys Gln Ser Pro Lys Leu Phe 1 5 10 <210> 55 <211> 10 <212> PRT <213> Artificial Sequence <220> <221> VARIANT <222> (1)..(10) <223> Peptides <400> 55 Ser Ala Lys Asp Phe Ser Gln Arg His Val 1 5 10 <210> 56 <211> 10 <212> PRT <213> Artificial Sequence <220> <221> VARIANT <222> (1)..(10) <223> Peptides <400> 56 Pro Pro Ala Trp His Arg Ala Gln Ser Gly 1 5 10 <210> 57 <211> 8 <212> PRT <213> Artificial Sequence <220> <221> VARIANT <222> (1)..(8) <223> Peptides <400> 57 Ser Ser Lys Arg His Gly Asp Gly 1 5 <210> 58 <211> 9 <212> PRT <213> Artificial Sequence <220> <221> VARIANT <222> (1)..(9) <223> Peptides <400> 58 Lys Leu Val Asn Ser Glu Trp Ser Glu 1 5 <210> 59 <211> 10 <212> PRT <213> Artificial Sequence <220> <221> VARIANT <222> (1)..(10) <223> Peptides <400> 59 Ser Ser Arg Pro Asn Gln Lys His Phe Ser 1 5 10 <210> 60 <211> 8 <212> PRT <213> Artificial Sequence <220> <221> VARIANT <222> (1)..(8) <223> Peptides <400> 60 Ser Ser Arg Ser Trp Ser Glu Gly 1 5 <210> 61 <211> 8 <212> PRT <213> Artificial Sequence <220> <221> VARIANT <222> (1)..(8) <223> Peptides <400> 61 Gly Ser Pro Leu Pro Glu Asp Gly 1 5 <210> 62 <211> 12 <212> PRT <213> Artificial Sequence <220> <221> VARIANT <222> (1)..(12) <223> Peptides <400> 62 Lys His Ala Ser Asp Asn Phe Glu Tyr His Ser Asp 1 5 10 <210> 63 <211> 11 <212> PRT <213> Artificial Sequence <220> <221> VARIANT <222> (1)..(11) <223> Peptides <400> 63 Ser Asp Phe Asp Gly Glu Pro Gln Arg Phe Gly 1 5 10 <210> 64 <211> 10 <212> PRT <213> Artificial Sequence <220> <221> VARIANT <222> (1)..(10) <223> Peptides <400> 64 Ser Ala Lys Asp Ser Tyr Lys His Phe Ser 1 5 10 <210> 65 <211> 7 <212> PRT <213> Artificial Sequence <220> <221> VARIANT <222> (1)..(7) <223> Peptides <400> 65 Ser Ser Arg Gly Phe Asp Gly 1 5 <210> 66 <211> 7 <212> PRT <213> Artificial Sequence <220> <221> VARIANT <222> (1)..(7) <223> Peptides <400> 66 Ser Ser Lys Gly Ser Glu Gly 1 5 <210> 67 <211> 8 <212> PRT <213> Artificial Sequence <220> <221> VARIANT <222> (1)..(8) <223> Peptides <400> 67 Gln Asp Tyr Asp Asp Pro Asp Gly 1 5 <210> 68 <211> 9 <212> PRT <213> Artificial Sequence <220> <221> VARIANT <222> (1)..(9) <223> Peptides <400> 68 Ala Glu Tyr Trp Lys Tyr Lys Lys Gly 1 5 <210> 69 <211> 10 <212> PRT <213> Artificial Sequence <220> <221> VARIANT <222> (1)..(10) <223> Peptides <400> 69 Val Lys Val Asp Asn Gln Leu Asn Phe Gly 1 5 10 <210> 70 <211> 9 <212> PRT <213> Artificial Sequence <220> <221> VARIANT <222> (1)..(9) <223> Peptides <400> 70 Gly Ala Pro Ala Gln Asp Arg Leu Gly 1 5 <210> 71 <211> 8 <212> PRT <213> Artificial Sequence <220> <221> VARIANT <222> (1)..(8) <223> Peptides <400> 71 Gly Ala Pro Tyr Gly Lys Leu Gly 1 5 <210> 72 <211> 5 <212> PRT <213> Artificial Sequence <220> <221> VARIANT <222> (1)..(5) <223> Peptides <400> 72 Ser Ser Lys Pro Gly 1 5 <210> 73 <211> 9 <212> PRT <213> Artificial Sequence <220> <221> VARIANT <222> (1)..(9) <223> Peptides <400> 73 Gly Ala Pro Ser Glu Asn Lys Val Asp 1 5 <210> 74 <211> 11 <212> PRT <213> Artificial Sequence <220> <221> VARIANT <222> (1)..(11) <223> Peptides <400> 74 Arg Ser Glu Ala Asp Tyr His Glu Tyr His Glu 1 5 10 <210> 75 <211> 7 <212> PRT <213> Artificial Sequence <220> <221> VARIANT <222> (1)..(7) <223> Peptides <400> 75 Gly Ala Pro Val His Phe Glu 1 5 <210> 76 <211> 9 <212> PRT <213> Artificial Sequence <220> <221> VARIANT <222> (1)..(9) <223> Peptides <400> 76 Gly Ala Val Leu Pro Leu Ala Asp Gly 1 5 <210> 77 <211> 8 <212> PRT <213> Artificial Sequence <220> <221> VARIANT <222> (1)..(8) <223> Peptides <400> 77 Ser Ser Lys Tyr Val Phe Glu Gly 1 5 <210> 78 <211> 8 <212> PRT <213> Artificial Sequence <220> <221> VARIANT <222> (1)..(8) <223> Peptides <400> 78 Ser Ser Lys Leu Gly Pro Glu Gly 1 5 <210> 79 <211> 8 <212> PRT <213> Artificial Sequence <220> <221> VARIANT <222> (1)..(8) <223> Peptides <400> 79 Tyr Asp Ser Ala Asp Gly Asp Gly 1 5 <210> 80 <211> 8 <212> PRT <213> Artificial Sequence <220> <221> VARIANT <222> (1)..(8) <223> Peptides <400> 80 Ser Ser Arg Phe Lys Val Asp Gly 1 5 <210> 81 <211> 9 <212> PRT <213> Artificial Sequence <220> <221> VARIANT <222> (1)..(9) <223> Peptides <400> 81 Leu Gly Lys Leu Leu Asp Ala Gln Gly 1 5 <210> 82 <211> 8 <212> PRT <213> Artificial Sequence <220> <221> VARIANT <222> (1)..(8) <223> Peptides <400> 82 Ser Ala Lys Ala Asp Phe Asp Gly 1 5 <210> 83 <211> 10 <212> PRT <213> Artificial Sequence <220> <221> VARIANT <222> (1)..(10) <223> Peptides <400> 83 Ala Ser Arg Pro Lys His Ser Asn Glu Asp 1 5 10 <210> 84 <211> 8 <212> PRT <213> Artificial Sequence <220> <221> VARIANT <222> (1)..(8) <223> Peptides <400> 84 Pro Pro Lys Trp Leu Arg Glu Gly 1 5 <210> 85 <211> 10 <212> PRT <213> Artificial Sequence <220> <221> VARIANT <222> (1)..(10) <223> Peptides <400> 85 Tyr Gln Lys Ser Val Pro Gln Glu Phe Gly 1 5 10 <210> 86 <211> 9 <212> PRT <213> Artificial Sequence <220> <221> VARIANT <222> (1)..(9) <223> Peptides <400> 86 Gly Ala Pro Arg Asn Gln Leu Asp Gly 1 5 <210> 87 <211> 7 <212> PRT <213> Artificial Sequence <220> <221> VARIANT <222> (1)..(7) <223> Peptides <400> 87 Ser Ser Arg Glu Tyr Glu Gly 1 5 <210> 88 <211> 9 <212> PRT <213> Artificial Sequence <220> <221> VARIANT <222> (1)..(9) <223> Peptides <400> 88 Ser Ala Lys Asn Val Glu Val Phe Gly 1 5 <210> 89 <211> 9 <212> PRT <213> Artificial Sequence <220> <221> VARIANT <222> (1)..(9) <223> Peptides <400> 89 Ser Ser Lys Asn His Val Phe Ser Glu 1 5 <210> 90 <211> 6 <212> PRT <213> Artificial Sequence <220> <221> VARIANT <222> (1)..(6) <223> Peptides <400> 90 Gly Ala Pro His Gln Gly 1 5 <210> 91 <211> 9 <212> PRT <213> Artificial Sequence <220> <221> VARIANT <222> (1)..(9) <223> Peptides <400> 91 Gly Ala Pro Leu Gln Val Leu Ser Asp 1 5 <210> 92 <211> 8 <212> PRT <213> Artificial Sequence <220> <221> VARIANT <222> (1)..(8) <223> Peptides <400> 92 Pro Pro Lys Trp Gly Trp Lys Asp 1 5 <210> 93 <211> 8 <212> PRT <213> Artificial Sequence <220> <221> VARIANT <222> (1)..(8) <223> Peptides <400> 93 Lys Gly Asn Asn Glu Lys Glu Asp 1 5 <210> 94 <211> 10 <212> PRT <213> Artificial Sequence <220> <221> VARIANT <222> (1)..(10) <223> Peptides <400> 94 Lys Asp Asn Tyr Phe Asp Gln Lys Val Ser 1 5 10 <210> 95 <211> 7 <212> PRT <213> Artificial Sequence <220> <221> VARIANT <222> (1)..(7) <223> Peptides <400> 95 Asn Asp Gly Asp Glu Gly Gly 1 5 <210> 96 <211> 9 <212> PRT <213> Artificial Sequence <220> <221> VARIANT <222> (1)..(9) <223> Peptides <400> 96 Val Lys Glu Ala Pro Ala Asn Phe Asp 1 5 <210> 97 <211> 7 <212> PRT <213> Artificial Sequence <220> <221> VARIANT <222> (1)..(7) <223> Peptides <400> 97 Asn Asn Ser His Arg Glu Gly 1 5 <210> 98 <211> 9 <212> PRT <213> Artificial Sequence <220> <221> VARIANT <222> (1)..(9) <223> Peptides <400> 98 Ser Ala Lys Ser Arg Val Phe Leu Asp 1 5 <210> 99 <211> 7 <212> PRT <213> Artificial Sequence <220> <221> VARIANT <222> (1)..(7) <223> Peptides <400> 99 Asn Asn Ser Ser Leu Glu Gly 1 5 <210> 100 <211> 9 <212> PRT <213> Artificial Sequence <220> <221> VARIANT <222> (1)..(9) <223> Peptides <400> 100 Gly Ala Pro Gln Asp Lys Arg Glu Asp 1 5 <210> 101 <211> 11 <212> PRT <213> Artificial Sequence <220> <221> VARIANT <222> (1)..(11) <223> Peptides <400> 101 Lys Leu Trp Gln Val Tyr Asn Glu Arg Ser Glu 1 5 10 <210> 102 <211> 9 <212> PRT <213> Artificial Sequence <220> <221> VARIANT <222> (1)..(9) <223> Peptides <400> 102 Ser Ala Lys Pro Leu Phe Glu Leu Gly 1 5 <210> 103 <211> 11 <212> PRT <213> Artificial Sequence <220> <221> VARIANT <222> (1)..(11) <223> Peptides <400> 103 Tyr Gln Lys Asp Arg Val Gly Asp Ala Val Glu 1 5 10 <210> 104 <211> 7 <212> PRT <213> Artificial Sequence <220> <221> VARIANT <222> (1)..(7) <223> Peptides <400> 104 Ser Pro Pro Arg Lys Leu Gly 1 5 <210> 105 <211> 9 <212> PRT <213> Artificial Sequence <220> <221> VARIANT <222> (1)..(9) <223> Peptides <400> 105 Gly Ala Pro Gly Ala Arg His Leu Gly 1 5 <210> 106 <211> 9 <212> PRT <213> Artificial Sequence <220> <221> VARIANT <222> (1)..(9) <223> Peptides <400> 106 Gly Ala Pro Pro Ala Trp Lys Ser Asp 1 5 <210> 107 <211> 8 <212> PRT <213> Artificial Sequence <220> <221> VARIANT <222> (1)..(8) <223> Peptides <400> 107 Ser Ser Arg Ser Gln Ser Asp Gly 1 5 <210> 108 <211> 9 <212> PRT <213> Artificial Sequence <220> <221> VARIANT <222> (1)..(9) <223> Peptides <400> 108 Gln Asn Glu Leu Asp Pro Gln Lys Asp 1 5 <210> 109 <211> 8 <212> PRT <213> Artificial Sequence <220> <221> VARIANT <222> (1)..(8) <223> Peptides <400> 109 Asp Val Lys Phe Phe Glu Asp Gly 1 5 <210> 110 <211> 9 <212> PRT <213> Artificial Sequence <220> <221> VARIANT <222> (1)..(9) <223> Peptides <400> 110 Ala Ser Ala Pro Lys Lys Leu Phe Ser 1 5 <210> 111 <211> 10 <212> PRT <213> Artificial Sequence <220> <221> VARIANT <222> (1)..(10) <223> Peptides <400> 111 Gly Ala Pro Ala Ser Glu Pro Asn Lys Asp 1 5 10 <210> 112 <211> 6 <212> PRT <213> Artificial Sequence <220> <221> VARIANT <222> (1)..(6) <223> Peptides <400> 112 Ser Ser Lys Asp Asp Gly 1 5 <210> 113 <211> 9 <212> PRT <213> Artificial Sequence <220> <221> VARIANT <222> (1)..(9) <223> Peptides <400> 113 Gly Ser Pro Val Pro Arg Val Phe Gly 1 5 <210> 114 <211> 9 <212> PRT <213> Artificial Sequence <220> <221> VARIANT <222> (1)..(9) <223> Peptides <400> 114 Asp Val Lys Phe Ala Arg His Phe Glu 1 5 <210> 115 <211> 7 <212> PRT <213> Artificial Sequence <220> <221> VARIANT <222> (1)..(7) <223> Peptides <400> 115 Ser Ser Lys Val Gln Asp Gly 1 5 <210> 116 <211> 8 <212> PRT <213> Artificial Sequence <220> <221> VARIANT <222> (1)..(8) <223> Peptides <400> 116 Asp Gln Val Lys Phe Val Glu Gly 1 5 <210> 117 <211> 8 <212> PRT <213> Artificial Sequence <220> <221> VARIANT <222> (1)..(8) <223> Peptides <400> 117 Gly Ala Pro Phe Ala Asn Gln Gly 1 5 <210> 118 <211> 6 <212> PRT <213> Artificial Sequence <220> <221> VARIANT <222> (1)..(6) <223> Peptides <400> 118 Gly Ala Pro His Lys Gly 1 5 <210> 119 <211> 7 <212> PRT <213> Artificial Sequence <220> <221> VARIANT <222> (1)..(7) <223> Peptides <400> 119 Gly Ala Pro Leu Leu Ser Gly 1 5 <210> 120 <211> 8 <212> PRT <213> Artificial Sequence <220> <221> VARIANT <222> (1)..(8) <223> Peptides <400> 120 Lys Ala Val Pro Leu Glu Ala Gly 1 5 <210> 121 <211> 10 <212> PRT <213> Artificial Sequence <220> <221> VARIANT <222> (1)..(10) <223> Peptides <400> 121 Ala Lys Gly Phe Pro Tyr Lys Trp Ser Gly 1 5 10 <210> 122 <211> 9 <212> PRT <213> Artificial Sequence <220> <221> VARIANT <222> (1)..(9) <223> Peptides <400> 122 Ser Ser Lys Trp Gln His Ser Asp Gly 1 5 <210> 123 <211> 9 <212> PRT <213> Artificial Sequence <220> <221> VARIANT <222> (1)..(9) <223> Peptides <400> 123 Ser Gln Lys Ser Phe Pro Val Phe Glu 1 5 <210> 124 <211> 7 <212> PRT <213> Artificial Sequence <220> <221> VARIANT <222> (1)..(7) <223> Peptides <400> 124 Ser Ser Arg Gly Glu Asp Gly 1 5 <210> 125 <211> 9 <212> PRT <213> Artificial Sequence <220> <221> VARIANT <222> (1)..(9) <223> Peptides <400> 125 Ser Ser Lys Ser Tyr Lys Phe Ser Gly 1 5 <210> 126 <211> 9 <212> PRT <213> Artificial Sequence <220> <221> VARIANT <222> (1)..(9) <223> Peptides <400> 126 Ala Asp Ser Ser Ala Lys Leu Ser Asp 1 5 <210> 127 <211> 9 <212> PRT <213> Artificial Sequence <220> <221> VARIANT <222> (1)..(9) <223> Peptides <400> 127 Val Ala Asp Val Ser Glu Tyr His Gly 1 5 <210> 128 <211> 9 <212> PRT <213> Artificial Sequence <220> <221> VARIANT <222> (1)..(9) <223> Peptides <400> 128 Ser Ala Lys Phe Ser Tyr Pro Ala Gly 1 5 <210> 129 <211> 8 <212> PRT <213> Artificial Sequence <220> <221> VARIANT <222> (1)..(8) <223> Peptides <400> 129 Gly Ser Pro Arg Pro Phe Glu Gly 1 5 <210> 131 <211> 8 <212> PRT <213> Artificial Sequence <220> <221> VARIANT <222> (1)..(8) <223> Peptides <400> 131 Lys Ala Glu Leu Gly Lys Arg Gly 1 5 <210> 130 <211> 9 <212> PRT <213> Artificial Sequence <220> <221> VARIANT <222> (1)..(9) <223> Peptides <400> 130 Ser Ser Val Pro Tyr Gln Lys Phe Glu 1 5 <210> 132 <211> 9 <212> PRT <213> Artificial Sequence <220> <221> VARIANT <222> (1)..(9) <223> Peptides <400> 132 Gly Ala Pro Asn Asp Pro His Ser Glu 1 5 <210> 133 <211> 12 <212> PRT <213> Artificial Sequence <220> <221> VARIANT <222> (1)..(12) <223> Peptides <400> 133 Lys Ser Asp Gly Ala Gln Tyr Asn Ala His Ser Gly 1 5 10 <210> 134 <211> 10 <212> PRT <213> Artificial Sequence <220> <221> VARIANT <222> (1)..(10) <223> Peptides <400> 134 Pro Val Pro Ala Glu Arg Leu Pro Val Gly 1 5 10 <210> 135 <211> 10 <212> PRT <213> Artificial Sequence <220> <221> VARIANT <222> (1)..(10) <223> Peptides <400> 135 Ser Ser Lys Arg Leu Ala Trp Arg Val Glu 1 5 10 <210> 136 <211> 9 <212> PRT <213> Artificial Sequence <220> <221> VARIANT <222> (1)..(9) <223> Peptides <400> 136 Ser Ser Lys His Gln His Val Ser Gly 1 5 <210> 137 <211> 9 <212> PRT <213> Artificial Sequence <220> <221> VARIANT <222> (1)..(9) <223> Peptides <400> 137 Lys Asp Arg Pro Gly Ala Arg Leu Phe 1 5 <210> 138 <211> 8 <212> PRT <213> Artificial Sequence <220> <221> VARIANT <222> (1)..(8) <223> Peptides <400> 138 Gly Arg Pro Val Pro Tyr Glu Gly 1 5 <210> 139 <211> 9 <212> PRT <213> Artificial Sequence <220> <221> VARIANT <222> (1)..(9) <223> Peptides <400> 139 Leu Ser His Gly Ala Ser Lys Phe Glu 1 5 <210> 140 <211> 9 <212> PRT <213> Artificial Sequence <220> <221> VARIANT <222> (1)..(9) <223> Peptides <400> 140 Ser Ala Arg Pro Phe Glu Ala His Asp 1 5 <210> 141 <211> 8 <212> PRT <213> Artificial Sequence <220> <221> VARIANT <222> (1)..(8) <223> Peptides <400> 141 Ser Ala Lys Ala Ser His Phe Glu 1 5 <210> 142 <211> 9 <212> PRT <213> Artificial Sequence <220> <221> VARIANT <222> (1)..(9) <223> Peptides <400> 142 Asn Asn Val Phe Lys Arg Ala Glu Asp 1 5 <210> 143 <211> 10 <212> PRT <213> Artificial Sequence <220> <221> VARIANT <222> (1)..(10) <223> Peptides <400> 143 Pro Val Pro Asn Asp Glu Gly Lys Ser Glu 1 5 10 <210> 144 <211> 10 <212> PRT <213> Artificial Sequence <220> <221> VARIANT <222> (1)..(10) <223> Peptides <400> 144 Ser Ala Lys His Phe Ser Gln Tyr Val Gly 1 5 10 <210> 145 <211> 9 <212> PRT <213> Artificial Sequence <220> <221> VARIANT <222> (1)..(9) <223> Peptides <400> 145 Gly Leu Gly Pro Lys Phe Val Glu Asp 1 5 <210> 146 <211> 7 <212> PRT <213> Artificial Sequence <220> <221> VARIANT <222> (1)..(7) <223> Peptides <400> 146 Pro Pro Arg Trp Ser Ser Glu 1 5 <210> 147 <211> 8 <212> PRT <213> Artificial Sequence <220> <221> VARIANT <222> (1)..(8) <223> Peptides <400> 147 Arg Asp Ser Arg Tyr Asn Val Gly 1 5 <210> 148 <211> 8 <212> PRT <213> Artificial Sequence <220> <221> VARIANT <222> (1)..(8) <223> Peptides <400> 148 Gly Ser Pro Pro Arg Phe Ser Gly 1 5 <210> 149 <211> 10 <212> PRT <213> Artificial Sequence <220> <221> VARIANT <222> (1)..(10) <223> Peptides <400> 149 Asn Asn Val Gln Lys Leu Gly Asn Glu Gly 1 5 10 <210> 150 <211> 11 <212> PRT <213> Artificial Sequence <220> <221> VARIANT <222> (1)..(11) <223> Peptides <400> 150 Glu Gln Phe Trp Lys His Gly Asn Arg Ser Gly 1 5 10 <210> 151 <211> 8 <212> PRT <213> Artificial Sequence <220> <221> VARIANT <222> (1)..(8) <223> Peptides <400> 151 Ser Ser Leu Pro Arg His Val Glu 1 5 <210> 152 <211> 7 <212> PRT <213> Artificial Sequence <220> <221> VARIANT <222> (1)..(7) <223> Peptides <400> 152 Ser Ser Lys Lys Phe Ser Gly 1 5 <210> 153 <211> 9 <212> PRT <213> Artificial Sequence <220> <221> VARIANT <222> (1)..(9) <223> Peptides <400> 153 Asn Asn Ser Ala Asn Asn His Glu Asp 1 5 <210> 154 <211> 8 <212> PRT <213> Artificial Sequence <220> <221> VARIANT <222> (1)..(8) <223> Peptides <400> 154 Glu Arg Arg Pro Phe Glu Asp Gly 1 5 <210> 155 <211> 10 <212> PRT <213> Artificial Sequence <220> <221> VARIANT <222> (1)..(10) <223> Peptides <400> 155 Lys Arg Asp Pro Asp Ser Ala Asn Phe Gly 1 5 10 <210> 156 <211> 7 <212> PRT <213> Artificial Sequence <220> <221> VARIANT <222> (1)..(7) <223> Peptides <400> 156 His Asn Ser Ser Gly Asp Gly 1 5 <210> 157 <211> 8 <212> PRT <213> Artificial Sequence <220> <221> VARIANT <222> (1)..(8) <223> Peptides <400> 157 Ala Asn Ser Ser Leu Asp Glu Gly 1 5 <210> 158 <211> 9 <212> PRT <213> Artificial Sequence <220> <221> VARIANT <222> (1)..(9) <223> Peptides <400> 158 Asn Asn Lys Pro Gln Arg Asp Phe Asp 1 5 <210> 159 <211> 7 <212> PRT <213> Artificial Sequence <220> <221> VARIANT <222> (1)..(7) <223> Peptides <400> 159 Gly Ser Pro Phe Lys His Gly 1 5 <210> 160 <211> 9 <212> PRT <213> Artificial Sequence <220> <221> VARIANT <222> (1)..(9) <223> Peptides <400> 160 Ser Ser Arg Leu Tyr Asn Arg His Glu 1 5 <210> 161 <211> 8 <212> PRT <213> Artificial Sequence <220> <221> VARIANT <222> (1)..(8) <223> Peptides <400> 161 Ser Asn Ser Ser Val Leu Ser Asp 1 5 <210> 162 <211> 10 <212> PRT <213> Artificial Sequence <220> <221> VARIANT <222> (1)..(10) <223> Peptides <400> 162 Gly Lys Val Gly Pro Arg His Arg Asp Gly 1 5 10 <210> 163 <211> 9 <212> PRT <213> Artificial Sequence <220> <221> VARIANT <222> (1)..(9) <223> Peptides <400> 163 Lys Gly Val Asn Asn Lys Phe Glu Asp 1 5 <210> 164 <211> 9 <212> PRT <213> Artificial Sequence <220> <221> VARIANT <222> (1)..(9) <223> Peptides <400> 164 Asn Asn Leu Val Pro Gln Lys Glu Asp 1 5 <210> 165 <211> 7 <212> PRT <213> Artificial Sequence <220> <221> VARIANT <222> (1)..(7) <223> Peptides <400> 165 Pro Pro Asn Trp Ser Ser Gly 1 5 <210> 166 <211> 10 <212> PRT <213> Artificial Sequence <220> <221> VARIANT <222> (1)..(10) <223> Peptides <400> 166 Asn Asp Ser Ser Asp Pro Arg Phe Ser Gly 1 5 10 <210> 167 <211> 9 <212> PRT <213> Artificial Sequence <220> <221> VARIANT <222> (1)..(9) <223> Peptides <400> 167 Asp Ser Asp Gly Glu Pro Phe Glu Asp 1 5 <210> 168 <211> 8 <212> PRT <213> Artificial Sequence <220> <221> VARIANT <222> (1)..(8) <223> Peptides <400> 168 Ser Arg Lys Asp Val Leu Phe Ser 1 5 <210> 169 <211> 8 <212> PRT <213> Artificial Sequence <220> <221> VARIANT <222> (1)..(8) <223> Peptides <400> 169 Gln Asp Phe Trp Lys Lys His Gly 1 5 <210> 170 <211> 9 <212> PRT <213> Artificial Sequence <220> <221> VARIANT <222> (1)..(9) <223> Peptides <400> 170 Glu Gly Gly His Leu Ser Arg Phe Glu 1 5 <210> 171 <211> 9 <212> PRT <213> Artificial Sequence <220> <221> VARIANT <222> (1)..(9) <223> Peptides <400> 171 Gly Ala Pro Lys Leu Asn Lys Phe Gly 1 5 <210> 172 <211> 10 <212> PRT <213> Artificial Sequence <220> <221> VARIANT <222> (1)..(10) <400> 172 Trp Arg Lys Val Leu Glu Glu Pro Lys Gly 1 5 10 <210> 173 <211> 7 <212> PRT <213> Artificial Sequence <220> <221> VARIANT <222> (1)..(7) <223> Peptides <400> 173 Phe Val Lys Ala Glu Asp Gly 1 5 <210> 174 <211> 9 <212> PRT <213> Artificial Sequence <220> <221> VARIANT <222> (1)..(9) <223> Peptides <400> 174 Asn Asn Lys Gln Pro Gln Ser Glu Gly 1 5 <210> 175 <211> 8 <212> PRT <213> Artificial Sequence <220> <221> VARIANT <222> (1)..(8) <223> Peptides <400> 175 Gly Asn Pro Asp Gln His Ser Gly 1 5 <210> 176 <211> 8 <212> PRT <213> Artificial Sequence <220> <221> VARIANT <222> (1)..(8) <223> Peptides <400> 176 Asn Asn Gly Gly Asn Gln Lys Asp 1 5 <210> 177 <211> 9 <212> PRT <213> Artificial Sequence <220> <221> VARIANT <222> (1)..(9) <223> Peptides <400> 177 Ser Ala Arg Val Pro Glu Gly Asn Asp 1 5 <210> 178 <211> 9 <212> PRT <213> Artificial Sequence <220> <221> VARIANT <222> (1)..(9) <223> Peptides <400> 178 Asn Asn Asn Asn Lys Val Leu Glu Gly 1 5 <210> 179 <211> 8 <212> PRT <213> Artificial Sequence <220> <221> VARIANT <222> (1)..(8) <223> Peptides <400> 179 Gly Ser Pro Ser Pro Glu Asp Gly 1 5 <210> 180 <211> 9 <212> PRT <213> Artificial Sequence <220> <221> VARIANT <222> (1)..(9) <223> Peptides <400> 180 Ala Lys Pro Phe Arg Ala Lys Asp Gly 1 5 <210> 181 <211> 10 <212> PRT <213> Artificial Sequence <220> <221> VARIANT <222> (1)..(10) <223> Peptides <400> 181 Val His Glu His Gly Pro Lys Phe Ser Asp 1 5 10 <210> 182 <211> 10 <212> PRT <213> Artificial Sequence <220> <221> VARIANT <222> (1)..(10) <223> Peptides <400> 182 Gln Arg Asn His Val Ser Lys Lys Glu Asp 1 5 10 <210> 183 <211> 7 <212> PRT <213> Artificial Sequence <220> <221> VARIANT <222> (1)..(7) <223> Peptides <400> 183 Phe Lys Lys Pro Ser Glu Asp 1 5 <210> 184 <211> 12 <212> PRT <213> Artificial Sequence <220> <221> VARIANT <222> (1)..(12) <223> Peptides <400> 184 Asn Asn Lys Phe Glu Asp Pro Val Gln Lys Arg Gly 1 5 10 <210> 185 <211> 10 <212> PRT <213> Artificial Sequence <220> <221> VARIANT <222> (1)..(10) <223> Peptides <400> 185 Ser Asn Ser Ala Asp Asn Gln Val Phe Gly 1 5 10 <210> 186 <211> 9 <212> PRT <213> Artificial Sequence <220> <221> VARIANT <222> (1)..(9) <223> Peptides <400> 186 Ser Gln Phe Trp Lys Ser Glu Asn Gly 1 5 <210> 187 <211> 7 <212> PRT <213> Artificial Sequence <220> <221> VARIANT <222> (1)..(7) <223> Peptides <400> 187 Gly Arg Pro Leu Leu Ser Glu 1 5 <210> 188 <211> 9 <212> PRT <213> Artificial Sequence <220> <221> VARIANT <222> (1)..(9) <223> Peptides <400> 188 Asn Pro Pro Ser Trp Asn Lys Arg Glu 1 5 <210> 189 <211> 9 <212> PRT <213> Artificial Sequence <220> <221> VARIANT <222> (1)..(9) <223> Peptides <400> 189 Lys Tyr Asp Pro Asp Ala Asn Arg Glu 1 5 <210> 190 <211> 8 <212> PRT <213> Artificial Sequence <220> <221> VARIANT <222> (1)..(8) <223> Peptides <400> 190 Asn Asn Ala Ser Gly Pro Ser Asp 1 5 <210> 191 <211> 8 <212> PRT <213> Artificial Sequence <220> <221> VARIANT <222> (1)..(8) <223> Peptides <400> 191 Val Asp Leu Val Asn Ser Glu Gly 1 5 <210> 192 <211> 9 <212> PRT <213> Artificial Sequence <220> <221> VARIANT <222> (1)..(9) <223> Peptides <400> 192 Asn Asn Glu Gly Pro Gly Asn Lys Gly 1 5 <210> 193 <211> 8 <212> PRT <213> Artificial Sequence <220> <221> VARIANT <222> (1)..(8) <223> Peptides <400> 193 Ser Ser Pro Ser Lys Val Glu Gly 1 5 <210> 194 <211> 8 <212> PRT <213> Artificial Sequence <220> <221> VARIANT <222> (1)..(8) <223> Peptides <400> 194 Gly Val Pro Val Lys Arg Asp Gly 1 5 <210> 195 <211> 8 <212> PRT <213> Artificial Sequence <220> <221> VARIANT <222> (1)..(8) <223> Peptides <400> 195 Lys Leu Val Ser Val Lys Asp Gly 1 5 <210> 196 <211> 9 <212> PRT <213> Artificial Sequence <220> <221> VARIANT <222> (1)..(9) <223> Peptides <400> 196 Asn Asn Asn Gly Gln Lys Arg Leu Asp 1 5 <210> 197 <211> 8 <212> PRT <213> Artificial Sequence <220> <221> VARIANT <222> (1)..(8) <223> Peptides <400> 197 Gly Lys Pro Val Pro Trp Glu Gly 1 5 <210> 198 <211> 9 <212> PRT <213> Artificial Sequence <220> <221> VARIANT <222> (1)..(9) <223> Peptides <400> 198 Phe Ser Lys Gln Val Ser Pro Ser Asp 1 5 <210> 199 <211> 8 <212> PRT <213> Artificial Sequence <220> <221> VARIANT <222> (1)..(8) <223> Peptides <400> 199 Pro Val Pro Asp Phe Ala Lys Gly 1 5 <210> 200 <211> 7 <212> PRT <213> Artificial Sequence <220> <221> VARIANT <222> (1)..(7) <223> Peptides <400> 200 Ser Ala Lys Gln Asp Glu Gly 1 5 <210> 201 <211> 9 <212> PRT <213> Artificial Sequence <220> <221> VARIANT <222> (1)..(9) <223> Peptides <400> 201 Ala Ser Asp Tyr Tyr Glu Tyr Gln Gly 1 5 <210> 202 <211> 10 <212> PRT <213> Artificial Sequence <220> <221> VARIANT <222> (1)..(10) <223> Peptides <400> 202 Asn Asn Gly Ala Pro Tyr Lys Arg Ser Asp 1 5 10 <210> 203 <211> 10 <212> PRT <213> Artificial Sequence <220> <221> VARIANT <222> (1)..(10) <223> Peptides <400> 203 Ser Ala Lys Gly Arg Val Glu Asn Glu Asp 1 5 10 <210> 204 <211> 10 <212> PRT <213> Artificial Sequence <220> <221> VARIANT <222> (1)..(10) <223> Peptides <400> 204 Lys Leu Val Gly Leu Ser Gly Lys Glu Gly 1 5 10 <210> 205 <211> 9 <212> PRT <213> Artificial Sequence <220> <221> VARIANT <222> (1)..(9) <223> Peptides <400> 205 Ser Asn Lys Ala Val Pro Arg Phe Asp 1 5 <210> 206 <211> 11 <212> PRT <213> Artificial Sequence <220> <221> VARIANT <222> (1)..(11) <223> Peptides <400> 206 Ala Asp Glu Gly Pro Tyr Lys Pro Lys Phe Gly 1 5 10 <210> 207 <211> 8 <212> PRT <213> Artificial Sequence <220> <221> VARIANT <222> (1)..(8) <223> Peptides <400> 207 Pro Asn Lys Ser Val Pro Phe Glu 1 5

Claims

1. A peptide chip, It is characterized in that The polypeptide chip includes characteristic polypeptides, and the characteristic polypeptides include The polypeptides shown in SEQ ID NOs: 1 to 123; or The polypeptides shown in SEQ ID NOs: 1 to 129; or The polypeptides shown in SEQ ID NOs: 1 to 163; or The polypeptides represented by SEQ ID NOs: 1 to 172; or The polypeptides shown in SEQ ID NOs: 1 to 207.

2. A method for constructing an AIDS prediction model. It is characterized in that The construction method comprises: a) obtaining the AIDS status of the sample population and biological samples, using a polypeptide chip containing the characteristic polypeptide of claim 1 to detect, image, and detect the fluorescence signal intensity of the biological samples, extracting the fluorescence signal intensity information of the characteristic polypeptide, and establishing a data set of the biological samples, The data set includes a mapping relationship between the AIDS status of the sample population and the fluorescence signal intensity information of the characteristic polypeptide in the biological sample; b) constructing a prediction model using the data set through a machine learning training method; The sample population includes healthy people and AIDS patients.

3. The construction method according to claim 2, It is characterized in that Before using the machine learning training, the data set is divided into a training set and a test set, the training set is used to build a classification model, and the test set is used to verify the model effect.

4. The construction method according to claim 2 or 3, It is characterized in that The machine learning training method includes any one or more of the following: a logistic regression model, a support vector machine model, or a naive Bayes model.

5. The construction method according to claim 2, It is characterized in that The biological sample includes serum or plasma.

6. An electronic device for AIDS diagnosis, It is characterized in that The electronic device comprises a diagnostic model, wherein the diagnostic model is constructed by using the construction method described in any one of claims 2 to 5. The diagnostic model uses the detection result of the polypeptide chip on the sample to be tested as the model input and outputs the diagnostic result. The detection result includes that obtained by detecting the sample to be tested using the polypeptide chip and extracting the fluorescence signal intensity information of the characteristic polypeptide.

7. The electronic device according to claim 6, It is characterized in that The sample to be tested includes serum or plasma.

8. A computer-readable storage medium, It is characterized in that The storage medium includes a stored program, wherein when the program is executed, the device where the storage medium is located is controlled to execute the construction method according to any one of claims 2 to 5.

9. A processor, It is characterized in that The processor is used to run a program, wherein the program executes the construction method described in any one of claims 2 to 5 when running.

10. Application of characteristic peptides in the preparation of products for diagnosing AIDS, It is characterized in that The characteristic polypeptides include polypeptides shown in SEQ ID NOs: 1 to 123; or polypeptides shown in SEQ ID NOs: 1 to 129; or polypeptides shown in SEQ ID NOs: 1 to 163; or polypeptides shown in SEQ ID NOs: 1 to 172; or polypeptides shown in SEQ ID NOs: 1 to 207.

11. The use according to claim 10, It is characterized in that The AIDS diagnostic product includes a polypeptide chip.

Citation Information

Patent Citations

  • Polypeptide bound with gp120 protein, polypeptide chip, and preparation method and application thereof

    CN104292304A

  • Polypeptide, composition thereof, kit containing polypeptide and application of polypeptide

    CN112574284A