Application of protein for predicting thyroid cancer risk in preparation of thyroid detection reagent

By detecting specific TCR β-strand CDR3 sequences in the peripheral blood of thyroid cancer, using high-throughput sequencing and machine learning algorithms, the limitations of early screening and diagnosis of thyroid cancer in the prior art are solved, and high sensitivity, specificity and low-cost detection effects are achieved.

CN120177798APending Publication Date: 2025-06-20920TH HOSPITAL OF THE JOINT LOGISTIC SUPPORT FORCE OF THE CHINESE PEOPLES LIBERATION ARMY
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202510395993.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-31
Publication Date
2025-06-20

AI Technical Summary

Technical Problem

The prior art has limitations in the early screening and diagnosis of thyroid cancer, and lacks convenient, inexpensive, and highly sensitive and specific methods.

Method used

By detecting specific TCR β-strand CDR3 sequences in the peripheral blood of thyroid cancer, a scoring system is constructed to judge the risk of thyroid cancer using high-throughput sequencing technology and machine learning algorithms.

Benefits of technology

High sensitivity and specific detection of early screening of thyroid cancer is achieved, reducing detection costs, and providing a non-invasive sampling method.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120177798A_ABST
    Figure CN120177798A_ABST
Patent Text Reader

Abstract

The invention discloses an application of a protein for predicting thyroid cancer risk in preparation of a thyroid detection reagent, and the protein or an RNA marker is a thyroid cancer tumor specific TCR sequence, which can be one of 50 sequences. The marker comprises at least one of 50 proteins shown in SEQ ID NO.1-50 or RNA sequences capable of being translated into the proteins. The kit comprises a protein marker reagent for detecting a blood sample of a subject and predicting the thyroid cancer risk. The high-throughput sequencing technology is adopted, meanwhile, a huge number of characteristic TCR sequences are compared, and compared with single detection of one or more markers, higher specificity and accuracy are achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of biomedicine, particularly to the production and composition of clinical diagnostic reagents, and specifically relates to the application of a protein for predicting the risk of thyroid cancer in the preparation of thyroid measurement reagents. Background Art

[0002] Thyroid cancer is a malignant tumor originating from thyroid follicular epithelial or parafollicular epithelial cells, and is also the most common malignant tumor in the endocrine system and head and neck. High-risk groups include women and those with genetic mutations or family history of thyroid cancer. Different pathological types of thyroid cancer have obvious differences in their pathogenesis, biological behavior, histological morphology, clinical manifestations, treatment methods, and prognosis. According to the origin and differentiation differences of the tumor, thyroid cancer is divided into the following types: Papillary thyroid carcinoma (PTC): The most common, accounting for about 85% - 90% of all thyroid cancers. Follicular thyroid carcinoma (FTC): Together with PTC, it is called differentiated thyroid cancer (DTC), and has a relatively good prognosis. Medullary thyroid carcinoma (MTC): The prognosis is between differentiated and undifferentiated types. And anaplastic thyroid cancer (ATC): It has a very high degree of malignancy, with a median survival time of only 7 - 10 months and a very poor prognosis.

[0003] The survival rate of thyroid cancer, especially differentiated thyroid cancer (including papillary and follicular carcinomas), is relatively high. Early diagnosis is the premise of standardized treatment. Early screening and diagnosis of thyroid cancer are of great significance in improving the quality of life of patients, reducing medical costs, and improving public health levels. Therefore, seeking simpler and more reliable early screening and diagnosis methods has great practical significance for the prevention and treatment of thyroid cancer.

[0004] Most early-stage thyroid cancers have no obvious clinical symptoms. Some patients may have painless lumps or nodules in the neck, usually detected by thyroid palpation and neck ultrasound during physical examinations. As the tumor grows, it may compress or invade adjacent organs or tissues, leading to manifestations such as difficulty breathing, swallowing difficulties, jugular vein distension, hoarseness of voice, flushing of the face, and tachycardia. Some patients may experience cervical lymph node metastasis and distant organ metastasis, mainly to the lungs, liver, and bones. MTC tumor cells secrete active substances such as calcitonin and serotonin, which can cause symptoms such as diarrhea, palpitations, and flushing of the face. Currently, the main methods for clinical diagnosis of thyroid cancer are as follows: Ultrasound examination: High-resolution ultrasound is the preferred method for evaluating thyroid cancer and can detect nodules as small as 2 mm in the early stage of the thyroid. Thyroid function tests: Include the detection of indicators such as thyroid hormones (T3, T4), thyroid-stimulating hormone (TSH), etc. Fine needle aspiration biopsy (FNAB): By aspirating cells from the thyroid nodule with a fine needle for pathological examination, the benign or malignant nature of the nodule can be determined. Imaging examinations: Include CT and MRI, which can provide detailed images of the thyroid and surrounding tissues to help doctors evaluate the extent of thyroid cancer, whether there is lymph node metastasis, etc. Laboratory diagnosis: Thyroid function, Tg, and thyroid antibody tests should be performed before surgery and used as the baseline assessment for dynamic monitoring.

[0005] All appeal testing methods have certain limitations. For example, in tumor marker detection: Thyroglobulin (Tg) is a specific protein produced by the thyroid gland. However, the measurement of serum Tg lacks specific value in differentiating benign and malignant thyroid nodules, so it is not recommended for preoperative diagnosis. In addition, the presence of Tg antibodies may interfere with the measurement of Tg, resulting in false-negative or false-positive results. In imaging examinations: Color Doppler ultrasound is the preferred method for diagnosing thyroid cancer, but it also has limitations. For example, the evaluation depends on the operator, and it cannot fully image deep anatomical structures and those structures affected by bone or air acoustic shadows. CT and MRI may be superior to ultrasound in terms of the sensitivity of examining central and lateral cervical lymph nodes, but in predicting extrathyroidal tumor extension and multifocal bilateral lobe lesions, color Doppler ultrasound is more accurate. PET-CT is not used as a routine examination method in the diagnosis of thyroid cancer, and PET cannot fully distinguish inflammatory lymph nodes from metastatic lymph nodes, reducing the specificity. Gene testing plays an important role in the diagnosis of thyroid cancer, but multiple factors need to be considered when selecting a suitable gene testing panel, including the purpose of clinical application and objective conditions. In addition, gene testing methods and variant interpretation need to be carried out in a laboratory that complies with regulations to ensure the accuracy of the results. Limitations of pathological examinations: Fine needle aspiration biopsy is a commonly used diagnostic method for thyroid nodules. Fine Needle Aspiration Biopsy (FNAB) is a commonly used method in the diagnosis of thyroid nodules, but it has some limitations, mainly including: the influence of the operator's experience, the quality of specimen collection, the limitations of cytology itself, the false-negative rate, diagnostic uncertainty, patient acceptance, and follow-up compliance. Therefore, there is an urgent need for a method that is more convenient and sensitive and can track the development of thyroid cancer. There is an urgent need to find a convenient, inexpensive method that has high sensitivity and specificity for the diagnosis, tracking of thyroid cancer. Summary of the Invention

[0006] An object of the present invention is to provide a peripheral blood TCR marker for thyroid cancer, a method for obtaining the same, and an application thereof.

[0007] Application of a protein for predicting the risk of thyroid cancer in the preparation of a reagent for thyroid detection, characterized in that: the protein marker is a thyroid cancer tumor-specific TCR sequence, which may include n sequences, and the marker includes at least one of the proteins shown in SEQ ID NOs. 1 to 50. The kit contains a reagent for detecting the protein marker related to predicting the risk of thyroid cancer in a subject sample, and the subject sample is the blood of a patient. The detection reagent is a reagent for quantitatively detecting the expression level of at least one of the proteins shown in SEQ ID NOs. 1 to 50. The steps for detecting the sequence are as follows: sampling and sample preservation, blood collection and anticoagulation treatment: ensure that the blood sample is freshly collected and anticoagulated using an EDTA anticoagulation tube, then perform RNA extraction and then perform immunosequencing. Immediately perform 5′RACE RT-PCR on the isolated RNA. Mix 1 μg of total RNA with a cDNA synthesis primer mixture, 10 μM each, incubate at 70°C for 2 min, then lower the incubation temperature to 4°C to anneal the oligo dT primer. After incubation, add a mixture containing 5× first-strand buffer, 20 mM DTT, 5′ template-switching oligonucleotide at a concentration of 10 μM, dNTP solution, and reverse transcriptase to the total RNA reaction with annealed primers, and incubate at 42°C for 60 min. During reverse transcription, unique molecular identifiers are incorporated to minimize amplification bias and allow for quantitative comparison of gene expression levels. Purify using the AMPure Size Select magnetic bead kit at a 1× ratio (Beckman Coulter) and elute with 20 μL of water. Subsequently, all purified products are used as templates for a 50 μL PCR amplification system, which contains 2× Q5 high-fidelity master mix (NEB), dNTPs (10 mM each), forward universal primer (10 μM), and reverse primer mixture (0.2 μM each in the heavy chain mixture or 0.2 μM in TCRB). Perform thermal cycling under the following conditions: initial denaturation at 98°C for 1 min 30 s, followed by 18 cycles with denaturation at 98°C for 10 s, annealing at 60°C for 20 s, extension at 72°C for 40 s, and finally extension at 72°C for 4 min. Next, use 1 μL of the first-round PCR product for a second-round 24-cycle PCR amplification, including 2× Q5 high-fidelity master mix, dNTPs (10 mM each), and P5 and P7 sequencing adapter primers, with the same PCR conditions as above. For the second PCR reaction, purify using the AMPure Size Select Magnetic Bead Kit at a 0.8× Beckman Coulter ratio.Fastq files generated by Illumina Nova-seq 6000 with paired-end 150 reads were processed using MIXCR 4.6.0 (39) to align each TCR, known V and J regions, and identify the CDR3. The output of this pipeline was the identity of all unique clonotypes with V, the amino acid sequence of CDR3, read counts, and frequencies. Data analysis and screening of thyroid cancer-specific TCR β-chain CDR3 sequences involved assembling and annotating the sequences of T cell receptors (TCRs) using the MIXCR software to determine the V and J gene regions as well as the CDR3 sequences. This process output all unique clonotypes containing V gene, CDR3 amino acid sequence, read counts, and frequencies. We calculated the relative proportions of V-J gene usage in each sample and compared only the V-J combinations that were detectable in all samples. Statistical analysis was performed using unpaired Student's t-test, and subsequent FDR adjustment of the P-values was carried out using the Benjamini-Hochberg method at 5%. Unpaired Student's t-test or Welch's t-test was selected as needed. Using machine learning algorithms, a scoring system was constructed, and a score above 80 was considered indicative of thyroid cancer. The scoring principle was based on the clones detected in new patients, and scores were assigned according to the overlap and similarity between this clone and the clones specific to thyroid cancer and the clones specific to nodules. Positive scores were given for overlap or similarity with thyroid cancer, and negative scores were given for overlap or similarity with nodules. A cumulative score exceeding 80 was considered indicative of thyroid cancer. The primer sequences were the strand displacement primer BIOTIN-ACACGACGCTCTTCCGATCT NNNNNNNN LNAGLNAGLNAG, the TCR-beta upstream primer: ACACTCTTTCCCTACACGACGCTCTTCCGATCT,. The TCR-beta downstream primer: GCACCTCCTTCCCATTCAC, and the adapter primer was the P5 primer: AATGATACGGCGACCACCGAGATCTACAC[I5]TCGTCGGCAGCGTCAGATGTGTATAAGAGACAG; the P7 primer CTGTCTCTTATACACATCTCCGAGCCCACGAGAC[I7]ATCTCGTATGCCGTCTTCTGCTTG.

[0008] Advantages of the present invention: 1. In the present invention, first, characteristic TCR β-chain CDR3 sequences of thyroid cancer were screened and determined using the control group samples of benign thyroid nodules and the TCR high-throughput sequencing data of thyroid cancer patients. By comparing with these characteristic TCR β-chain CDR3 sequences of thyroid cancer, it is possible to clearly determine whether there are individuals with a high risk of thyroid cancer in the test samples; 2. Analyzing TCR changes through high-throughput sequencing can detect very early-stage thyroid cancer. Analyzing the response of T cells in the human immune system to thyroid cancer based on the characteristic TCR sequences of thyroid cancer is a novel detection method. 3. In the present invention, due to the adoption of high-throughput sequencing technology and the simultaneous comparison of a huge number of characteristic TCR sequences, it has higher specificity and accuracy compared to the individual detection of one or several markers. 4. The cost of the high-throughput sequencing instrument used in the present invention is lower than that of large-scale imaging equipment and can be outsourced to a third party. In addition, the labor costs for sampling and processing are lower than those for the simultaneous detection of multiple markers and also lower than those for a large number of cytological detections. Therefore, the present invention greatly reduces the detection cost. 5. The present invention only requires a small amount of peripheral blood to be taken. The sampling is simple and safe, and it is a non-invasive testing method. 6. The characteristic TCR β-chain CDR3 sequence of thyroid cancer described in the present invention can also be used for the immunotherapy of thyroid cancer. BRIEF DESCRIPTION OF THE DRAWINGS

[0009] Figure 1 is the schematic diagram of immune group sequencing; Figure 2 is the immune map (T cell clone expansion map), where: a. normal healthy people; b. patients with benign thyroid nodules; c. patients with thyroid cancer; Figure 3 is the evaluation using the ROC model, and the area under the curve is 0.77, indicating that the model has a high degree of accuracy. DETAILED DESCRIPTION OF THE INVENTION

[0010] The present invention will be further described below, but it is not limited in any way. Any transformation or replacement based on the teachings of the present invention falls within the protection scope of the present invention.

[0011] The principle of the present invention is: Cancer poses a huge threat to human health, and there is an urgent need for more effective methods of diagnosis, monitoring, and treatment. In recent years, with the successful application of a series of immune checkpoint inhibitors (ICIs) and engineered immune cells in cancer treatment, immunotherapy has become a promising and powerful tool against cancer. However, only a small fraction of people show an effective immune response to immunotherapy, indicating that the treatment effect may be influenced by certain factors, especially immune factors. Therefore, identifying cancer immunity and immune characteristics not only helps in diagnosis but also in predicting and optimizing the treatment response for tumor eradication. B lymphocytes and T lymphocytes in the human body are two important types of cells in the acquired immune system. B cells recognize antigens through the B cell receptor (BCR) on the cell surface. Later, when BCR is expressed as an antibody during the differentiation of B cells into plasma cells, it is secreted extracellularly. T cells recognize antigens through the T cell receptor (TCR) on the cell surface, mediate cell-mediated immune responses, and play a key role in effective anti-tumor immune responses as an essential component for the activation of humoral immune responses.

[0012] The diversity of BCR and TCR is the basis for establishing the acquired immune system. In oncology, the diversity of B cell receptor (BCR) and T cell receptor (TCR) is very important for tumors, mainly reflected in their impact on tumor immune surveillance, as biomarkers for predicting treatment efficacy and prognosis, and their indication of patient response in immunotherapy, etc. There is a region in BCR / TCR called the Complementary Determining Region (CDR), which consists of CDR1, CDR2, and CDR3. CDR1 and CDR2 are both encoded by the V gene, while CDR3 is encoded by three genes, V, D, and J. CDR3 has the highest degree of diversity and plays a key role in antigen recognition. Therefore, it is not difficult to understand that the higher the diversity of CDR3, the stronger the ability to recognize antigens, and the more responsive the immune system is to "foreign" molecules. Therefore, studying the gene sequence of the CDR3 region is of great significance for disease diagnosis, monitoring minimal residual disease (MRD), and immune assessment, etc.

[0013] In the human body, after cells age, their proteins will undergo a degradation process, and the generated fragments will be transported to the cell surface. These fragments are presented to T cells in the immune system through major histocompatibility complex II (MHC II). In normal cells, due to the immune tolerance mechanism, these presented antigen fragments usually do not trigger an immune response. However, when cells become cancerous, due to gene mutations leading to the production of abnormal proteins, once the fragments of these abnormal proteins are presented on the cell surface, they will trigger a specific immune response of the human immune system. Therefore, analyzing the changes in B cell receptor (BCR) or T cell receptor (TCR) helps us monitor the occurrence and development of tumors.

[0014] Example 1

[0015] Collect peripheral blood (10 mL per person) from 1 healthy person, 1 patient with benign thyroid nodules, and 1 patient with thyroid cancer. Collect and store the samples according to the following procedures, extract RNA, perform immune group sequencing, and analyze data to screen for thyroid cancer-specific TCR β-chain CDR3 sequences.

[0016] I. Sampling and Sample Preservation Blood collection and anticoagulation treatment: Ensure that the blood sample is freshly collected and use an EDTA anticoagulation tube for anticoagulation treatment.

[0017] Temporary preservation: If the anticoagulated blood sample cannot be processed immediately, you can store it at 4°C, but the storage time should not exceed 48 hours.

[0018] Preparation before processing: Before processing the blood stored at 4°C, invert the sample 20 times to avoid the influence of white blood cell sedimentation on RNA extraction.

[0019] Blood sample storage: Use a pipette or syringe to aspirate 3 ml of blood from the anticoagulated blood and add it to the sample preservation solution. Be sure to shake well to ensure mixing, and then store the sample under the appropriate conditions. We recommend long-term storage at -80°C.

[0020] II. RNA Extraction Preparation before the experiment 1. Sample preparation: Take out the blood samples to be extracted (48) from the -80°C refrigerator and melt them overnight at -20°C.

[0021] 2. Prepare the anhydrous ethanol plate Plate preparation: Use a 24 / 48-well plate as the storage plate and pre-fill it with anhydrous ethanol (1 ml / well).

[0022] Pour the anhydrous ethanol into the dedicated pipetting trough for ethanol. Use a 1 ml pipette gun, adjust it to 1 ml / gun, and add anhydrous ethanol to the 24 / 48-well plate.

[0023] Seal the plate: Seal the plate with an aluminum foil sealing film and store it at -20°C.

[0024] After sealing the plate with an aluminum foil sealing film, place it in the -20°C refrigerator and press a heavy object on the plate to reduce ethanol evaporation (it can be stored for 1 - 2 weeks).

[0025] 3. Prepare the washing solution Operation: Dilute the 4× washing solution stock solution to 1× washing solution.

[0026] For example, to prepare 200 ml of 1× washing solution (50 ml of stock solution + 150 ml of absolute ethanol): Add 50 ml of 4× washing solution to a 50 ml washing solution tube (on the test tube rack on the side of the table lamp), and pour it into the 1× washing solution bottle; Add 50 ml of absolute ethanol to a 50 ml washing solution tube (on the test tube rack on the side of the table lamp), and pour it into the 1× washing solution bottle; Repeat the second step twice, adding a total of 150 ml of absolute ethanol; It can be used only after mixing evenly.

[0027] 4. Preparation of purification-related consumables (during the first centrifugation): Purification column number: The same as the extraction sample number (i.e., "date + number") Receiving tube: No numbering is required 1.5 ml EP tube: The number is the same as the sample number, and both the side wall and the tube cap are numbered (i.e., "date + number").

[0028] Experimental operations 1. Numbering: The number of the RNA extraction tube should correspond to the blood sample number, and then take a photo for record.

[0029] 2. Pretreatment: For the mixture of 1 ml of blood and 2 ml of preservation solution, add 100 ul of BCP, mix well by shaking at room temperature (2700) for 10 min, and then centrifuge.

[0030] 3. Centrifugation at room temperature: Centrifugation conditions: For the 401 large centrifuge: JA-14 rotor, 9000 rpm, room temperature, 10 min.

[0031] 4. Arrange the samples to conform to the loading order of the absolute ethanol plate.

[0032] 5. Use a 1 ml pipette to aspirate the supernatant (about 1.5 ml in Scheme 1, mixed with 1 ml of absolute ethanol; about 2.5 ml in Scheme 2, mixed with 1.5 ml of absolute ethanol) into the corresponding wells of the absolute ethanol plate, and directly discard the remaining samples.

[0033] Remark 1: The amount of supernatant aspiration can be reduced to ensure that no precipitate is aspirated! Remark 2: The supernatant should be clear and transparent. If a turbid supernatant is encountered, please report it in time! 6. Connect the numbered purification column to the negative pressure suction filter and turn on the negative pressure suction filter.

[0034] 7. Use a 1 ml pipette to add all the mixed solutions in the absolute ethanol plate to the 24 purification columns respectively.

[0035] 8. After the stepwise filtration is completed, close the valve, use a wash bottle to add 1 ml of cleaning solution per well (no need for accurate quantification, just fill it up), and then open the valve for filtration. This process is regarded as one wash.

[0036] 9. Repeat the washing 4 times.

[0037] 10. After the filtration is completed, insert the waste liquid receiving tube, and centrifuge at 12000 rpm for 1 min in a small centrifuge.

[0038] 11. Discard the waste liquid receiving tube, insert the corresponding 1.5 ml EP tube, and use a continuous pipette*, add 50 μl of eluent to each tube in the purification column, and let it stand at room temperature for 3 min.

[0039] *Continuous pipette: The largest pipette tip (12.5 ml), confirm that it is adjusted to 50 μl / drop; *If using a continuous pipette, it is necessary to confirm that 50 μl of eluent has been dropped into each tube. Finally, check again. If not dropped, repeat this step.

[0040] 12. Centrifuge at 12000 rpm for 1 min in a small centrifuge. Discard the purification column and collect the flow-through in the 1.5 ml EP tube.

[0041] 13. Add 1 μl of RNase inhibitor to every 50 μl of RNA flow-through (1.5 ml EP tube).

[0042] 14. Cover the tube cap and place it in a 4°C refrigerator.

[0043] 15. Detect the RNA concentration with a nanodrop.

[0044] 16. Store it frozen in an -80°C refrigerator.

[0045] III. Immune group sequencing For the batch TCR library, immediately perform 5′RACE RT-PCR on the isolated RNA. Focusing on the T cell receptor beta chain (TCR B), we made the following improvements: 1. Mix 1 μg of total RNA with a cDNA synthesis primer mixture (10 μM each), incubate at 70 °C for 2 min, then lower the incubation temperature to 4 °C to anneal the oligo dT primer. After incubation, add a mixture containing 5× First Strand Buffer (TAKARA), 20 mM DTT, 5′ Template Switching Oligonucleotide (10 μM), dNTP solution (10 mM each), and 10× SMARTScribe Reverse Transcriptase (TAKARA) to the total RNA reaction with annealed primers, and incubate at 42 °C for 60 min. During reverse transcription, unique molecular identifiers (UMIs) are incorporated to minimize amplification bias and allow quantitative comparison of gene expression levels.

[0046] 2. First strand synthesis reaction Purify using the AMPure Size Select magnetic bead kit at a 1× ratio (Beckman Coulter) and elute with 20 μL of water. Subsequently, all purified products are used as templates for a 50 μL PCR amplification system that contains 2× Q5 High-Fidelity Master Mix (NEB), dNTPs (10 mM each), forward universal primer (10 μM), and reverse primer mixture (0.2 μM each in the heavy chain mixture or 0.2 μM in TCRB). Perform thermal cycling under the following conditions: Initial denaturation at 98 °C for 1 min 30 s, Followed by 18 cycles with denaturation at 98 °C for 10 s, annealing at 60 °C for 20 s, extension at 72 °C for 40 s, and a final extension at 72 °C for 4 min.

[0047] Next, use 1 μL of the first round PCR product for a second round of 24-cycle PCR amplification, including 2× Q5 High-Fidelity Master Mix (NEB), dNTPs (10 mM each), and P5 and P7 sequencing adapter primers, with the same PCR conditions as above.

[0048] 3. Second PCR reaction Purify using the AMPure Size Select Magnetic Bead Kit at a 0.8× ratio (Beckman Coulter). Fastq files generated by Illumina Nova-seq 6000 with paired-end 150 reads MIXCR 4.6.0 (39) are used to assemble and annotate the reads to align and identify each TCR, known V and J regions, and CDR3. The output of this pipeline is the identity of all unique clonotypes with V, the amino acid sequence of CDR3, read counts, and frequencies.

[0049] IV. Data analysis and screening of thyroid cancer-specific TCR β-chain CDR3 sequences The sequences of the T cell receptor (TCR) were assembled and annotated using the MIXCR software to determine the V and J gene regions as well as the CDR3 sequences. This process output all unique clonotypes containing the V gene, CDR3 amino acid sequence, read counts, and frequencies. Then, we calculated the relative proportions of V-J gene usage in each sample and only compared the V-J combinations that could be detected in all samples. Statistical analysis was performed using an unpaired Student's t-test, and the P-values were subsequently adjusted for 5% FDR using the Benjamini-Hochberg method. Statistical comparisons were carried out using the rstatix package in R, and unpaired Student's t-test or Welch's t-test was selected as needed.

[0050] Primers used for amplification: Reverse transcription primer: TTTTTTTTTTTTTTTTTTTTTTTTTTTTTTTVN Strand displacement primer: BIOTIN- ACACGACGCTCTTCCGATCT NNNNNNNN LNAGLNAGLNAG TCR-beta upstream primer: ACACTCTTTCCCTACACGACGCTCTTCCGATCT TCR-beta downstream primer: GCACCTCCTTCCCATTCAC The sequencing adapter primer is the P5 primer: AATGATACGGCGACCACCGAGATCTACAC[I5]TCGTCGGCAGCGTCAGATGTGTATAAGAGACAG; P7 primer CTGTCTCTTATACACATCTCCGAGCCCACGAGAC[I7]ATCTCGTATGCCGTCTTCTGCTTG.

[0051] Example 2

[0052] In this study, peripheral blood samples were collected from 11 patients with benign thyroid nodules larger than 4 cm, 19 patients with early PTC, and 12 healthy adults. All participants were aged 40 - 55 years and had no history of chronic diseases, infectious diseases, or autoimmune diseases. Sample collection was approved by the hospital ethics committee (number: 202402301), and informed consent was obtained from all participants. The entire experimental method was carried out according to the protocol of Example 1.

[0053] The present invention proposes a thyroid cancer diagnosis scoring system based on the Multiple Instance Learning (MIL) framework. The core of this method lies in calculating the overlap and similarity between the clone features detected in new patients and the known thyroid cancer-specific clones and nodule-specific clones, and scoring accordingly.

[0054] The specific method is as follows: 1. Data preprocessing: Collect the clone feature data of thyroid cancer patients and benign nodule patients, and use them as the reference data sets for thyroid cancer-specific clones and nodule-specific clones respectively.

[0055] 2. Construction of the multiple instance learning model: Regard the clone features as "instances" in the multiple instance learning framework, and regard the patient samples as "bags". Use the labeled thyroid cancer and nodule data to train the model so that the model can learn the characteristic patterns of thyroid cancer and nodule clones.

[0056] 3. Scoring mechanism: For the clone features of new patients, the model calculates their overlap and similarity with thyroid cancer-specific clones. The thyroid cancer-specific clone features are at least one of the proteins shown in SEQ ID NO.1-50 or the RNA sequence that can be translated into this protein, and a positive score is given; calculate their overlap and similarity with nodule-specific clones. The nodule-specific clone features are at least one of the proteins shown in SEQ ID NO.51-98 or the RNA sequence that can be translated into this protein, and a negative score is given. Through the accumulation of positive and negative scores, the total score of the patient is obtained.

[0057] 4. Diagnostic judgment: When the cumulative score exceeds 80 points, it is judged that the patient has thyroid cancer; otherwise, it is judged as non-thyroid cancer.

[0058] 5. The protein sequences of SEQ ID NO.1-50 or the RNA sequences that can be translated into this protein sequence. The RNA sequences are corresponding according to the universal codons of eukaryotic cells. SEQ ID NO.51-98 is for thyroid face type hyperplasia, and its detection does not represent thyroid cancer.

[0059] Through the multiple instance learning framework, the present invention combines the calculation of the overlap and similarity of clone features, realizes the efficient differentiation between thyroid cancer and benign nodules, and has high accuracy and clinical application value. The results are shown in the appendix Figure 3 .

Claims

1. Use of a protein for predicting the risk of thyroid cancer in preparing a reagent for thyroid testing, characterized in that: The protein or RNA marker is a thyroid cancer tumor-specific TCR sequence, which includes a total of 50 sequences. The marker set is one or several of the proteins shown in SEQ ID NO.1 to 50 or an RNA sequence that can be translated into the protein. The kit contains reagents for detecting protein markers or RNA markers in subject samples and predicting the risk of thyroid cancer. The subject sample is the patient's blood.

2. The use according to claim 1, characterized in that: The detection reagent is a reagent for quantitatively detecting the expression amount of at least one of the proteins or RNAs with sequences shown in SEQ ID NO. 1 to 50, or a reagent for quantitatively detecting an RNA sequence that can be translated into the protein.

3. The use according to claim 1, characterized in that: The steps and reagents of the detection sequence are as follows: sampling and sample preservation, blood collection and anticoagulation: Please ensure that the blood sample is freshly collected and anticoagulated using EDTA anticoagulation tubes, then RNA extraction and then immunohistochemistry sequencing, the isolated RNA is immediately subjected to 5′RACE RT-PCR. The total RNA and cDNA synthesis primers are used as a reverse transcription primer mixture, each 10 μM, incubated at 70°C for 2 minutes, and then the incubation temperature is reduced to 4°C to anneal the oligo dT primer. After incubation, a mixture containing 5× first strand buffer, 20 mM DTT, 5′ template exchange oligonucleotide, i.e., a strand displacement primer concentration of 10 μM, dNTP solution and reverse transcriptase is added to the primer annealing total RNA reaction, and incubated at 42°C for 60 minutes. During the reverse transcription process, a unique molecular identifier is incorporated to minimize amplification bias and allow quantitative comparison of gene expression levels. Purification is performed using the AMPure SizeSelect magnetic bead kit at a 1× ratio and eluted with 20 μL of water. Subsequently, all purified products were used as templates for PCR amplification system containing 2×Q5 high-fidelity master mix, dNTPs, forward universal primer (TCR-beta upstream primer) and reverse primer mix (TCR-beta downstream primer), each 0.2 μM in heavy chain mix or 0.2 μM in TCRB. Thermal cycling was performed under the following conditions: initial denaturation at 98°C for 1 min 30 s, followed by 18 cycles of denaturation at 98°C for 10 s, annealing at 60°C for 20 s, extension at 72°C for 40 s, and final extension at 72°C for 4 min. Next, 1 μL of the first-round PCR product was used for a second-round PCR amplification of 24 cycles, including 2×Q5 high-fidelity master mix, dNTPs, each 10 mM, and P5 and P7 sequencing adapter primers containing P5 primer; P7 primer, the same as the above PCR conditions. The second PCR reaction was purified using the AMPure SizeSelect Magnetic Bead Kit at a ratio of 0.8× Beckman Coulter. Fastq files generated by Illumina Nova-seq 6000 with paired-end 150 reads MIXCR4.6.0 (39) were used to align the identities of each TCR, known V and J regions, and CDR3. The output of the pipeline was all unique clonotypes with V identities, CDR3 amino acid sequences, read counts, and frequencies. Data analysis and screening of thyroid cancer-specific TCR β chain CDR3 sequences The sequences of T cell receptors (TCRs) were assembled and annotated using MIXCR software to identify V and J gene regions and CDR3 sequences. The pipeline outputs all unique clonotypes containing V genes, CDR3 amino acid sequences, read counts, and frequencies.We calculated the relative proportion of VJ gene usage in each sample and compared only the VJ combinations that could be detected in all samples. The unpaired Student's t test was used for statistical analysis, and the P value was subsequently adjusted for 5% FDR by the Benjamini-Hochberg method. Unpaired Student's t test or Welch's t test was selected as needed. The clone detected in the new patient was scored according to the overlap and similarity between this clone and the unique clone in thyroid cancer and the unique clone in the nodule. The overlap or similarity with thyroid cancer was scored positively, and the overlap or similarity with the nodule was scored negatively. If the cumulative score exceeded 80 points, it was judged as thyroid cancer.

4. The use according to claim 3, characterized in that: The kit includes the following primer combinations: strand displacement primer BIOTIN- ACACGACGCTCTTCCGATCT NNNNNNNN LNAGLNAGLNAG, TCR-beta upstream primer: ACACTCTTTCCCTACACGACGCTCTTCCGATCT, TCR-beta downstream primer: GCACCTCCTTCCCATTCAC, sequence adapter primer is P5 primer: AATGATACGGCGACCACCGAGATCTACAC[I5]TCGTCGGCAGCGTCAGATGTGTATAAGAGACAG; P7 primer CTGTCTCTTATACACATCTCCGAGCCCACGAGAC[I7]ATCTCGTATGCCGTCTTCTGCTTG.

Citation Information

Patent Citations

  • Peripheral blood TCR marker for prostate cancer, and detection kit and application thereof

    CN111679074A

  • Construction method of single cell transcriptome sequencing library and applications thereof

    CN113444770A

  • Composition, kit and sequencing method for implementing immune repertoire sequencing

    CN114934096A

  • Glypican-3-specific t-cell receptors and their uses for immunotherapy of hepatocellular carcinoma

    WO2015173112A1