Gene marker combinations for assessing risk of hlh and uses thereof

By screening 155 SNP mutation sites through whole exome sequencing, a combination of gene markers was identified, which solved the problem of low coverage of gene testing packages in HLH diagnosis and enabled more accurate HLH risk assessment and early identification.

CN118086488BActive Publication Date: 2026-06-02GUANGZHOU KINGMED TRANSFORMATIVE MEDICINE INST CO LTD +2

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
GUANGZHOU KINGMED TRANSFORMATIVE MEDICINE INST CO LTD
Filing Date
2024-03-27
Publication Date
2026-06-02

AI Technical Summary

Technical Problem

Current HLH diagnosis suffers from low coverage of gene testing packages, making it difficult to fully reflect the polygenic heterogeneity of the disease, leading to diagnostic difficulties and poor prognosis.

Method used

Genome-wide association analysis using whole-exome sequencing was employed to screen for a combination of gene biomarkers consisting of 155 SNP mutation sites. The risk of HLH was assessed by calculating a polygenic risk score (PRS).

Benefits of technology

It improves the sensitivity and specificity of HLH gene testing, provides cost-effective testing products, reduces the economic burden, and supports the early identification of HLH patients.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118086488B_ABST
    Figure CN118086488B_ABST
Patent Text Reader

Abstract

The application discloses a gene marker combination for evaluating HLH risk and use thereof, and belongs to the technical field of gene detection. The gene marker combination is obtained by whole genome association analysis of whole exome sequencing, covers more extensive genetic information, solves the narrowness of the prior art, and provides more comprehensive analysis of the polygenic heterogeneity of HLH. Moreover, the cumulative effect of alleles is comprehensively considered, the genetic susceptibility characteristics of an individual to HLH can be more accurately reflected, and the sensitivity and specificity of gene detection can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of gene detection technology, and in particular to a combination of gene biomarkers for assessing HLH risk and their uses. Background Technology

[0002] Hemophagocytic lymphohistiocytosis (HLH), also known as hemophagocytic syndrome (HPS), is an excessive inflammatory response syndrome caused by abnormal activation, proliferation, and secretion of large amounts of inflammatory cytokines by lymphocytes, monocytes, and macrophages due to hereditary or acquired abnormal immune regulation. It can easily lead to severe tissue inflammation and end-stage multi-organ damage.

[0003] The clinical phenotype of HLH is broad and nonspecific, making it difficult to differentiate from diseases such as severe sepsis and persistent systemic inflammatory response syndrome (SIRS). Furthermore, the disease progresses rapidly and has a high mortality rate. Therefore, early identification of high-risk individuals for HLH is crucial for developing personalized treatment plans and reducing mortality.

[0004] Currently, the diagnostic criteria for HLH are based on the HLH-2004 diagnostic guidelines developed by the International Histiocytic Society, which clearly state that gene defects are the gold standard for HLH diagnosis. Therefore, gene sequencing testing remains the primary method for HLH diagnosis. With a deeper understanding and more extensive research into the disease, the number of diagnostic genes listed in the HLH-2004 guidelines has gradually increased. Currently, in most medical institutions, considering factors such as the time and cost of gene sequencing testing and the rapid progression of the disease, targeted diagnostic gene sequencing is the first choice for the clinical differential diagnosis of HLH.

[0005] However, mounting research evidence suggests that HLH is a highly heterogeneous polygenic disease with multiple potential molecular pathways leading to it, including primary immunodeficiency, dysregulated immune activation, and proliferative disorders. Therefore, a large number of unknown genes still exist that increase the risk of HLH. Currently, HLH genetic testing packages available both domestically and internationally typically only include a few known pathogenic genes, with a clinical diagnostic gene detection rate of less than 20%. This makes HLH diagnosis difficult and significantly impacts patient prognosis. Summary of the Invention

[0006] To address the aforementioned difficulties in the clinical diagnosis of HLH, this invention provides a combination of gene markers for assessing HLH risk. By detecting this combination of gene markers, the genetic risk of HLH can be predicted, which helps to promote the early identification of patients' HLH risk and provides certain auxiliary judgment basis for early clinical diagnosis of HLH.

[0007] This invention discloses a combination of gene biomarkers for assessing the risk of HLH, the combination of gene biomarkers comprising SNP mutations in the following genes: AARD, ACOT12, ADGRF5, DPM1, AGAP7P, ANKMY1, ANKRD36, ANKRD36B, ANKRD36C, APOBR, ARFGEF2, ARHGAP18, ATAD3B, ATL1, ATMIN, AVL9, BAIAP2L1, BCLAF1, BDKRB2, C1S, C2, C2CD2L, CBX3, CD82, CDC27, CISD2, CLCN1, CNGB1, COL4A3, CSNK1G2, CTAG E15, CYP27A1, CYP2B6, DENND5A, DPP6, DRICH1, DTNBP1, EGFR, EP300, EPHA1, EVA1C, FAHD2B, FAT2, FCGBP, FCGR2B, FOLH1, G3BP2, GCSH, GLIPR1L2, GLRX 5. GON4L, GPR84, GYPA, GYPB, GZMM, HARBI1, HDAC3, HEATR1, HLA-C, HLA-DQA1, HMCN1, HNRNPA3, HNRNPCL2, HRNR, IGSF3, ING1, INKA2, ITFG1, ITGA4, KCNJ 12. KIF2A, LCE3D, LMO3, AP1G2, LRRIQ1, MCPH1, MDC1, MICAL2, MIER1, MIR205HG, MME, MORN1, MRPS25, MSH2, MST1, MUC12, MUC16, MUC19, MUC3A, MYO5B, N ACAD, NATD1, NBPF3, NECTIN3, NEK11, NFATC1, NPIPB15, NRDE2, NUTM2B, OSBPL10, PABPC1, PAPOLA, PDE6B, PDS5A, PIK3R1, PLA2G2D, PLAGL1, PLD3, POLR1 E, POT1, PPP5C, PRAMEF13, PRSS3, PRSS57, PTPRU, PUM1, RGPD1, RGPD3, RNVU1-18, SAMSN1, SERGEF, SLC35G5, SLC35G6, SLC38A11, SLC9B1, SLFN13, SPAT A1, SPATS2, SPRR3, SULT1A2, SVIL, TAF4, TASOR2, ​​TIMM23, TMOD4, TRANK1, TUBA3D, UBN1, UNC13D, VIPAS39, VWF, ZADH2, ZFP62, ZNF70, ZNF705A and ZSWIM6.

[0008] In previous studies, the inventors discovered that multiple mutations in the same or different pathways can lead to an increased risk of HLH in an individual. This suggests that the cumulative effect of alleles is a genetic susceptibility trait of an individual to HLH. Among these alleles, those with larger gene effects can be called major genes, and those with smaller gene effects can be called minor genes.

[0009] Furthermore, based on the theoretical basis that the cumulative effect of the aforementioned polygenic mutations leads to an increased risk of HLH, the inventors analyzed the differences in genetic background between the control group and HLH patients, calculated the risk effect value of genetic loci and the polygenic risk score of the samples, and screened out 155 risk loci. This set of high-risk loci is distributed in the aforementioned gene combinations, and these gene markers and risk loci can be combined to form a gene panel. Applying this panel helps to promote the early identification of HLH risk in patients and provides a certain basis for clinical early HLH diagnosis.

[0010] In some of these schemes, the SNP mutations include the following SNP sites located based on hg19 in the reference genome: 1:1407372, 1:13183567, 1:13448283, 1:20442054, 1:21808159, 1:29650289, 1:31423116, 1:67451602, 1:85018771, 1:112269872, 1:117158972, 1:149224578, 1:151142978, 1:152191578, 1:152552279, 1:152975763, 1:155823124, 1:161643889, 1:18 5894153, 2:219679800, 2:228121101, 2:241451350, 3:15092233, 3:31704928, 3:36880027, 3:49723274, 3:110912166, 3:130992424, 3:154890037, 4 :664028, 4:39864448, 4:76584154, 4:103822298, 4:144922436, 4:145049120, 5:61648394, 5:67522851, 7:124486980, 7:143043240, 7:143105804, 7 :143270220, 7:153750014, 8:6484505, 8:11189529, 8:101721812, 8:117954790, 9:33798440, 9:37501844, 10:5766389, 10:29779789, 10:51465552, 10:51623061, 10:81472483, 11:9192336, 11:12263714, 11:18034592, 11:44640252, 16:4908667, 16:28506399, 16:28603585, 16:47292699, 16:5796 5793, 16:74422332, 16:81069449, 16:81129887, 17:7386114, 17:7386280, 17:21156530, 17:21319868, 17:33769034, 17:45219332, 17:45234710, 17 :73827216, 18:47364009, 18:72913590, 7:97937258, 4:103808519, 1:236730215, 1:209602719, 2:47698132, 2:87169316, 2:96521928, 2:96521944,2:96593000, 2:96601312, 2:96626292, 2:97749576, 2:97829707, 2:97830053, 2:97833399, 2:97879162, 2:98129690, 2:107051270, 2:132237927, 2:165802112, 2:178080396, 2:182344846, 5:80655709, 5:141004743, 5:150891902, 5:180278369, 6: 15593316, 6:30671206, 6:31239455, 6:31903786, 6:32605366, 6:46822406, 6:129955151, 6:136590640, 6:144263089, 7:32602381, 7:45123210, 7:55238535, 7:100551122, 7:100639419, 11:46626256, 11:49197416, 11:118978411, 12:6172183, 12:7 175872, 12:8327888, 12:16713487, 12:40879698, 12:49883204, 12:54757149, 12:75804273, 12:85449874, 13:111365593, 14:24030430, 14:51062357, 14:77901758, 14:90782967, 14:96001383, 14:96707777, 14:96998645, 19:685781, 19:1975435, 1 9:9006437, 19:40370296, 19:40883844, 19:41518245, 19:46857021, 20:47615074, 20:49551780, 20:60640030, 21:15889228, 21:33785343, 22:23959752, 22:24087319, 22:41547944, 18:77229409, 19:549091, 1:2322188, 5:60628574, and 7:26246130.

[0011] In some of these schemes, the SNP mutations include the following SNP mutation sites based on hg19 localization in the reference genome: 1:1407372:T:C, 1:13183567:C:T, 1:13448283:A:G, 1:20442054:T:C, 1:21808159:G:A, 1:29650289:C:G, 1:31423116:T:A, 1:67451602:A:C, 1:85018771:G:GA, 1:112269872:G:A, 1:117158972:A:G, 1:149224578:CT:C, 1:151142978:T:C, 1:152 191578:G:A, 1:152552279:C:A, 1:152975763:C:A, 1:155823124:G:A, 1: 161643889:G:T, 1:185894153:GT:G, 2:219679800:G:A, 2:228121101:G:T , 2:241451350:C:T, 3:15092233:C:T, 3:31704928:TCA:T, 3:36880027:C: T, 3:49723274:G:A, 3:110912166:ACAT:A, 3:130992424:T:C, 3:15489003 7:C:A, 4:664028:ATT:AT, 4:39864448:G:A, 4:76584154:C:T, 4:1038222 98:C:T, 4:144922436:T:G, 4:145049120:G:A, 5:61648394:A:AT, 5:67522 851:A:C, 7:124486980:T:C, 7:143043240:C:T, 7:143105804:C:G, 7:1432 70220:C:T, 7:153750014:G:A, 8:6484505:C:CA, 8:11189529:T:C, 8:1017 21812:G:A, 8:117954790:C:G, 9:33798440:C:A, 9:37501844:G:A, 10:57 66389:A:G, 10:29779789:C:T, 10:51465552:A:T, 10:51623061:G:A, 10:8 1472483:G:T, 11:9192336:T:TAC, 11:12263714:G:T, 11:18034592:G:C, 1 1:44640252:C:T, 16:4908667:A:G, 16:28506399:T:C, 16:28603585:T:A,16:47292699:TA:T,16:57965793:G:T,16:74422332:A:G,16:81069449:A:G,16:81129887:C:A,17:7386114:A:G,17:7386280:G:A,17:21156530:G:GC,17:21319868:G:T,17:33769034:G:A,17:45219332:T:G,17:45234710:A:C,17:73827216:C:T,18:47364009:G:A,18:72913590:C:T,7:97937258:T:C,4:103808519:C:T,1:236730215:T:C,1:209602719:C:T,2:47698132:A:G,2:87169316:T:G,2:96521928:C:T,2:96521944:C:A,2:96593000:A:G,2:96601312:G:A,2:96626292:C:T,2:97749576:C:G,2:97829707:C:T,2:97830053:A:G,2:97833399:C:T,2:97879162:G:C,2:98129690:T:C,2:107051270:T:C,2:132237927:C:A,2:165802112:G:A,2:178080396:T:A,2:182344846:G:GT,5:80655709:T:C,5:141004743:T:C,5:150891902:T:A,5:180278369:T:C,6:15593316:T:C,6:30671206:T:C,6:31239455:T:C,6:31903786:C:G,6:32605366:T:C,6:46822406:A:G,6:129955151:GA:G,6:136590640:A:C,6:144263089:C:A,7:32602381:A:AT,7:45123210:G:A,7:55238535:CT:C,7:100551122:T:G,7:100639419:G:A,11:46626256:T:C,11:49197416:A:G,11:118978411:G:C,12:6172183:C:G,12:7175872:T:C,12:8327888:A:G,12:16713487:A:AT,12:40879698:C:T,12:49883204:G:T, 12:54757149:G:T, 12:75804273:G:A, 12:85449874:G:A, 13:111365593:T:TG, 14:24030430:G:A, 14:51062357:G: A, 14:77901758:G:GAA, 14:90782967:G:T, 14:96001383:A:C, 14:96707777:G:A, 14:96998645:T:G, 19:685781:G:A, 19:1975435:C:T, 19:9006437:G:A, 19:40370296:G:C, 19:40883844:A:C, 19:41518245:G:C, 19:46857021:C:T, 20:47615074:TA:T, 20:49551780:T:TA , 20:60640030:A:G, 21:15889228:G:T, 21:33785343:G:A, 22:23959752:G:A, 22:24087319:A:C, 22:41547944:A:G, 18:77229409:C:G, 19:549091:C:A, 1:2322188:G:GCTCTTTTTTTTTTTTTGAGACTGAGTCTTGCTCTGTCGCC, 5:60628574:CCGGCCGCAACCTCGG:C, and 7:26246130:A:ATGCTGACAATACTTG;,

[0012] In each SNP mutation site, the chromosome number and the site location are indicated before and after the first colon, respectively, and the bases before and after the third colon are indicated before and after the mutation, respectively.

[0013] The present invention also discloses the application of the above-mentioned combination of gene markers in the preparation of reagents for assisting in the assessment of HLH risk.

[0014] On the other hand, the present invention also discloses a gene detection kit for assisting in the assessment of HLH risk, comprising reagents for detecting the above-mentioned combination of gene markers.

[0015] On the other hand, the present invention also discloses a detection system for assisting in the assessment of HLH risk, comprising:

[0016] The detection device is used to detect the above-mentioned combinations of gene markers in the sample to be evaluated and obtain the gene information of each SNP locus.

[0017] The analysis device acquires the mutation status of each SNP site, calculates the polygenic risk score (PRS) for the sample using the summation method, compares this PRS with a preset threshold, and determines the risk of HLH.

[0018] An output device for outputting the above-mentioned HLH disease risk results.

[0019] In some of these schemes, the polygenic risk score (PRS) of the sample to be evaluated is calculated using the following formula:

[0020]

[0021] Where: n is the number of SNPs, and the Effect Size is... i Let be the weight value of the i-th SNP, and Allele Count be the genotype value of the SNP. When there is no mutation in the SNP, Allele Count is 0; when there is 1 mutation in the SNP, Allele Count is 1; when there are 2 mutations in the SNP, Allele Count is 2.

[0022] The weight value for each SNP is taken from the following β value ± 10% range:

[0023]

[0024]

[0025]

[0026]

[0027] In some of these schemes, n=155, that is, 155 SNPs are selected to calculate the polygenic risk score PRS.

[0028] In some of these schemes, the analysis device compares and judges according to the following method: when the polygenic risk score PRS ≥ a threshold, it is judged as high risk of HLH, wherein the threshold is 0.719798±0.07, preferably 0.719798±0.035, and more preferably 0.719798.

[0029] On the other hand, the present invention also discloses a method for assessing HLH risk for non-diagnostic and therapeutic purposes, comprising the following steps:

[0030] Genetic testing: Detect the above-mentioned combinations of genetic markers in the sample to be evaluated to obtain genetic information for each SNP locus;

[0031] Polygenic risk analysis: Obtain the mutation status of each SNP site mentioned above, calculate the polygenic risk score (PRS) of the sample using the summation method, compare the PRS with a preset threshold to determine the risk of HLH.

[0032] On the other hand, the present invention also discloses a computer device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the program, it performs the following steps: analyzing the gene information of each SNP site of the above-mentioned gene marker combination in the sample to be evaluated according to the above-mentioned HLH risk assessment method.

[0033] Based on common knowledge in the field, the above-mentioned preferred conditions can be combined arbitrarily to obtain various preferred embodiments of the present invention.

[0034] All reagents and raw materials used in this invention are commercially available, and the methods used in this invention, unless otherwise specified, are conventional methods.

[0035] The positive and progressive effects of this invention are as follows:

[0036] This invention provides a combination of gene biomarkers for assessing HLH risk. Compared to current HLH gene testing panels, this method utilizes genome-wide association analysis based on whole-exome sequencing, covering a broader range of genetic information and overcoming the limitations of existing technologies. This provides a more comprehensive analysis of the polygenic heterogeneity of HLH. Furthermore, this invention comprehensively considers the cumulative effect of alleles, enabling a more accurate reflection of an individual's genetic susceptibility to HLH, thus improving the sensitivity and specificity of gene testing.

[0037] Meanwhile, this invention employs whole-genome sequencing and genome-wide association analysis, and through reasonable technical selection, screened 155 SNPs as HLH detection panels, providing patients with cost-effective detection products, solving the problem of excessively expensive whole-genome sequencing and whole-genome sequencing, and reducing the economic burden on patients. Attached Figure Description

[0038] Figure 1 The ROC curve and AUC value are for the optimal SNP combination in Example 1.

[0039] Figure 2 The ROC curve and AUC value are shown in Example 2.

[0040] Figure 3 This is a scatter plot from Example 2.

[0041] Figure 4 The ROC curve and AUC value are shown in Example 3.

[0042] Figure 5 This is a scatter plot from Example 3. Detailed Implementation

[0043] The present invention is further illustrated below by way of embodiments, but the invention is not limited to the scope of the embodiments described herein. Experimental methods in the following embodiments that do not specify specific conditions were performed according to conventional methods and conditions, or as selected according to the product instructions.

[0044] Example 1

[0045] This embodiment analyzes whole-exome sequencing data from HLH patients and control groups (non-HLH patients) to screen for gene mutations that increase the risk of HLH.

[0046] 1. HLH sample collection and mutation data processing

[0047] We collected whole-exome sequencing data from 1990 HLH patients using Illumina NovaSeq 6000. To ensure the accuracy and reliability of the results, all WES data were reanalyzed using the following method.

[0048] (1) Raw base quality was filtered using fastp (v 0.20.0) and primer adapters were removed to obtain high-quality clean reads. Due to the rarity of the disease, the sample collection time span was large. The capture chip used for a small portion of the sequencing was based on the human hg19 reference genome. Therefore, all clean reads were aligned to the hg19 reference genome using the bwa mem method to generate SAM format result files.

[0049] (2) Select SNV and Indel as the mutation types and use the following GATK procedure for detection.

[0050] ① First, use samtools to convert the sam format result file into a bam file.

[0051] ② Use the SortSam tool to sort the entries corresponding to the same chromosome in the bam file in ascending order of coordinates.

[0052] ③ Use the MarkDuplicates tool to mark duplicate sequences in the sorted bam to remove duplicate sequences.

[0053] ④ Use BaseRecalibrator to correct the base quality in the bam file. The correction may be due to systematic errors in base quality caused by library preparation, sequencing process, flow cell manufacturing errors of the sequencer, reagents used in sequencing, etc.

[0054] ⑤ After indexing the BAM file, use HaplotypeCaller to simultaneously detect SNP and INDEL variants and generate a single-sample GVCF file.

[0055] ⑥ After merging the GVCF files of 1990 samples using CombineGVCFs, GenotypeGVCFs were used to perform joint genotyping to generate the population vcf, which is a merged vcf containing the mutation sites of 1990 HLH samples.

[0056] ⑦ Since the filtering parameters for SNP and INDELETE are different in the next step (hard filtering), they need to be filtered separately. Therefore, the SelectVariants tool is used to extract SNP and INDELETE from the merged VCF file separately, generating snp.vcf and indel.vcf files.

[0057] ⑧ To ensure the accuracy of mutations, all mutations were hard-filtered using the VariantFiltration tool. The specific filtering criteria for SNP were QD<2.0||MQ<40.0||FS>60.0||SOR>3.0||MQRankSum<-12.5||ReadPosRankSum<-8.0, and for INDEL it was QD<2.0||FS>200.0||SOR>10.0||MQRankSum<-12.5||ReadPosRankSum<-20.0. In addition, regarding depth, the minimum coverage depth for SNV was 10 layers, and the minimum coverage depth for INDEL was 20 layers.

[0058] ⑨ Use MergeVcfs to merge the filtered SNP and INdel VCF files to generate a corrected and filtered population VCF.

[0059] (3) On this basis, the artificially introduced errors are removed.

[0060] This experiment ultimately yielded 265,488 high-quality SNP loci.

[0061] 2. Determination of SNP combinations and calculation of their effect sizes

[0062] (1) Preparation of the basic dataset

[0063] Since there is currently no comprehensive genome-wide association studies (GWAS) data for hemophagocytic lymphohistiocytosis (HLH), a GWAS dataset of population genetic background needs to be established before calculating polygenic risk scores to assess the risk effect of all SNPs in the sample on HLH disease. This study randomly selected 808 HLH patient samples from the total sample as the sample set for GWAS analysis, which also serves as the basis for subsequent PRS analysis.

[0064] GWAS analysis was performed on 808 HLH samples and control samples. The analysis tool was the Firth logistic regression model of Regenie v3.1.1 software, and the Firth likelihood ratio test was used.

[0065] The first step in the analysis process is to estimate the empty model using the stacked block ridge regression method. The second step is to perform association tests, and finally obtain the GWAS summary data results file, which includes detailed information on 259,525 SNPs in 808 HLH samples, the HLH prevalence risk effect value (BETA) of each SNP, and the corresponding p-value of the association test.

[0066] (2) Target dataset

[0067] This experiment included 450 HLH patient samples as the target dataset. To ensure the stability and accuracy of the model, the target dataset did not overlap with the base dataset. Genotype and phenotype files for the target dataset were prepared: the genotype file was in PLINK binary format, and the phenotype file consisted of three columns: FID, IID, and phenotype (0 or 1). This invention uses the C+T method to construct the PRS model, first performing clumping analysis.

[0068] (3) Clumping analysis

[0069] When calculating the polygenic risk score (PRS), it is necessary to select representative SNPs in the sample population to calculate the individual's PRS score. Therefore, from the total 259,525 SNPs in the basic dataset, a subset of SNPs that are genetically distant and do not have any associations is extracted based on the linkage disequilibrium (LD) between each pair of SNPs.

[0070] This experiment uses PLINK clumping analysis, which is based on the r of LD between SNPs.2 The p-values ​​obtained from GWAS analysis are used to screen for representative SNPs (highest importance) within the LD region. Using representative SNPs yields more accurate polygenic risk scores. Clumping analysis first sorts SNPs by importance from smallest to largest based on their p-values ​​from GWAS analysis. Then, the first SNP in the sorted list is selected, and the r-value is calculated between this SNP and other SNPs in the LD region. 2 When a high correlation is detected, the less important (high p-value) SNP pair is removed. The first SNP selected in this process will always be retained. Simultaneously, the function of the gene containing the SNP, the clinical symptoms of HLH, and the currently known pathogenesis are considered comprehensively to retain the highly correlated SNP. After this, the next step is to select the next SNP after ranking by p-value and repeat this process.

[0071] The specific clumping analysis process in this patent is as follows: clumping analysis is performed on 259,525 SNPs in the GWAS summary data generated in step (1). The clumping conditions are: --clump-kb 250kb, --clump-p1 1, --clump-r2 0.1. The --clump-kb option of 250kb sets the genomic distance threshold considered in the clustering algorithm to 250kb, indicating that only SNPs with a distance of less than 250kb from the index SNP (the SNP at the center of each clump) are considered when performing clumping. If the distance between two variant sites is less than 250kb, they may be clustered in the same cluster. `--clump-p1` being 1 means that sites with p-values ​​less than 1 in the GWAS results, i.e., all variant sites, will be considered index variant sites; `--clump-r2` being 0.1 sets the minimum LD threshold considered in the clustering algorithm to 0.1, meaning that if a variant site has a p-value less than 1 compared to an index variant site, then the p-value will be less than 1. 2 A value greater than 0.1 may be classified into the cluster of the indexed variant site. LD value r 2 It is a statistic used to measure whether there is linkage disequilibrium between two SNP loci.

[0072] The calculation method is as follows: Consider two biallelic loci, A and B, with alleles A1 / A2 and B1 / B2, respectively. In a given population, these two loci can form four possible haplotypes: A1B1, A1B2, A2B1, and A2B2, with frequencies p11, p12, p21, and p22, respectively. Then r 2 It can be calculated using the following formula:

[0073] r 2=(p11p22-p12p21) 2 / [pA*(1-pA)pB(1-pB)]

[0074] Where: pA = p11 + p12 (frequency of A1 allele), pB = p11 + p21 (frequency of B1 allele), r 2 The value of r ranges from 0 to 1. If r 2 =0 indicates that the two SNP sites are completely independent and there is no linkage disequilibrium. If r 2 =1, indicating a complete linkage disequilibrium between these two SNP sites. The median value is 0. <r 2 <1 indicates partial chain imbalance.

[0075] Finally, we obtained 137,492 indexed SNPs through clumping analysis for PRS calculation.

[0076] The following example illustrates the clumping process: Suppose we have an index SNP with a GWAS p-value of 7.35E-07, which is lower than our set p-threshold of 1, so it will be selected as the starting point for the clustering index SNP.

[0077] First, PLINK examines other SNPs within a 250kb range (depending on the `--clump-kb 250` setting) around the index SNP. Let's assume five other SNPs are found within this range, labeled SNP1, SNP2, SNP3, SNP4, and SNP5. Next, PLINK calculates the least-needle ratio (LD) * r between these five SNPs and the index SNP. 2 Assume the following results:

[0078] SNP1 and the index SNP r 2 =0.8

[0079] SNP2 and the index SNP r 2 =0.05

[0080] SNP3 and index SNP r 2 =0.7

[0081] SNP4 and the index SNP r 2 =0.2

[0082] SNP5 and the index SNP r 2 =0.6

[0083] According to the `--clump-r2 0.1` setting, PLINK will include all r values ​​related to the index SNP. 2SNPs with a value ≥ 0.1 cluster together. Therefore, in this example, SNP1, SNP3, SNP4, and SNP5 will be clustered into the cluster centered on the index SNP, and SNP2's r... 2 =0.05<0.1, so it will not be included in this cluster. Ultimately, this cluster centered on the index SNP with a p-value of 7.35E-07 contains four other SNPs: SNP1, SNP3, SNP4, and SNP5. That is, the index SNP is the representative SNP of the cluster, and only the representative SNP of the cluster is used for PRS calculation.

[0084] (4) Calculation of PRS

[0085] PRSice 2.3.5 software was used to perform multigene risk score analysis on the target dataset of 450 HLH patient samples and 2432 control samples. PRSice 2 provides a variety of models and formulas for calculating PRS. This experiment used the summation method for analysis. This model obtains the PRS of each individual (sample to be evaluated) by multiplying the individual genotype by the SNP weight (beta value) and summing the results.

[0086] The calculation formula is:

[0087]

[0088] Where: n is the number of SNPs, and the Effect Size is... i Let be the weight value of the i-th SNP, and Allele Count be the genotype value of the SNP. When there is no mutation in the SNP, Allele Count is 0; when there is 1 mutation in the SNP, Allele Count is 1; and when there are 2 mutations in the SNP, Allele Count is 2.

[0089] The PRS score calculation process for each individual (sample to be evaluated) is as follows:

[0090] ① Calculate the PRS of SNPs within multiple P-value ranges.

[0091] Using the p-value of GWAS abstract statistics as a threshold, different SNP combinations were included by setting different thresholds. To reduce the low-low correlation (LD) among selected SNP combinations, a clumping method was used to retain representative SNPs with low LD while removing other SNPs that were highly correlated with these representative SNPs within the set LD threshold and physical distance. The PRS score of the corresponding SNP combination was calculated based on the representative SNPs included according to different p-values ​​and the genotype of each individual.

[0092] This experiment set 9990 P thresholds (5e-8 to 1), included SNPs ranging from 6 to 137492, and performed 9990 PRS calculations for each individual.

[0093] ② Select the best SNP that can distinguish the sample population from the control population.

[0094] Since the optimal P-value threshold was unknown before the experiment, after calculating 9990 PRS scores, regression analysis was performed on the 9990 PRS scores to find the most suitable P-value threshold, and the P-value threshold that could explain the most phenotypic variance was selected.

[0095] In this analysis, the optimal p-value threshold for GWAS summary statistics is 0.00065005, and the R-value at this threshold is... 2 The value is 0.0652116, and there are 155 SNPs that are optimal for PRS (as shown in Table 1 below).

[0096] Table 1. Optimal SNP combinations and their effect values

[0097]

[0098]

[0099]

[0100]

[0101] In each of the above SNP risk alleles, the first colon indicates the chromosome number and locus location, respectively, and the third colon indicates the pre-mutation and post-mutation bases, respectively. This SNP localization is based on the reference genome hg19.

[0102] (5) Validate the effectiveness of the optimal SNP combination and determine the optimal PRS threshold.

[0103] After calculating the optimal PRS for each individual using the best SNP combination, the performance was verified using ROC curves, and the results are as follows: Figure 1 As shown, the AUC reached 0.64. To determine the optimal PRS threshold, we used the `coords` function from the `pROC` package to calculate the cutoff point on the ROC curve that optimally combines sensitivity and specificity. This cutoff point, 0.719798, is the optimal PRS threshold. This threshold of 0.719798 provides the best case / control classification accuracy on this dataset. This optimal SNP combination consists of 155 SNPs, which can form a biomarker panel for analyzing and assessing HLH risk.

[0104] Example 2

[0105] This embodiment newly includes data from 366 HLH patient samples and 3000 control group samples. The risk of HLH in the HLH patient samples and control group samples was calculated using the 155 SNPs and their effect values ​​that indicate a high risk of HLH obtained in Example 1.

[0106] The results are as follows Figure 2-3 As shown, Figure 2 The ROC curve and AUC value for this embodiment are shown. Figure 3 The results are presented as a scatter plot. The results show that the AUC value of the ROC curve using this combination of 155 high-risk SNPs for HLH reached 0.653 (95% CI: 0.621–0.684), indicating that this combination of high-risk SNPs for HLH has moderate to good discriminative power in the application of this combination in 366 HLH patient samples and the control group.

[0107] Example 3

[0108] This embodiment newly incorporates data from 366 HLH patient samples and 3000 control group samples, which are different from those in Embodiment 2. The risk of HLH in the HLH patient samples and control group samples is calculated using the 155 SNPs and their effect values ​​that indicate a high risk of HLH obtained in Embodiment 1.

[0109] The results are as follows Figure 4-5 As shown, Figure 4 The ROC curve and AUC value for this embodiment are shown. Figure 5 The results are presented as a scatter plot. The results show that the AUC value of the ROC curve using this combination of 155 high-risk SNPs for HLH prediction reached 0.671 (95% CI: 0.6419-0.7005, slightly higher than in Examples 1 or 2), indicating that this combination of 155 high-risk SNPs has a certain discriminative ability and predictive value for HLH risk prediction, reflecting that this combination of high-risk SNPs for HLH can effectively capture the genetic characteristics of HLH patients.

[0110] In summary, this invention, through GWAS and multi-gene risk score analysis, identified a combination of 155 SNPs as gene biomarkers, providing a more discriminative indicator for the early diagnosis of HLH. This innovative design helps improve the accuracy of predicting patient risk and provides a theoretical basis for developing personalized treatment plans in clinical practice.

Claims

1. The use of reagents for detecting mutation sites in the preparation of kits for assessing the risk of hemophagocytic lymphohistiocytosis, characterized in that, The mutation sites include the following mutation sites located based on hg19 in the reference genome: 1:1407372:T:C,1:13183567:C:T,1:13448283:A:G,1:20442054:T:C,1:21808159:G:A,1:29650289:C:G,1:31423116:T:A,1:67451602:A:C,1:85018771:G:GA,1:112269872:G:A,1:117158972:A:G,1:149224578:CT:C,1:151142978:T:C,1:152191578:G:A,1:152552279:C:A,1:152975763:C:A,1:155823124:G:A,1:161643889:G:T,1:185894153:GT:G,2:219679800:G:A,2:228121101:G:T,2:241451350:C:T,3:15092233:C:T,3:31704928:TCA:T,3:36880027:C:T,3:49723274:G:A,3:110912166:ACAT:A,3:130992424:T:C,3:154890037:C:A,4:664028:ATT:AT,4:39864448:G:A,4:76584154:C:T,4:103822298:C:T,4:144922436:T:G,4:145049120:G:A,5:61648394:A:AT,5:67522851:A:C,7:124486980:T:C,7:143043240:C:T,7:143105804:C:G,7:143270220:C:T,7:153750014:G:A,8:6484505:C:CA,8:11189529:T:C,8:101721812:G:A,8:117954790:C:G,9:33798440:C:A,9:37501844:G:A,10:5766389:A:G,10:29779789:C:T,10:51465552:A:T,10:51623061:G:A,10:81472483:G:T,11:9192336:T:TAC,11:12263714:G:T,11:18034592:G:C,11:44640252:C:T,16:4908667:A:G,16:28506399:T:C,16:28603585:T:A,16:47292699:TA:T,16:57965793:G:T,16:74422332:A:G,16:81069449:A:G,16:81129887:C:A,17:7386114:A:G,17:7386280:G:A,17:21156530:G:GC,17:21319868:G:T,17:33769034:G:A,17:45219332:T:G,17:45234710:A:C,17:73827216:C:T,18:47364009:G:A,18:72913590:C:T,7:97937258:T:C,4:103808519:C:T,1:236730215:T:C,1:209602719:C:T,2:47698132:A:G,2:87169316:T:G,2:96521928:C:T,2:96521944:C:A,2:96593000:A:G,2:96601312:G:A,2:96626292:C:T,2:97749576:C:G,2:97829707:C:T,2:97830053:A:G,2:97833399:C:T,2:97879162:G:C,2:98129690:T:C,2:107051270:T:C,2:132237927:C:A,2:165802112:G:A,2:178080396:T:A,2:182344846:G:GT,5:80655709:T:C,5:141004743:T:C,5:150891902:T:A,5:180278369:T:C,6:15593316:T:C,6:30671206:T:C,6:31239455:T:C,6:31903786:C:G,6:32605366:T:C,6:46822406:A:G,6:129955151:GA:G,6:136590640:A:C,6:144263089:C:A,7:32602381:A:AT,7:45123210:G:A,7:55238535:CT:C,7:100551122:T:G,7:100639419:G:A,11:46626256:T:C,11:49197416:A:G,11:118978411:G:C,12:6172183:C:G,12:7175872:T:C,12:8327888:A:G,12:16713487:A:AT,12:40879698:C:T,12:49883204:G:T,12:54757149:G:T,12:75804273:G:A,12:85449874:G:A,13:111365593:T:TG,14:24030430:G:A,14:51062357:G: A,14:77901758:G:GAA,14:90782967:G:T,14:96001383:A:C,14:96707777: G:A,14:96998645:T:G,19:685781:G:A,19:1975435:C:T,19:9006437:G:A, 19:40370296:G:C,19:40883844:A:C,19:41518245:G:C,19:46857021:C:T,2 0:47615074:TA:T,20:49551780:T:TA,20:60640030:A:G,21:15889228:G:T ,21:33785343:G:A,22:23959752:G:A,22:24087319:A:C,22:41547944:A:G, 18:77229409:C:G,19:549091:C:A,1:2322188:G:GCTCTTTTTTTTTTTTTTTTGAGACTGAGTCTTGCTCTGTCGCC,5:60628574:CCGGCCGCAACCTCGG:C,and 7:26246130:A: ATGCTGACAATACTTG, In each mutation site, the chromosome number and the site location are indicated before and after the first colon, respectively, and the bases before and after the third colon are indicated before and after the mutation, respectively.

2. A gene detection kit for assisting in the assessment of the risk of hemophagocytic lymphohistiocytosis, characterized in that, Includes reagents for detecting mutation sites as described in claim 1.

3. A detection system for assisting in the assessment of the risk of hemophagocytic lymphohistiocytosis, characterized in that, include: A detection device for detecting mutation sites in a sample to be evaluated for the purpose described in claim 1, and obtaining gene information for each mutation site; The analysis device acquires the mutation status of each mutation site mentioned above, calculates the polygenic risk score (PRS) of the sample using the summation method, compares the PRS with a preset threshold, and obtains the risk of hemophagocytic lymphohistiocytosis. as well as An output device is used to output the above-mentioned risk results for hemophagocytic lymphohistiocytosis.

4. The detection system according to claim 3, characterized in that, The polygenic risk score (PRS) of the sample to be evaluated is calculated using the following formula: , Where: n is the number of mutation sites, and Effect Size i is the weight value of the i-th mutation site, and Allele Count is the genotype value of the mutation site. When there is no mutation at the mutation site, Allele Count is 0; when there is 1 mutation at the mutation site, Allele Count is 1; when there are 2 mutations at the mutation site, Allele Count is 2. The weight of each mutation site is taken from the following β value ± 10% range: ; ; ; 。 5. The detection system according to claim 4, characterized in that, In the analytical device, the following comparison and judgment are made: when the polygenic risk score PRS ≥ the threshold, it is judged as a high risk of hemophagocytic lymphohistiocytosis, and the threshold is 0.719798±0.

07.

6. The detection system according to claim 5, characterized in that, The threshold value is 0.719798 ± 0.

035.

7. The detection system according to claim 6, characterized in that, The threshold is 0.719798.

8. The detection system according to claim 4, characterized in that, The detection system performs the following steps: Genetic testing: Detecting mutation sites in the sample to be evaluated for the purpose described in claim 1, and obtaining genetic information for each mutation site; Polygenic risk analysis: Obtain the mutation status of each mutation site mentioned above, calculate the polygenic risk score (PRS) of the sample using the summation method, compare the PRS with a preset threshold to determine the risk of hemophagocytic lymphohistiocytosis.

9. A computer device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the program, it performs the following steps: analyzing the gene information of each mutation site in the sample to be evaluated for the purpose described in claim 1, according to the execution steps in the detection system described in claim 8.