A method for detecting viral load based on single-cell transcriptome sequencing
By employing a comparison strategy based on protein sequence similarity and evolutionary distance weighting, combined with host immune response and cell type sensitivity, an infection propensity index model was constructed. This solved the problem of detecting unknown viruses and viruses with high mutation rates in single-cell transcriptome sequencing, achieving efficient and accurate viral load detection and infection probability estimation.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- BEIJING GEZHI BOYA BIOTECHNOLOGY CO LTD
- Filing Date
- 2026-04-10
- Publication Date
- 2026-07-28
AI Technical Summary
Existing single-cell transcriptome sequencing technology cannot identify unknown or distantly related viruses, cannot adapt to viruses with high mutation rates, and lacks quantitative models for the degree of infection, resulting in low detection accuracy and inaccurate judgment of infection probability.
A comparison strategy based on protein sequence similarity combined with evolutionary distance weighting was adopted. Amino acid sequences were generated through dynamic translation and compared with viral protein databases in a weighted manner. Combined with host immune response index and cell type sensitivity, a quantitative model of infection susceptibility index was constructed to realize the detection of viral load.
It improves robustness against viruses with high mutation rates, significantly reduces false negative rates, can quickly identify distant and unknown viruses, and enables quantitative estimation of single-cell infection probability, thereby improving the accuracy and efficiency of detection.
Smart Images

Figure CN121999871B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the fields of single-cell transcriptomics and microbiome technology, specifically to a method for detecting viral load based on single-cell transcriptome sequencing. Background Technology
[0002] With the rapid development of high-throughput sequencing technology, single-cell RNA sequencing (scRNA-seq) has become an important tool for analyzing cellular heterogeneity, immune responses, multicellular system composition, and disease progression. Compared with traditional tissue transcriptome sequencing, single-cell transcriptome sequencing can analyze the expression characteristics of each cell and reveal intercellular differences. It has been widely used in research areas such as the tumor microenvironment, immune system development, infectious disease models, and multi-tissue atlas construction.
[0003] In recent years, an increasing number of studies have shown that scRNA-seq data not only contains host transcriptional information but may also carry transcriptional fragments derived from exogenous nucleic acids such as viruses. In viral infection-related research, the ability to identify virus-infected cells at single-cell resolution, analyze viral distribution patterns across different cell types, and study infection-induced changes in host gene expression is of significant scientific importance and promising application prospects. With the widespread adoption of single-cell transcriptome sequencing technology, identifying viral infections at the single-cell level has become a crucial requirement for host-pathogen interaction research and emerging virus surveillance. However, existing scRNA-seq virus detection technologies still have several limitations. Traditional virus detection methods primarily rely on reference genomes of known viruses, matching sequencing data with reference sequences through nucleotide alignment to identify infection events. However, these methods cannot detect unknown or distantly related viruses, are insensitive to mutated viruses, and are computationally intensive and slow.
[0004] For example, patent CN115512767A discloses a method for detecting and analyzing viral expression levels in single-cell transcriptome sequencing data. This method involves downloading reference genome sequences of the host and a specific virus, and then adding the viral genome to the host genome as a special chromosome to construct a "host-virus" composite reference genome. Single-cell sequencing reads are then compared with this composite reference genome to obtain the coverage and expression levels of different viral genes, and the expression correlation between viral and host genes is further calculated to analyze virus-host interactions. While this technology can perform quantitative analysis of known viral transcripts at the single-cell level, it relies on the complete reference genome of known viruses and cannot identify viruses without reference sequences, nor can it discover and identify unknown or novel viral sequences from single-cell sequencing data. Therefore, this method has significant limitations when dealing with unknown pathogens, mixed infections, or novel viruses in complex tissue microenvironments. Patent CN116758988A, a method and apparatus for analyzing microbial information from single-cell transcriptome data, also uses reference genome alignment to determine microbial information.
[0005] In summary, the existing methods have the following problems:
[0006] a) Unable to identify unknown or distantly related viruses. Current technologies generally rely on reference genomes of known viruses or virus-specific genetic markers, limiting their ability to detect viruses already recorded in databases. They cannot detect new viruses, unknown variants, or distantly related viruses that differ significantly from known viruses.
[0007] b) It is not sensitive to RNA viruses with high mutation rates. Existing nucleotide alignment methods are difficult to adapt to the high mutation rate of RNA viruses. Base changes can easily lead to alignment failures, which can result in false positives for uninfected viruses, reduce detection accuracy, and lead to insufficient robustness in virus identification.
[0008] c) Lack of quantitative models for the degree of infection. Existing methods usually only determine whether an infection has occurred, and lack a modeling approach that integrates information such as viral load, cell type sensitivity, and co-infection with multiple viruses, making it impossible to accurately infer the degree or probability of infection.
[0009] Therefore, developing a new technology that does not rely on a complete viral reference genome, can tolerate viral mutations, and can rapidly identify viral infections at single-cell resolution is an urgent need in the fields of virology and single-cell sequencing. Summary of the Invention
[0010] To address the problems existing in the prior art, this invention proposes a method for detecting viral load based on single-cell transcriptome sequencing.
[0011] Technical solution: The first aspect of this invention proposes a method for detecting viral load based on single-cell transcriptome sequencing, comprising the following steps:
[0012] S1 extracts single-cell transcriptome sequencing data to form a single-cell gene expression matrix, and reads data that is not aligned to the host reference genome from the file compared with the host genome. It records the cell barcode and unique molecular identifier of these data and extracts the corresponding reading data.
[0013] S2 dynamically translates the reading data extracted in step S1 to generate an amino acid sequence;
[0014] S3 compares and scores the amino acid sequence with the viral protein sequence in the virus database, and obtains a weighted alignment value after distance weight correction. This weighted alignment value and the reading data in step S1 form a virus sequence-cell matrix.
[0015] S4 is based on the virus sequence-cell matrix in step S3 to obtain the virus reading data V in a single cell; based on the single cell gene expression matrix in step S1, the host cell immune response index H is obtained according to the IFN pathway gene set, and a cell type sensitivity score C is assigned according to the cell type.
[0016] S5 inputs the virus reading data V, host cell immune response index H, and cell type sensitivity score C obtained in step S4 into the infection tendency index quantitative model to obtain the cell infection probability.
[0017] Furthermore, in step S3, the formula for calculating the distance weight correction is: ; In the formula, Alignment score adjusted for distance weight;
[0018] Based on the comparison score;
[0019] The evolutionary distance between two nodes in a phylogenetic tree constructed for a viral protein reference database;
[0020] Let be the evolution distance weight function, where:
[0021] .
[0022] Furthermore, when A value less than 0.05 indicates an unknown virus.
[0023] Preferably, in step S4, the cell type sensitivity score C ranges from 0 to 1.
[0024] Furthermore, in step S5, the quantitative model for the infection propensity index is as follows:
[0025] ;
[0026] In the formula,
[0027] The model is a sigmoid function; the model parameters a, c, d, and e are the parameters estimated by maximum likelihood, the parameters learned through supervised learning, the parameters learned through the EM algorithm, and the fixed empirical parameters, respectively; V CPM CPM normalization was performed on viral reads of quantity V to reflect the absolute expression level of viral RNA; H was the host immune index normalization to capture the host immune response; and C was the cell type sensitivity normalization to characterize the susceptibility differences caused by cell type.
[0028] Preferably, the ILI value ranges from 0 to 1.
[0029] A second aspect of this invention provides a device for detecting viral load based on single-cell transcriptome sequencing, comprising:
[0030] The positive and negative strand dynamic translation module is used to dynamically translate the reading frames of reads data on the positive and negative strands to generate amino acid sequences;
[0031] The viral protein alignment module based on evolutionary distance weighting is used to calculate the phylogenetic distance between each viral protein. The amino acid sequence is compared with the alignment score by the rapid alignment software, and then weighting coefficients are set according to the distance to realize viral sequence identification and obtain the viral sequence-cell matrix.
[0032] The ILI value calculation module uses the viral load characteristic V obtained from the virus sequence-cell matrix, the host immune response index H, and the cell type susceptibility score C to obtain the ILI value of the cell according to the infection tendency index quantitative model.
[0033] A third aspect of the present invention provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the method for detecting viral load based on single-cell transcriptome sequencing.
[0034] A fourth aspect of the present invention provides a non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that the computer program, when executed by a processor, implements the method for detecting viral load based on single-cell transcriptome sequencing.
[0035] The fifth aspect of the present invention provides a computer program product, comprising a computer program, characterized in that, when the computer program is executed by a processor, it implements the method for detecting viral load based on single-cell transcriptome sequencing.
[0036] Beneficial effects: This invention proposes a method for detecting viral load based on single-cell transcriptome sequencing. It employs protein sequence similarity combined with evolutionary distance weighting to avoid the technical problem of poor sensitivity to RNA viruses with high mutation rates. Furthermore, it proposes an Infection Propensity Index (ILI) quantitative model to achieve quantitative estimation of single-cell infection probability, suitable for detecting distantly related viruses or novel viruses not yet included in the database. The beneficial effects are as follows:
[0037] a) It exhibits higher robustness against viruses with high mutation rates. Unlike nucleic acid alignment that relies on base matching, this invention uses protein sequence similarity combined with evolutionary distance weighting. Protein sequences are protected by codon degeneracy, thus maintaining alignment stability even with high viral nucleic acid mutation rates and significantly reducing the false negative rate.
[0038] b) High computational speed, adaptable to large-scale single-cell data. This invention uses DIAMOND protein alignment, which is tens to hundreds of times faster than other software and can process millions to tens of millions of reads.
[0039] c) Ability to detect new viruses. This invention proposes a weighted comparison score; a score less than 0.05 indicates an unknown virus sequence.
[0040] d) Propose a quantitative model for ILI (Infection Propensity Index). Existing technologies can only determine whether an infection has occurred, but cannot provide the degree of infection. This invention constructs an ILI model that integrates viral expression intensity (V... CPM The system uses host immune response (H) and cell type susceptibility (C) to achieve a quantitative estimate of the probability of single-cell infection.
[0041] In summary, this invention solves the technical problems that urgently need to be addressed in the prior art, and is suitable for quantifying the degree of cell infection. It integrates multiple variables such as the number of viral reads, host antiviral gene expression, and co-infection information to accurately determine the true infection status of different cells, which is conducive to its widespread application. Attached Figure Description
[0042] Figure 1 This is a schematic diagram of the method flow of the present invention;
[0043] Figure 2 A schematic diagram of the structure of the electronic device provided by this invention;
[0044] Standards in the diagram: 310, processor; 320, communication interface; 330, memory; 340, communication bus. Detailed Implementation
[0045] To further illustrate the technical means and effects of this invention, the following description, in conjunction with embodiments and accompanying drawings, provides a further explanation of the invention. It is understood that the specific embodiments described herein are merely illustrative of the invention and not intended to limit it.
[0046] This invention addresses the limitations of existing virus detection methods, which generally rely on reference genomes of known viruses and identify viruses through nucleotide sequence alignment. This approach is insufficient for detecting distantly related viruses or novel viruses not yet included in the genome sequence, thus limiting its ability to explore unknown viruses. Secondly, most methods analyze tissue transcriptome sequencing data, failing to pinpoint the specific cells where viral infection occurs and thus hindering research on virus-host interactions at the single-cell level. Furthermore, traditional methods, dependent on nucleotide alignment, are less sensitive to RNA viruses with high mutation rates, and the high computational complexity of nucleotide alignment makes it unsuitable for large single-cell sequencing datasets. Simultaneously, existing methods generally lack the ability to quantify the degree of cellular infection, failing to integrate multiple variables such as viral read count, host antiviral gene expression, and co-infection information, making it difficult to accurately determine the true infection status of different cells.
[0047] The first aspect of this invention proposes a method for detecting viral load based on single-cell transcriptome sequencing, comprising the following steps:
[0048] S1 extracts single-cell transcriptome sequencing data to form a single-cell gene expression matrix, and reads data that is not aligned to the host reference genome from the file (BAM) that is compared with the host genome. It records the cell barcode and unique molecular identifier (UMI) of these data and extracts the corresponding reads data.
[0049] Reads that do not align to the host reference genome (unmap) may originate from viruses, therefore these unmapped reads undergo step S2. These unmapped reads contain sequence information and corresponding cell barcode records. If the virus in the unmapped read is identified, then it can be determined which cell the virus resides in.
[0050] S2 dynamically translates the reading data extracted in step S1 through three reading frames on each of the positive and negative strands to generate an amino acid sequence; each unmapped read is dynamically translated through three reading frames on each of the positive and negative strands to generate an amino acid sequence.
[0051] S3 compares and scores the amino acid sequence with the viral protein sequence in the virus database, and obtains a weighted alignment value after distance weight correction. This weighted alignment value and the reading data in step S1 form a virus sequence-cell matrix.
[0052] Existing nucleotide alignment methods are ill-suited to the high mutation rate of RNA viruses. Base variations can easily lead to alignment failures, resulting in false positives for non-infection, reduced detection accuracy, and insufficient robustness in virus identification. This invention proposes an innovative alignment strategy based on protein sequence similarity combined with evolutionary distance weighting. Because protein sequences are buffered by codon degeneracy, even with high-frequency mutations at the nucleic acid level, the corresponding amino acid sequences maintain high conservation. Therefore, even with viral genomes exhibiting high mutation rates, this strategy maintains the stability and reliability of the alignment process, significantly reducing the risk of false negatives due to sequence variations, improving detection sensitivity, and enhancing the robustness of virus identification.
[0053] Specifically, a phylogenetic tree is calculated for each sequence in a viral database (such as RdRP). Viral databases (such as RdRP) are derived from public databases such as NCBI RefSeq and Viral Orthologous Groups, which contain viral sequences from human, animal, plant, and environmental sources. Based on the phylogenetic tree constructed from these databases, the phylogenetic distance between viral proteins is calculated, covering all viruses from different hosts. Evolutionary distance reflects the phylogenetic relationship between viruses, not the relationship between the virus and its host. Therefore, the method of this invention is applicable to the analysis of viruses from human, animal, plant, and environmental sources. This invention uses DIAMOND (v2.1.16) fast alignment software to obtain alignment scores. Then set the weighting coefficient according to the distance. This weighting strategy assigns higher weight to viral sequences that are more evolutionarily similar to the original sequence in the alignment score. This significantly improves the detection capability of distantly related and unknown viruses while reducing the risk of mismatches with host-origin noise sequences, thus achieving higher specificity and sensitivity in viral sequence identification. The final alignment score between the translated sequence and the viral protein sequence is corrected using evolutionary distance weighting, calculated as follows:
[0054]
[0055] In the formula,
[0056] Alignment score adjusted for distance weight;
[0057] Based on the comparison score;
[0058] The evolutionary distance between two nodes in a phylogenetic tree constructed for a viral protein reference database;
[0059] The evolution distance weighting function uses an exponential decay function:
[0060] .
[0061] In this embodiment, RdRP data was used, which contains 296,561 viral sequences. Clustering was used to obtain 37 representative viral sequences for phylogenetic tree calculation, and the maximum evolutionary distance was obtained. The value is 3.04293219, therefore the model uses 3 as the threshold, i.e., the evolutionary distance. Values greater than 3 are considered different viruses. e to the power of -3 is 0.05, therefore when... A value less than 0.05 indicates an unknown virus. Therefore, this invention can discover new viruses, solving the technical problem that existing methods mainly rely on the reference genome of known viruses and cannot discover unknown or distantly related viruses.
[0062] The 37 viral sequences are encoded as u1001, u1010, u10011, u10015, u100000, u100001, u100002, u100004, u100007, u100011, u100012, u100017, u100019, u100024, u100026, u100028, u100031, u100041, u100 The specific sequences are as follows: 048, u100049, u100057, u100074, u100076, u100093, u100111, u100116, u100117, u100152, u100153, u100154, u100173, u100177, u100251, u100296, u100631, u100644, and u102324.
[0063] The serial number of u1001 is SEQ ID NO:1, and the sequence is:
[0064] ITGDNTKWNENQNPRMFLAMITYITRNQPEWFRNILSMAPIMFSNKMARLGKGYMFESKKMKLRTQIPAEMLASIDLKYFNESTRKRIEKIRPLLIDGTASLSPGMMMGMFNMLSTVLGVSILNLGQKKYTRTTYWWDGLQSSDDFAL;
[0065] The serial number of u1010 is SEQ ID NO:2, and the sequence is:
[0066] YADDTAGWDTRITEDDLQNEAKITDIMEPEHALLAKSIFKLTYQNKVVRVQRPAKNGTVMDVISRRDQRGSGQVGTYGLNTFTNMEAQLIRQMESEGIFSPSELETPNLAERALDWLEKHGVERLKRMAISGDDCVV;
[0067] The serial number of u10011 is SEQ ID NO:3, and the sequence is:
[0068] FSYDTRCFDSTVTEHDIMTEESIYQSCDLEPEARVAIRSLTQRLYCGGPMYNSKGQQCGYRRCRASGVFTTSMGNTMTCYIKALASCRAAKLQDCTLLVCGDDLVA;
[0069] The serial number of u10015 is SEQ ID NO:4, and the sequence is:
[0070] FCFDYDDFNSQHSIASMYTVLIAFRDAFFRNMSAQQKEAMDWVCESVKHMWVLDPDTKEWYQLRGTLLSGWRLTTFMNTVLNWAYMKIAGVFDIDDVQDSVHNGDDVMI;
[0071] The serial number of u100000 is SEQ ID NO:5, and the sequence is:
[0072] VGLDASRFDQHVSAEALKYEHSLYNMIFADKELEEYLQWQIKNVGYANFSDGSLKYTVQGVRGSGDMNTALGNVFLMCSITHHYLEGLGIKYRFINDGDDCGV;
[0073] The serial number of u100001 is SEQ ID NO:6, and the sequence is:
[0074] YSYDLSAATDRLPVDLQVDLLSELMGKKLALLWKGLLVSRPYKLPKIAKSYNLGFSEVKYEVGQPMGALSSWAMLALTHHAIVQFAASQVGAKQPKGWFTGYAVLGDDIVI;
[0075] The serial number of u100002 is SEQ ID NO:7, and the sequence is:
[0076] GTIDLESASDSLSLSMLREYIPSDQLAWFEFLRSPKTFIPSQGWTELHILSTMGNGFTFPLQTVLFACVVAAVYDLYDLPLRRPRGSNAGNFGVFGDDIIV;
[0077] The serial number of u100004 is SEQ ID NO:8, and the sequence is:
[0078] VSGDYSSATDNLPIEVVEVILDTMQKTACFVPPAIFKFAQKILRPLMWCFEPTLEIDWEVQVVRGQMMGSSLLSFPMLCLQNRLGLEWACFAEGREVPPCLINGDDILY;
[0079] The serial number of u100007 is SEQ ID NO:9, and the sequence is:
[0080] NSGDYKAATDMINKSLSEVCVEELSKFGELELYETFMFIKSLTGHQVEYPDGDLLKQVWGQLMGSVTSFPILCIINIAVNRYFLEVTEGHSYSLKSLPMLVNGDDIVF;
[0081] The serial number of u100011 is SEQ ID NO:10, and the sequence is:
[0082] YCFDLSSASDRIPAKMQRYRLQLMGGLHVADNWYKVMTKRNFYIKPLGKSIRWSVGQPLGLLSSFPSFALWHHDIIQFAANYERFSNGKPLKFFKNYKLLGDDVVI;
[0083] The sequence number of u100012 is SEQ ID NO:11, and the sequence is:
[0084] STIDLKQASDSLSRGLVSLCVPELWLSLLEAFRVPKIKIDGKTVRMNKFSGMGNGYTFELESVLYMSIAQEVCKMLGLPSTANHDVYVFGDDIIV;
[0085] The serial number of u100017 is SEQ ID NO:12, and the sequence is:
[0086] GTLDLSEASDRVSVQHVRDLLRTTPSLREAVDSCRRSRTASVPSYGVNSVKELHLAKFASMGSALCFPFEAMVFLTVVFCGIEKQLGRPLTLRDIKSFREQVRVYGDDLIV;
[0087] The serial number of u100019 is SEQ ID NO:13, and the sequence is:
[0088] DGGDTTKMDAHMRPAQLRLVFHIVKWLFQEAYWDELDRSLQHITNIPLLIGEDEMLVGTHGLASGSSWTQLSETVLQMFMAYVSGTHGQGIGDDFYW;
[0089] The sequence number of u100024 is SEQ ID NO:14, and the sequence is:
[0090] VCTDFSAFDQHFNSSLSSVSKSLLEYLFVDSDDTHRWLETVFPIKYHIPLAYDWGALKTGEHGMGSGSGGTNADETLAHRALQYEAAIRNGKRLNLTSQCLGDDGLL;
[0091] The sequence number of u100026 is SEQ ID NO:15, and the sequence is:
[0092] YSIDSSAFDSSISRFLIRQAFRIIKTWFNLDDVEPTTGVAVRDILKQIEVYFSTCPLVMPDGRIYYGRRHGVPSGSFFTQIVDSIVNTIVAGTISARFHMSVGRSEIFVLGDDLLI;
[0093] The serial number of u100028 is SEQ ID NO:16, and the sequence is:
[0094] LSMDYSAFDSSLSCNMIIAAGEILKSWFSPDAEPIINLLVQSLCTTALVTPAGVFTGRNGGMPSGSVLTNLVDTLCNLIAGHYSALRNGTSIRRFAVLGDDSVY;
[0095] The serial number of u100031 is SEQ ID NO:17, and the sequence is:
[0096] VSGDDSFVVVSWEGQLFTFEGDATMYDQSESFGPLGVAWKACRRLGVDSSTIKLLSELAHNSYTVRSKDPTDQATLRINKRNRPIRDTGGADTSLGNSLVM;
[0097] The sequence number of u100041 is SEQ ID NO:18, and the sequence is:
[0098] IESDADKWDGHYASPSREFERMVLLALYKVIYHDRVIAAHRKQYNLKGITSFGVIYDLLFSRGSGELGTGSFNSLSNMKIAYCGYREQGLTPEQAWESLGIFGGDDALN;
[0099] The serial number of u100048 is SEQ ID NO:19, and the sequence is:
[0100] HSIDLSGATDYFPLALQEHLLRRMFPDIEVDLFVDLSKASWYMPREGEISWRRGQPLGLYPSFGAFALTHGTLLLGLLNKGWNNQFFVLGDDVVI;
[0101] The serial number of u100049 is SEQ ID NO:20, and the sequence is:
[0102] LSGDFKAATDNLESWVSEQICNDIADELGLFPVERRKLIESLTRHIFVDHEGIEHIQTKGQLMGSITSFIVLCIANVTCLRWACEIDQRRQMSLESAPIMVNGDDGAA;
[0103] The serial number of u100057 is SEQ ID NO:21, and the sequence is:
[0104] VSTDLTRATDLLPSDLCGAVVDGLELSGRLSPTEIQVLRSLTGPQILDYPDGTVLQNTRGILMGLPTSWAVLSLIHLFWLSEVKATSPVGVEKPHLHKFSICGDDALL;
[0105] The sequence number of u100074 is SEQ ID NO:22, and the sequence is:
[0106] YSYDLTAATDRIPIQLYEIMLEHLLDDRDYAQAWKRLMIGIPFSYEGSAVSYAAGQPMGMYSSWPLMAMSHHIVVRYAAXSVGIPDFSDYRLLGDDIVI;
[0107] The serial number of u100076 is SEQ ID NO:23, and the sequence is:
[0108] ITMDYSRYDSRVRKYQLHGEHSLYLKAFNNDRTLMRLLSYQINNVGKDYSGNISYKVEGTRATGEMNTLIGNSLINAAMLIGIAKYVGIKIDLMVCGDDSII;
[0109] The serial number of u100093 is SEQ ID NO:24, and the sequence is:
[0110] FNLDWSEFDMRVYFSMWLDIINEVSTYFCFCGKYCPTRTYPNSQTSPKRITNLWNRITDMYFNMPCVTPMGNVFNRLYAGMPSGIFCTQFYDSIYNGVMIVTCLLALGVPIPDDFPIKLMGDDVLF;
[0111] The serial number of u100111 is SEQ ID NO:25, and the sequence is:
[0112] VMLDYDDFNSRHSNENMVMIFEELFSHIGHTSILADKVCRSFHNSYMYNPGTDSHDPIAGTLMSGHRATTFINSVLNFAYIYACAPEIADMRSMHVGDDILI;
[0113] The sequence number of u100116 is SEQ ID NO:26, and the sequence is:
[0114] KYMDIDDCDVLMMRRLWAILCQAEHVVLDEVWYCSQGVFSGNPFTDCINTMVNNIYVRCAFIGCCTKYAPHLASLEDGVYNLRNFRKHVRVVAYGDDLII;
[0115] The serial number of u100117 is SEQ ID NO:27, and the sequence is:
[0116] CTIDYSNFGPGFNAVVAQQAYDLMIKWLMENVRGMSKTELDCIKEECLNSYHICINTVYYQKCGSPSGAPITTIINTLVNQLLVLIAWDYLMKDHPLLADKFSTEVYQQHVALFCYGDDLIM;
[0117] The serial number of u100152 is SEQ ID NO:28, and the sequence is:
[0118] FSFDLSAATDRLPLALQKQILTIAVSYEFAEAXGTVLTGRSYHLFYKKQNYDLTYSVGQPMGALSSXGMLALTHHVIVQIAASRSGRSELFLDYALLGDDICI;
[0119] The serial number of u100153 is SEQ ID NO:29, and the sequence is:
[0120] TSGDYKSASDNLSIEVAEAILDVAWSTSMRVPSNVFRFALAAQRPSLAYEDENGLLEHFVPTRGQMMGSYLCFPLLCLQNYIAFRYAEKASGLSGTPVLINGDDILF;
[0121] The sequence number of u100154 is SEQ ID NO:30, and the sequence is:
[0122] VELDWAKFDRERPAEDIAFVIDVVLSCFEPRNAREQRLLEGYGLMMRRALVERLVIMDEGGVFGIDGMVPSGSLWTGWLDTALNILYIRAACVEAGFGASGFKPMCAGDDNLT;
[0123] The serial number of u100173 is SEQ ID NO:31, and the sequence is:
[0124] VSGDYESATDNLNISLSHLMIDCMRQTSTHIPRTVWREARKALVAKFSDGRAQSRGQLMGSLLSFPLLCLANFLAFKWAVKRKVPLKINGDDIVF;
[0125] The serial number of u100177 is SEQ ID NO:32, and the sequence is:
[0126] LGLDASKFDMHVCIWLLMYEHSFYTALFPGSEELAELLEQQLHNQGIAFAPDGSVKFKVQGTRASGDLNTSLGNCILMCAMVWVLLKKKGINAELVNNGDDCVV;
[0127] The serial number of u100251 is SEQ ID NO:33, and the sequence is:
[0128] VALDFSAFDASLSSRLIDDAFGILKTHVDLSDPAEARLWDRLVNEFIHSRIVLPDQSIWQVHRGVPSGSAFTSLIDSVCNLIIINYVWLKVLGYAPDPDGVMILGDDSIV;
[0129] The serial number of u100296 is SEQ ID NO:34, and the sequence is:
[0130] LAGDFSNFDGSLNSMVLWRIFDVIEHWYKMNDPNYSIDHYMVRQTLWAHIVNSVHIYKNTLYQWTHSQPSGNPFTVIINSIYNSMILRCAYLTCMRDNCKEDPRDHSLYTLEAFDKHVAVITYGDDNRL;
[0131] The serial number of u100631 is SEQ ID NO:35, and the sequence is:
[0132] VNTDFSRFDATVGVLHDLLFKSCMMRAFAPEYHSELARLMSKETFAKAYTRYGAKYNTGTKTNSGSAATANRNSVINAFTSYVALRHSLTENEAYEALGMYGGDDGAT;
[0133] The sequence number of u100644 is SEQ ID NO:36, and the sequence is:
[0134] LSSDLKEATDNIPHEVAKQLLRGFIDGFGLQSHLINLAVELLCSPRRLEHQDEVWVSRRGVMMGEPLTKVTLTLLNLVVEEEAIRFYLNKPEAPIHVPWRCFHVGGDDHIA;
[0135] The sequence number of u102324 is SEQ ID NO:37, and the sequence is:
[0136] VTGDYKNFGPGLMKKCVTKAKNIILSWYRQYGASLEHINVLNIMFEEICDAFYLGGNLIYQTPCGIPSGSPITAPLNSLVNSLYIRYVYGCLVDRNFSKFGENVFVCTYGDDVIM.
[0137] S4 is based on the virus sequence-cell matrix in step S3 to obtain the virus reading data V in a single cell; based on the single cell gene expression matrix in step S1, the host cell immune response index H is obtained according to the IFN pathway gene set, and a cell type sensitivity score C is assigned according to the cell type, with C ranging from 0 to 1.
[0138] This step requires obtaining the core input features:
[0139] Viral read data V: Based on the viral gene expression identified in step S3, the number of sequencing reads mapped to the viral genome in each cell is directly counted, serving as the most direct evidence of the presence of the virus within the cell. V is used to identify the viral reads obtained.
[0140] Host cell immune response index H: Based on the IFN pathway gene set (IFIT1, IFITM3, ISG15, IFI44L, OAS1, etc.), the host cell immune response index H is calculated using the addmodule module of the Seurat analysis tool. H is used to reflect host immune expression, because viral infection will inevitably trigger an immune response.
[0141] Cell type susceptibility score C: Based on known biological knowledge, a priori values (range 0-1) are assigned to the basic infection susceptibility of different cell types. For example, viruses generally infect epithelial cells more easily (score close to 1), epithelial cells > immune cells > fibroblasts, while susceptibility to immune cells or fibroblasts decreases in that order (scores decrease accordingly). C is used to reflect the virus's susceptibility to host cells.
[0142] S5 inputs the virus reading data V, host cell immune response index H, and cell type sensitivity score C obtained in step S4 into the infection tendency index quantitative model to obtain the cell infection probability.
[0143] The formula for calculating the Infection Proneness Index (ILI) quantitative model is as follows:
[0144]
[0145] In the formula, The model is a sigmoid function; the model parameters a, c, d, and e are the parameters estimated by maximum likelihood, the parameters learned through supervised learning, the parameters learned through the EM algorithm, and the fixed empirical parameters, respectively; VCPM The viral reads, V, are normalized to CPM (counts per million) to reflect the absolute expression level of viral RNA; H is the host immune index normalization, used to capture the host immune response; C is the cell type sensitivity normalization, used to characterize the susceptibility differences caused by cell type. The ILI value ranges from 0 to 1, representing the infection probability of the corresponding cell, and can be used for high-precision identification of infected cells. The training process for model parameters a, c, d, and e is as follows: the original data is uniformly normalized, and training is performed using four columns of information, V... CPM Columns H and C are used as feature variables, column E contains fixed empirical features constructed based on domain experience, and ILI is used as the target variable. A linear model is constructed using the Ridge Regression algorithm. Search parameters are pre-set using a grid search, which is a range within which an optimal value is found as the model training result. The final optimal weight combination (a, c, d, e) is obtained, enabling the model to achieve its best predictive performance on the validation set. The trained model object, normalizer, and optimal parameter configuration are saved for direct prediction on new data.
[0146] The ILI model proposed in this invention addresses the problem that existing technologies can only determine whether an infection has occurred, but cannot provide the degree of infection. The ILI model constructed in this invention integrates viral expression intensity (V... CPM This allows for the quantitative estimation of single-cell infection probability by considering host immune response (H) and cell type susceptibility (C). This enables further in-depth mechanistic studies or serves as a precise indicator for clinical diagnosis.
[0147] This invention also proposes a method and analysis device for detecting viral load based on single-cell transcriptome sequencing, which includes at least a positive and negative strand dynamic translation module, an evolutionary distance-weighted viral protein alignment module, and an ILI value calculation module connected in sequence.
[0148] The positive and negative strand dynamic translation module is used to dynamically translate the unmapped reads data in step S1 into three reading frames for each of the positive and negative strands to generate amino acid sequences.
[0149] The viral protein alignment module based on evolutionary distance weighting is used to calculate the phylogenetic distance between each viral protein. The alignment score is obtained through rapid alignment software, and then weighting coefficients are set according to the distance, thereby achieving higher specificity and higher sensitivity of viral sequence identification.
[0150] The ILI value calculation module is based on viral load characteristic V, host immune response index H and cell type susceptibility score C, and obtains the cell's ILI (infection susceptibility index) according to the infection susceptibility index quantitative model.
[0151] Example 1
[0152] A method for detecting viral load based on single-cell transcriptome sequencing, comprising the following steps:
[0153] S1 extracts single-cell transcriptome sequencing data to form a single-cell gene expression matrix, and compares it with the host alignment file to extract reading data;
[0154] The dataset used was a PBMC single-cell transcriptome sequencing dataset (7149 cells) of porcine epidemic diarrhea virus (PEDV, GenBank: PV098397.1), obtained from self-testing after PEDV challenge. Sequencing platform: DNBelabC4.
[0155] Reads that were not aligned to the pig genome were extracted from the BAM file in the DNBC4tools results. The statistical results are shown in Table 1 below.
[0156] Table 1 Statistical Results
[0157] S2 performs dynamic reading frame translation on each read in Table 1 to generate an amino acid sequence;
[0158] S3 compares and scores the amino acid sequence with the viral protein sequence in the virus database, and obtains a weighted alignment value after distance weight correction. This weighted alignment value and the reading data in step S1 form a virus sequence-cell matrix.
[0159] The viral protein database was compared using DIAMOND, and a weighted comparison value was calculated by incorporating phylogenetic distance. 23 cells were detected carrying viral reads, all of which were PEDV. This demonstrates that this embodiment accurately identified the 23 cells carrying viral sequences (all PEDV) from 7149 cells, without erroneously detecting other viruses. This indicates that the virus detection specificity of the method of this invention is reliable.
[0160] S4 is based on the virus sequence-cell matrix in step S3 to obtain the virus reading data V in a single cell; based on the single cell gene expression matrix in step S1, the host cell immune response index H is obtained according to the IFN pathway gene set, and a cell type sensitivity score C is assigned according to the cell type.
[0161] S5 inputs the parameters obtained in step S4 into the cell infection propensity index calculation formula, and obtains the cell infection probability through the infection propensity index quantitative model. Three key feature values are calculated for each cell and substituted into the ILI model to obtain the cell infection probability. Example results for five cells are shown in Table 2 below.
[0162] Table 2 shows the results for five cells.
[0163] As shown in Table 2, Cell_1 had the highest infection probability, while Cell_4 had the lowest. Specifically, cells with viral loads (Cell_1, Cell_2) had significantly higher ILI values than cells without viral loads (Cell_4, Cell_5). Even without viral loads, the model correctly integrated immune activation (H value) and cell sensitivity (C value), providing a reasonable probability ranking (Cell_3 > Cell_5). Suppressed cells (Cell_4, Cell_5) had the lowest probability. Traditional methods would consider Cell_3 negative, but the model judged it to have a moderate infection risk, indicating that the method of this invention can capture potentially infected cells that might have been missed.
[0164] In summary, the embodiments of the present invention have successfully detected cells infected with a specific virus (PEDV) and demonstrated that the infection probability output by the model can accurately distinguish between positive and negative cells and is consistent with common immunological knowledge, thus fully demonstrating the feasibility and effectiveness of the method for detecting viral load based on single-cell transcriptome sequencing.
[0165] Example 2
[0166] The method for detecting viral load based on single-cell transcriptome sequencing, as described in this invention, is applied to human nasal epithelial single cells containing an unknown virus. The steps are as follows:
[0167] S1 extracts single-cell transcriptome sequencing data to form a single-cell gene expression matrix, and compares it with the host alignment file to extract reading data;
[0168] We used a human nasal epithelial single-cell transcriptome sequencing dataset (approximately 12,000 cells) derived from a patient's viral infection model, containing known viral reads (respiratory respiratory virus, RSV). Sequencing platform: 10x Genomics; known infected cell percentage: approximately 3–5%.
[0169] Reads that did not align to the human genome were extracted from the BAM files in the CellRanger results. The statistical results are shown in Table 3 below.
[0170] Table 3 Statistical Results
[0171] S2 performs dynamic reading frame translation on each read in Table 3 to generate an amino acid sequence;
[0172] S3 compares and scores the amino acid sequence with the viral protein sequence in the virus database, and obtains a weighted alignment value after distance weight correction. This weighted alignment value and the reading data in step S1 form a virus sequence-cell matrix.
[0173] The DIAMOND viral protein database was used for alignment, and phylogenetic distance was incorporated to calculate weighted alignment values. 823 cells were found to contain viral reads, of which 27 cells had a weighted alignment score less than 0.05 and were marked as potentially unknown viral sequences.
[0174] S4 is based on the virus sequence-cell matrix in step S3 to obtain the virus reading data V in a single cell; based on the single cell gene expression matrix in step S1, the host cell immune response index H is obtained according to the IFN pathway gene set, and a cell type sensitivity score C is assigned according to the cell type.
[0175] S5 inputs the parameters obtained in step S4 into the cell infection propensity index calculation formula, and obtains the cell infection probability through the infection propensity index quantitative model. Three key feature values are calculated for each cell and substituted into the ILI model to obtain the cell infection probability. Example results for three cells are shown in Table 4 below.
[0176] Table 4 shows the results for three cells.
[0177] As shown in Table 4, Cell_A has the highest infection probability, while Cell_C has the lowest infection probability.
[0178] In this embodiment, viral reads were detected in 823 cells out of approximately 12,000 cells. 823 / 12000 ≈ 6.9%, slightly higher than known information (RSV infection, infected cell ratio 3-5%). This demonstrates that the method of this invention can accurately and quantitatively detect expected known viruses (RSV) and identify unknown viruses. Among these 823 cells, 27 cells had a weighted alignment score of less than 0.05 for their viral reads. This low score means that although these reads resemble viral sequences, their similarity to all known viral protein sequences in the database is very low. They are thus marked as "potentially unknown viral sequences." This proves that the method of this invention can not only be used for the precise quantification of known viruses but also has the potential to discover unknown or emerging viruses.
[0179] Figure 2 An example is a schematic diagram of the physical structure of an electronic device, such as... Figure 2As shown, the electronic device may include a processor 310, a communications interface 320, a memory 330, and a communication bus 340, wherein the processor 310, communications interface 320, and memory 330 communicate with each other via the communication bus 340. The processor 310 can call logical instructions in the memory 330 to execute a method for detecting viral load based on single-cell transcriptome sequencing.
[0180] Furthermore, the logical instructions in the aforementioned memory 330 can be implemented as software functional units and, when sold or used as independent products, can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, essentially, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0181] On the other hand, the present invention also provides a computer program product, which includes a computer program that can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer is able to execute the methods provided above for detecting viral load based on single-cell transcriptome sequencing.
[0182] In another aspect, the present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, is implemented to perform the methods for detecting viral load based on single-cell transcriptome sequencing provided by the above methods.
[0183] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Those skilled in the art can understand and implement this without any creative effort.
[0184] Through the above description of the embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus necessary general-purpose hardware platforms, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solutions, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods described in the various embodiments or some parts of the embodiments.
[0185] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A method for detecting viral load based on single-cell transcriptome sequencing, characterized in that, The steps are as follows: S1 extracts single-cell transcriptome sequencing data to form a single-cell gene expression matrix, and reads data that is not aligned to the host reference genome from the file compared with the host genome. It records the cell barcode and unique molecular identifier of these data and extracts the corresponding reading data. S2 dynamically translates the reading data extracted in step S1 to generate an amino acid sequence; S3 compares and scores the amino acid sequence with the viral protein sequence in the virus database, and obtains a weighted alignment value after distance weight correction. This weighted alignment value and the reading data in step S1 form a virus sequence-cell matrix. The formula for calculating the distance weight correction is as follows: ; In the formula, Alignment score adjusted for distance weight; Based on the comparison score; The evolutionary distance between two nodes in a phylogenetic tree constructed for a viral protein reference database; Let be the evolution distance weight function, where: ; S4 derives the virus readout data V in a single cell based on the virus sequence-cell matrix in step S3; based on the single cell gene expression matrix in step S1, the host cell immune response index H is obtained according to the IFN pathway gene set; and a cell type sensitivity score C is assigned according to the cell type. S5 inputs the virus reading data V, host cell immune response index H and cell type sensitivity score C obtained in step S4 into the infection tendency index quantitative model to obtain the cell infection probability. The quantitative model for the infection tendency index is as follows: ; In the formula, The model is a sigmoid function; the model parameters a, c, d, and e are the parameters estimated by maximum likelihood, the parameters learned through supervised learning, the parameters learned through the EM algorithm, and the fixed empirical parameters, respectively; V CPM CPM normalization was performed on viral reads of quantity V to reflect the absolute expression level of viral RNA; H was the host immune index normalization to capture the host immune response; and C was the cell type sensitivity normalization to characterize the susceptibility differences caused by cell type.
2. The method for detecting viral load based on single-cell transcriptome sequencing according to claim 1, characterized in that, when A value less than 0.05 indicates an unknown virus.
3. The method for detecting viral load based on single-cell transcriptome sequencing according to claim 1, characterized in that, In step S4, the cell type sensitivity score C ranges from 0 to 1.
4. The method for detecting viral load based on single-cell transcriptome sequencing according to claim 1, characterized in that, The ILI value ranges from 0 to 1.
5. An apparatus for implementing the method for detecting viral load based on single-cell transcriptome sequencing as described in claim 1, characterized in that, include: The positive and negative strand dynamic translation module is used to dynamically translate the reading frames of reads data on the positive and negative strands to generate amino acid sequences; The viral protein alignment module based on evolutionary distance weighting is used to calculate the phylogenetic distance between each viral protein. The amino acid sequence is compared with the alignment score by the rapid alignment software, and then weighting coefficients are set according to the distance to realize viral sequence identification and obtain the viral sequence-cell matrix. The ILI value calculation module uses viral read data V obtained from the viral sequence-cell matrix, as well as the host immune response index H and cell type susceptibility score C, to obtain the ILI value of the cell according to the infection tendency index quantitative model.
6. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the program, it implements the method for detecting viral load based on single-cell transcriptome sequencing as described in any one of claims 1 to 4.
7. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the method for detecting viral load based on single-cell transcriptome sequencing as described in any one of claims 1 to 4.
8. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by the processor, it implements the method for detecting viral load based on single-cell transcriptome sequencing as described in any one of claims 1 to 4.