Method and system for analyzing hbv protein translation level based on ribo-seq

By constructing an optimized linear HBV reference genome sequence and combining it with signal separation and correction strategies, the problem of superimposed HBV multiple translation signals was solved, enabling accurate quantification of HBV protein translation levels and breaking through the limitations of traditional Ribo-seq analysis.

CN120950846BActive Publication Date: 2025-12-05TIANJIN SECOND PEOPLES HOSPITAL
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511485139.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-10-17
Publication Date
2025-12-05
Estimated Expiration
2045-10-17

AI Technical Summary

Technical Problem

Existing Ribo-seq analysis tools cannot effectively handle signal superposition caused by multiple and frameshift translation in hepatitis B virus (HBV), resulting in the inability to accurately identify reading frames and quantify the translation levels of overlapping proteins, thus limiting in-depth research on the HBV translation regulatory mechanism.

Method used

An optimized linear HBV reference genome sequence was constructed, and by combining median signal density, signal leakage rate, and signal-to-noise ratio, a signal separation and correction strategy was adopted to achieve accurate separation and quantification of HBV protein translation signals.

Benefits of technology

It significantly improves the accuracy and reliability of HBV overlapping translation signal analysis, providing an effective tool for the study of HBV translation regulation mechanisms and the development of antiviral drugs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120950846B_ABST
    Figure CN120950846B_ABST
Patent Text Reader

Abstract

The application provides a method and system for analyzing HBV protein translation level based on Ribo-seq, and belongs to the technical field of bioinformatics and virology. The method comprises the following steps: constructing a linear HBV reference genome sequence and optimizing the start point thereof; aligning Ribo-seq sequencing data to the linear HBV reference genome sequence to obtain P-site signals of each nucleotide position; separating the P-site signals according to reading frames to obtain signal intensity distribution of each reading frame; calculating the median signal density and signal leakage rate of the protein in the pure area of the HBV protein; correcting the observed signal of the target reading frame in the mixed area of the overlapping protein; and outputting the corrected translation signal intensity of the HBV protein. The method and system for analyzing HBV protein translation level based on Ribo-seq can accurately and quantitatively analyze the translation level of the HBV protein.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of bioinformatics and virology, and in particular to a method and system for analyzing HBV protein translation level based on Ribo-seq. BACKGROUND

[0002] Ribosome profiling (Ribo-seq) is a high-throughput sequencing technology that can capture the mRNA fragments (RPFs) being translated in cells with single-nucleotide precision, and identify the translation activity by inferring the ribosome P-site position. In typical eukaryotic genes, the translation region usually presents a clear "three-nucleotide periodicity" signal, which appears on the first nucleotide of the codon and forms a "high-low-low" periodic pattern. This periodicity is a key feature to distinguish real translation from background noise, and is also the core basis for existing Ribo-seq analysis tools (such as RiboWave) to identify translation events and quantify translation efficiency.

[0003] However, the Hepatitis B Virus (HBV) genome has a special structure. Its circular DNA of about 3.2 kb contains four major genes that are in different reading frames (ORFs) and overlap with each other, which encode core protein (C), polymerase protein (P), surface antigen (S), and X protein (X), respectively. Figure 1 These ORFs are in different reading frames and share part of the RNA sequence, which leads to multiple translation events on the same nucleic acid sequence. This signal superposition caused by multiple, frameshifted translation completely masks and destroys the three-nucleotide periodicity characteristic of single translation event, making the traditional Ribo-seq analysis tools that rely on finding a single dominant periodicity completely ineffective. They cannot distinguish which reading frame is the real translation, let alone accurately quantify the translation level of each overlapping protein. Figure 2

[0004] Existing Ribo-seq analysis methods are usually based on the assumption that "one coding region corresponds to one dominant periodicity", and cannot handle the mixed signals caused by multiple, frameshifted translation in HBV. Therefore, traditional tools often fail when analyzing HBV data, and cannot accurately identify the reading frame of each ORF, let alone quantify the translation level of each overlapping protein, greatly limiting the in-depth study of the translation regulation mechanism of HBV.

[0005] Therefore, there is an urgent need in the art for a Ribo-seq data analysis method and system that can effectively separate the HBV overlapping translation signals and accurately analyze the translation activity of HBV proteins. SUMMARY

[0006] ​The purpose of the present application is to provide a method and system for analyzing the translation level of HBV protein based on Ribo-seq, by constructing an optimized linear HBV reference genome sequence, and based on the median signal density, signal leakage rate and signal-to-noise ratio, the mixed Ribo-seq signal is accurately separated, so as to output the independent and accurate translation signal strength of HBV protein; not only significantly improve the accuracy and reliability of the analysis of overlapping translation signal, but also provide an effective tool for in-depth study of HBV translation regulation mechanism and development of antiviral drugs.

[0007] To achieve the above purpose, the present application provides a method for analyzing the translation level of HBV protein based on Ribo-seq, comprising the following steps:

[0008] Step S1, constructing a linear HBV reference genome sequence and optimizing its starting point;

[0009] Step S2, aligning the Ribo-seq sequencing data to the linear HBV reference genome sequence to obtain the P-site signal of each nucleotide position; according to the reading frame attribution of each nucleotide position on the linear HBV reference genome sequence, separating the P-site signal into three reading frames to obtain the signal strength distribution of each reading frame;

[0010] Step S3, for each HBV protein, in the pure region of its genome which only encodes a single protein, based on the signal strength distribution of the three reading frames, calculating the median signal density of the protein and its signal leakage rate to other reading frames;

[0011] Step S4, in the mixed region of overlapping proteins, for the target reading frame of the target nucleotide position, performing:

[0012] Taking the P-site signal of the target reading frame at this position as the observed signal;

[0013] Based on the median signal density and the signal leakage rate, estimating the expected signal of the target protein and the expected interference noise of the overlapping protein, and calculating the signal-to-noise ratio;

[0014] Based on the ratio of the signal-to-noise ratio to the preset threshold, selecting a signal correction strategy to correct the observed signal;

[0015] Step S5, outputting the translation signal strength of the corrected HBV protein.

[0016] Preferably, in step S2, for any nucleotide position, the reading frame to which it belongs is:

[0017] ;

[0018] Wherein, represents the reading frame, represents any nucleotide position, represents a remainder operation.

[0019] Preferably, in step S3, the formula for calculating the signal leakage rate is:

[0020] ;

[0021] wherein, represents the signal leakage rate, represents the main reading frame, represents other reading frames, represents a pure region on the genome that only encodes a single protein, represents the observed signal at any nucleotide position .

[0022] Preferably, in step S4, the formula for calculating the signal-to-noise ratio is:

[0023] ;

[0024] ;

[0025] ;

[0026] wherein, represents the expected signal of the target protein A, represents the codon length of the overlapping region A, represents the median signal density of the target protein A, represents the expected interference noise of the overlapping protein B, represents the mixed region of the overlapping protein, represents the observed signal at any nucleotide position of the overlapping protein B reading frame, represents the signal leakage rate of the overlapping protein B to the target protein A, represents the signal-to-noise ratio.

[0027] Preferably, in step S4, the signal correction strategy includes:

[0028] When the signal-to-noise ratio is greater than or equal to the preset threshold, directly subtract the interference noise for signal correction:

[0029] ;

[0030] When the signal-to-noise ratio is less than the preset threshold, use the density estimation method for signal correction:

[0031] ;

[0032] wherein, represents the corrected signal, observed signal at a target nucleotide position of a target protein A reading frame, Poisson distribution of the median signal density of the target protein A.

[0033] Preferably, the HBV protein comprises at least one of a core protein, a polymerase protein, a surface antigen protein, and an X protein.

[0034] The present application also provides a system for analyzing the translation level of HBV protein based on Ribo-seq, comprising:

[0035] a data input module for constructing a linear HBV reference genome sequence and optimizing its start point;

[0036] a signal separation module for aligning Ribo-seq sequencing data to the linear HBV reference genome sequence to obtain P-site signals at each nucleotide position; and separating the P-site signals into three reading frames according to the reading frame attribution of each nucleotide position on the linear HBV reference genome sequence to obtain signal intensity distribution of each reading frame;

[0037] a benchmark modeling module for calculating the median signal density and signal leakage rate of each HBV protein to other reading frames based on the signal intensity distribution of the three reading frames in the pure region of the genome that only encodes a single protein for each HBV protein;

[0038] a signal correction module for correcting the observed signal of the target reading frame based on the median signal density and signal leakage rate in the mixed region of overlapping proteins;

[0039] an output module for outputting the corrected translation signal intensity of the HBV protein.

[0040] Preferably, the signal correction module comprises:

[0041] a signal observation unit for taking the P-site signal on the target reading frame of the target nucleotide position as the observed signal;

[0042] a signal-to-noise ratio analysis unit for estimating the expected signal of the target protein and the expected interference noise of the overlapping protein based on the median signal density and signal leakage rate, and calculating the signal-to-noise ratio;

[0043] a signal correction strategy execution unit for selecting a signal correction strategy based on the ratio of the signal-to-noise ratio to a preset threshold to correct the observed signal.

[0044] Therefore, the present application adopts the above-mentioned method and system for analyzing the translation level of HBV protein based on Ribo-seq, and has the following beneficial technical effects:

[0045] ​(1) The present application can effectively separate the mixed Ribo-seq signals of multiple overlapping ORFs of HBV by constructing an optimized linear HBV reference genome sequence, combining median signal density, signal leakage rate and signal-to-noise ratio, breaking through the limitations of traditional methods, and realizing independent and accurate quantification of HBV protein translation level.

[0046] (2) The signal correction strategy is selected based on the ratio of signal-to-noise ratio to preset threshold, which not only ensures the direct denoising of strong signal area, but also uses a robust density estimation method in weak signal area, improving the adaptability of the algorithm and providing technical support for HBV translation regulation mechanism research and antiviral drug development. BRIEF DESCRIPTION OF DRAWINGS

[0047] Figure 1 is a schematic diagram of the HBV protein open reading frame in the background art;

[0048] Figure 2 is a sequencing signal diagram of ordinary Ribo-seq and a sequencing signal diagram of double reading frame shift translation Ribo-seq in the background art;

[0049] Figure 3 is a comparison of the decomposition and correction signal intensity of the HBV protein translation group Ribo-seq simulation signal in Example 1 with the actual translation signal intensity and error evaluation, wherein, Figure 3 A in the above formula is the comparison of the actual signal and the decomposition and correction signal intensity of each open reading frame, Figure 3 B in the above formula is the error of the decomposition and correction signal of each open reading frame;

[0050] Figure 4 is a signal intensity distribution diagram of the HBV protein translation group Ribo-seq simulation signal and the predicted signal after decomposition in Example 1, wherein, Figure 4 A in the above formula is the original signal (simulation), Figure 4 B in the above formula is the decomposition and correction signal;

[0051] Figure 5 is a distribution diagram of the original sequencing signal and the decomposition signal of the Ribo-seq sequencing results of the Huh7.5.1 cell transfected with PrcccDNA plasmid in Example 1, wherein, Figure 5 A in the above formula is the HBV translation original signal, Figure 5 B in the above formula is the decomposition and correction signal. DETAILED DESCRIPTION

[0052] The technical solutions of the present application are further described below by means of the accompanying drawings and examples.

[0053] Unless otherwise defined, the technical terms or scientific terms used in the present application shall have the usual meaning understood by those skilled in the art to which the present application belongs.

[0054] Embodiment one

[0055] The method for analyzing the translation level of HBV proteins based on Ribo-seq comprises the following steps:

[0056] Step S1, constructing a linear HBV reference genome sequence and optimizing its start point.

[0057] The use of a linear HBV reference genome sequence with the same length as the genome (1X) instead of an ultra-long (such as 1.2X or 2X) linear HBV reference genome sequence that would cause serious multiple alignment problems fundamentally avoids ambiguity in signal sources. The traditional start point with EcoRI enzyme cutting site will truncate the ORFs of polymerase protein (P) and surface antigen (S) protein. The present application selects 1801 nt (taking HBV ayw strain as an example) as the new start point of the linear HBV reference genome sequence, and the optimized position is the position that keeps the main overlapping open reading frames intact and avoids multiple alignment. On the optimized linear HBV reference genome sequence, the translation signals of the three main proteins of core protein (C), polymerase protein (P) and surface antigen (S) are all complete, and only the X protein (X) is slightly truncated, which is an optimized trade-off to ensure the accuracy of the overall analysis.

[0058] Step S2, aligning the Ribo-seq sequencing data to the linear HBV reference genome sequence to obtain the P-site signal of each nucleotide position; according to the reading frame attribution of each nucleotide position on the linear HBV reference genome sequence, separating the P-site signal into three reading frames to obtain the signal intensity distribution of each reading frame.

[0059] For any nucleotide position, its belonging reading frame is:

[0060] ;

[0061] Wherein, represents the reading frame, represents any nucleotide position, represents the modulo operation.

[0062] Step S3, for each HBV protein, in the pure region on its genome that only encodes a single protein, based on the signal intensity distribution of the three reading frames, calculate the median signal density of the protein and its signal leakage rate to other reading frames.

[0063] (1) Median signal density: represents the typical translation efficiency of the protein under no interference condition.

[0064] (2) Signal leakage rate: calculated by the ratio of the signal of other reading frame to the signal of the main reading frame in the "pure region". For example, for a protein with main reading frame , the leakage rate to other reading frame is calculated as:

[0065] ;

[0066] wherein, represents the signal leakage rate, represents the main reading frame, represents the other reading frame, represents the pure region on the genome that only encodes a single protein, represents the observed signal at any nucleotide position .

[0067] Step S4, in the mixed region of the overlapping protein (encoding multiple proteins), for the target reading frame of the target nucleotide position, the following operations are performed:

[0068] (1) The P-site signal on the target reading frame at this position is taken as the observed signal.

[0069] ;

[0070] wherein, represents the true translation signal of protein C, represents the true translation signal of protein P, represents the signal leakage, represents the background noise.

[0071] (2) Based on the median signal density and the signal leakage rate, the expected signal of the target protein A and the expected interference noise from the overlapping protein B are estimated.

[0072] ;

[0073] ;

[0074] wherein, represents the expected signal of the target protein A, represents the codon length of the overlapping region A, represents the median signal density of the target protein A, represents the expected interference noise of the overlapping protein B, represents the mixed region of the overlapping protein, represents the observed signal at any nucleotide position of the reading frame of the overlapping protein B, represents the signal leakage rate of the overlapping protein B to the target protein A.

[0075] (3) Calculate the signal-to-noise ratio.

[0076] ;

[0077] wherein, represents the signal-to-noise ratio.

[0078] (4) Select a signal correction strategy based on the ratio of the signal-to-noise ratio to a preset threshold value, and correct the observed signal of each reading frame. The preset threshold value , which can be determined by the user, and the default value in this embodiment is 3.

[0079] The signal correction strategy includes:

[0080] 1) When , it indicates that the target signal is strong enough to directly subtract the interference noise for signal correction.

[0081] ;

[0082] wherein, represents the corrected signal, represents the observed signal of any nucleotide position of the reading frame of the target protein A.

[0083] 2) When , it indicates that the target signal is weak, and a more robust density estimation method is used for signal correction.

[0084] ;

[0085] wherein, represents the Poisson distribution of the median signal density of the target protein A.

[0086] Step S5, output the corrected translation signal intensity of the HBV protein.

[0087] The present application will be further described below through specific examples.

[0088] To verify the accuracy of the algorithm of the present application, a calculation model capable of generating HBV Ribo-seq simulation data is constructed, which is prior art and will not be described in detail. The model can preset the real translation signal intensity of each protein, and introduce various random disturbances such as signal attenuation, local burst, signal leakage and background noise.

[0089] In multiple independent simulation tests, different simulation signals were tested, including 9 different cases, namely: 4 average protein expression, C high expression, P high expression, S high expression, X high expression, C low expression, P low expression, S low expression, X low expression. As shown in Table 1, the error of each protein signal level separated and estimated by the process of the application compared with the preset true signal level is always controlled within ± 25%. As shown in Figure 3 、 Figure 4 the simulation experiment results strongly prove that even in the extreme case of serious signal overlap and full of noise, the algorithm of the application can still reliably split and accurately quantify the translation level of each protein of HBV.

[0090] Table 1 Error level of HBV Ribo-seq signal decomposition and protein translation level estimation under different expression levels

[0091] ;

[0092] To verify the application value of the application in real biological samples, the GSE135860 data set (Ribo-seq data of Huh7.5.1 cells transfected with HBV cccDNA plasmid) was downloaded from the public database GEO.

[0093] After preprocessing the original data, it was applied to the model in this embodiment. As shown in Figure 5 , the results show that this embodiment successfully separates the original highly mixed and unreadable translation signal into the core protein (C), polymerase protein (P), surface antigen (S) and X protein (X) of HBV four proteins, revealing their respective independent and clear three-nucleotide periodic translation profiles.

[0094] The results of this preclinical data analysis prove that the method in this embodiment is effective and can accurately analyze real and complex biological sample data, which cannot be achieved by traditional analysis methods.

[0095] Embodiment Two

[0096] The system for analyzing HBV protein translation level based on Ribo-seq comprises:

[0097] (1) A data input module for constructing a linear HBV reference genome sequence and optimizing its starting point.

[0098] (2) a signal separation module, configured to align Ribo-seq sequencing data to a linear HBV reference genome sequence to obtain P-site signals of each nucleotide position; and separate the P-site signals into three reading frames according to the reading frame attribution of each nucleotide position on the linear HBV reference genome sequence to obtain signal intensity distribution of each of the three reading frames.

[0099] (3) a benchmark modeling module, configured to calculate, for each HBV protein, a median signal density and a signal leakage rate to other reading frames of the protein based on the signal intensity distribution of the three reading frames in a pure region on the genome of the protein that only encodes a single protein.

[0100] (4) a signal correction module, configured to correct an observed signal of a target reading frame based on the median signal density and the signal leakage rate in a mixed region of overlapping proteins; and specifically comprising the following units:

[0101] a signal observation unit, configured to take P-site signals on the target reading frame of a target nucleotide position as the observed signal.

[0102] a signal-to-noise ratio analysis unit, configured to estimate an expected signal of the target protein and an expected interference noise of the overlapping proteins based on the median signal density and the signal leakage rate, and calculate a signal-to-noise ratio.

[0103] a signal correction strategy execution unit, configured to select a signal correction strategy based on a ratio of the signal-to-noise ratio to a preset threshold to correct the observed signal.

[0104] (5) an output module, configured to output a corrected translation signal intensity of the HBV protein.

[0105] It should be noted that the contents not elaborated in the present application are all prior art and are well known to those skilled in the art.

[0106] Therefore, the present application adopts the above-mentioned method and system for analyzing HBV protein translation level based on Ribo-seq, constructs an optimized linear HBV reference genome sequence, and based on the median signal density, the signal leakage rate and the signal-to-noise ratio, accurately separates the mixed Ribo-seq signals to output the independent and accurate translation signal intensity of the HBV protein, which not only significantly improves the accuracy and reliability of the overlapping translation signal analysis, but also provides an effective tool for in-depth research on HBV translation regulation mechanism and anti-viral drug development.

[0107] It should be pointed out finally that the above examples are only used to illustrate the technical solutions of the present application but not to limit it, and although the present application has been described in detail with reference to the preferred embodiments, it should be understood by those skilled in the art that the technical solutions of the present application can still be modified or replaced equivalently, and these modifications or equivalent replacements should not make the modified technical solutions deviate from the spirit and scope of the technical solutions of the present application.

Claims

1. A method for analyzing HBV protein translation levels based on Ribo-seq, characterized in that, Includes the following steps: Step S1: Construct a linear HBV reference genome sequence and optimize its starting point; Step S2: Align the Ribo-seq sequencing data to the linear HBV reference genome sequence to obtain the P-site signal at each nucleotide position; according to the reading frame assignment of each nucleotide position on the linear HBV reference genome sequence, separate the P-site signal into three reading frames to obtain the signal intensity distribution of each of the three reading frames. Step S3: For each HBV protein, within a pure region on its genome that encodes only a single protein, calculate the median signal density of the protein and its signal leakage rate to other reading frames based on the signal intensity distribution of the three reading frames. Step S4: In the mixed region of overlapping proteins, for the target reading frame at the target nucleotide position, perform: The P-site signal on the target reading frame at this location is used as the observation signal; Based on the median signal density and signal leakage rate, the expected signal of the target protein and the expected interference noise of overlapping proteins are estimated, and the signal-to-noise ratio is calculated. The signal correction strategy is selected based on the ratio of signal-to-noise ratio to a preset threshold to correct the observed signal; Step S5: Output the corrected translation signal intensity of HBV protein.

2. The method for analyzing HBV protein translation levels based on Ribo-seq according to claim 1, characterized in that, In step S2, for any nucleotide position, its corresponding reading frame is: ; in, This indicates a reading box. Indicates any nucleotide position. This indicates the modulo operation.

3. The method for analyzing HBV protein translation levels based on Ribo-seq according to claim 1, characterized in that, In step S3, the formula for calculating the signal leakage rate is: ; in, Indicates the signal leakage rate. This indicates the main reading box. Indicates other reading boxes, This refers to a pure region on the genome that encodes only a single protein. Indicates any nucleotide position The observed signal.

4. The method for analyzing HBV protein translation levels based on Ribo-seq according to claim 3, characterized in that, In step S4, the formula for calculating the signal-to-noise ratio is: ; ; ; in, This indicates the expected signal for target protein A. Indicates the codon length of the overlapping region A. This represents the median signal density of target protein A. This represents the expected interference noise of overlapping protein B. This indicates the mixing region of overlapping proteins. Indicates any nucleotide position in the reading frame of overlapping protein B. The observed signal, This indicates the signal leakage rate from overlapping protein B to target protein A. This indicates the signal-to-noise ratio.

5. The method for analyzing HBV protein translation levels based on Ribo-seq according to claim 4, characterized in that, In step S4, the signal correction strategy includes: When the signal-to-noise ratio is greater than or equal to the preset threshold, interference noise is directly subtracted for signal correction. ; When the signal-to-noise ratio is less than a preset threshold, a density estimation method is used for signal correction. ; in, Indicates the signal after correction. This indicates any nucleotide position within the reading frame of target protein A. The observed signal, The Poisson distribution represents the median signal density of target protein A.

6. The method for analyzing HBV protein translation levels based on Ribo-seq according to claim 1, characterized in that, HBV proteins include at least one of the following: core protein, polymerase protein, surface antigen protein, and X protein.

7. A system for analyzing HBV protein translation levels based on Ribo-seq, characterized in that, Performing the method for resolving HBV protein translation levels based on Ribo-seq as described in any one of claims 1 to 6, comprising: The data input module is used to construct a linear HBV reference genome sequence and optimize its starting point; The signal separation module is used to align Ribo-seq sequencing data to a linear HBV reference genome sequence to obtain the P-site signal at each nucleotide position; based on the reading frame assignment of each nucleotide position in the linear HBV reference genome sequence, the P-site signal is separated into three reading frames to obtain the signal intensity distribution of each of the three reading frames. The baseline modeling module is used to calculate the median signal density of each HBV protein and its signal leakage rate to other reading frames based on the signal intensity distribution of three reading frames within a pure region on its genome that encodes only a single protein. The signal correction module is used to correct the observed signal of the target reading frame in the mixed region of overlapping proteins based on the median signal density and signal leakage rate. The output module is used to output the corrected translation signal intensity of the HBV protein.

8. The system for analyzing HBV protein translation levels based on Ribo-seq according to claim 7, characterized in that, The signal correction module includes: The signal observation unit uses the P-site signal on the target reading frame at the target nucleotide position as the observation signal; The signal-to-noise ratio (SNR) analysis unit estimates the expected signal of the target protein and the expected interference noise of overlapping proteins based on the median signal density and signal leakage rate, and calculates the SNR. The signal correction strategy execution unit selects a signal correction strategy based on the ratio of the signal-to-noise ratio to a preset threshold and corrects the observed signal.

Citation Information

Patent Citations

  • Method for regulating and controlling translation efficiency of mRNAs based on ribosome S1 protein acylation modification engineering

    CN115820699A

  • RNA replicons for multifunctional and efficient gene expression

    CN115927467A