A method for sequence quality assessment after magnetic bead-nucleic acid separation
By combining the base frequency, entropy value and low-quality region weighted evaluation methods, the problem of failure to comprehensively evaluate the quality of nucleic acid sequences in traditional methods is solved, and the accurate quality evaluation and processing strategies of nucleic acid sequences are realized, which improves the reliability and scientificity of the evaluation.
Patent Information
- Application Number
- CN202510733137.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-04
- Publication Date
- 2025-08-12
- Estimated Expiration
- 2045-06-04
AI Technical Summary
The traditional nucleic acid quality evaluation method fails to comprehensively consider the base distribution uniformity and entropy value after magnetic bead separation, resulting in insufficient identification of pollution and sequencing errors and lack of accurate quality evaluation.
Through the sequence integrity evaluation algorithm and the sequence consistency scoring algorithm, combined with base frequency, entropy value and low-quality region weighted evaluation, the comprehensive quality score of nucleic acid sequence is calculated, different quality levels are divided and corresponding processing strategies are formulated.
It realizes accurate identification of nucleic acid sequence pollution and sequencing errors, reduces misjudgment, and improves the reliability and scientificity of nucleic acid sequence quality evaluation.
Smart Images

Figure CN120260681B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of sequence quality assessment, and in particular to a method for sequence quality assessment after magnetic bead-nucleic acid separation. Background Art
[0002] Nucleic acid extraction and separation are critical steps in biomolecular analysis, particularly in genomics and transcriptomics. Nucleic acid quality directly impacts the accuracy and reliability of downstream analyses. Magnetic bead separation, a commonly used method for nucleic acid extraction, has been widely used for nucleic acid separation and purification due to its ease of use, high efficiency, and low contamination. However, despite its numerous advantages, quality assessment of nucleic acid samples after magnetic bead-nucleic acid separation remains an urgent challenge. Currently, traditional nucleic acid quality assessment methods, including UV spectrophotometry, electrophoresis, and fluorescent dyes, have demonstrated promising applications but exhibit significant limitations. These methods focus on single indicators, such as concentration, purity, and integrity, and fail to comprehensively assess nucleic acid sample quality. This is particularly true for nucleic acid samples after magnetic bead separation, where interference from the magnetic beads and potential nucleic acid variability make current quality assessment methods inadequate. Therefore, developing a method for sequence quality assessment after magnetic bead-nucleic acid separation would provide researchers with timely and accurate quality control data, ensuring experimental reliability and reproducibility, and possessing significant scientific and clinical significance.
[0003] However, the above-mentioned existing method for sequence quality assessment after magnetic bead-nucleic acid separation still has the following technical problems: due to the huge amount of data and the continuous evolution of sequencing technology, nucleic acid sequence quality assessment often relies on a single standard and lacks a comprehensive and accurate quality assessment method; contamination or sequencing errors have a great impact on nucleic acid sequence data, but traditional quality assessment methods fail to effectively combine the uniformity of base distribution and the degree of contamination, and are prone to overlooking potential contamination or sequencing errors. Summary of the Invention
[0004] The present invention provides a method for sequence quality assessment after magnetic bead-nucleic acid separation, so as to solve the technical problems that traditional methods only rely on quality score thresholds when identifying contamination and errors in sequences, but fail to comprehensively consider the base distribution uniformity and entropy value of the sequence; existing technologies mostly rely on simple low-quality base counting, but ignore the impact of low-quality regions and the influence of sequence length on quality scores; traditional sequence quality assessment methods fail to provide a comprehensive quality score and have no targeted solution strategies when quality problems arise.
[0005] The present invention provides a method for sequence quality assessment after magnetic bead-nucleic acid separation, which specifically includes the following technical solutions:
[0006] A method for sequence quality assessment after magnetic bead-nucleic acid separation, comprising the following steps:
[0007] S1. Sequence the nucleic acid sample generated after magnetic bead separation to obtain a nucleic acid sequence set and its corresponding quality score; preprocess the nucleic acid sequence set to obtain a preprocessed nucleic acid sequence set; and calculate the nucleic acid sequence integrity score based on the preprocessed nucleic acid sequence set using a sequence integrity assessment algorithm;
[0008] S2. Based on the preprocessed nucleic acid sequence set and its corresponding quality scores, a sequence consistency score is calculated using a sequence consistency scoring algorithm. An overall quality score is calculated by combining the completeness score and consistency score of the nucleic acid sequences, and the nucleic acid sequences are classified into different quality levels based on the overall quality score.
[0009] Preferably, the S1 specifically includes:
[0010] Extract nucleic acid sequences one by one and record the length of the nucleic acid sequence, that is, the number of bases in the nucleic acid sequence. Use a sequencer to read the base symbols in the nucleic acid sequence, and mark the bases with a quality score lower than the quality score threshold as unknown bases. Each nucleic acid sequence is composed of bases. For nucleic acid sequences containing unknown bases, calculate the effective length of the nucleic acid sequence.
[0011] Preferably, the S1 specifically includes:
[0012] The sequence integrity assessment algorithm is based on the analysis of the base frequencies of the nucleic acid sequence. The frequency of occurrence of each base will affect the integrity of the nucleic acid sequence. When the base frequency is abnormal, there will be signs of contamination and sequencing errors.
[0013] Preferably, the S1 specifically includes:
[0014] In the process of implementing the sequence integrity assessment algorithm, the proportion of unknown bases in each nucleic acid sequence is calculated, and the proportion of unknown bases in the nucleic acid sequence is introduced into the integrity assessment of the nucleic acid sequence to quantify the degree of base contamination.
[0015] Preferably, the S1 specifically includes:
[0016] In the implementation of the sequence integrity assessment algorithm, the entropy value of the nucleic acid sequence is introduced as a measure of the complexity of the nucleic acid sequence, and the nucleic acid sequence integrity score is calculated by combining the entropy value of the nucleic acid sequence with the contamination ratio.
[0017] Preferably, the S2 specifically includes:
[0018] The sequence consistency scoring algorithm identifies all continuous low-quality regions based on the quality scores corresponding to the bases in each nucleic acid sequence, and for each low-quality region, calculates the length and average quality score of the low-quality region.
[0019] Preferably, the S2 specifically includes:
[0020] The sequence consistency scoring algorithm weights the length of the low-quality region and the average quality score to calculate the impact value of the low-quality region, and summarizes the impact values of all low-quality regions to obtain a total low-quality region weighted score.
[0021] Preferably, the S2 specifically includes:
[0022] The sequence consistency scoring algorithm performs normalization processing by combining the effective length of the nucleic acid sequence, so that the total weighted score of the low-quality region is associated with the effective length of the nucleic acid sequence, and introduces a penalty factor to adjust the influence of the low-quality region in the nucleic acid sequence consistency score.
[0023] Preferably, the S2 specifically includes:
[0024] The comprehensive quality score is calculated by combining the integrity score and consistency score of the nucleic acid sequence. After obtaining the comprehensive quality score of each nucleic acid sequence, the nucleic acid sequence is divided into different quality levels according to the comprehensive quality score, and a corresponding processing strategy is formulated for each quality level of nucleic acid sequence.
[0025] The beneficial effects of the technical solution of the present invention are:
[0026] 1. By introducing base frequency analysis, contamination penalty items, and entropy calculation of nucleic acid sequences, the contamination and sequencing errors in nucleic acid sequences can be accurately reflected. In particular, the entropy-based contamination penalty can adaptively adjust the penalty intensity for contamination according to the complexity of the nucleic acid sequence, effectively avoiding the false positives in traditional methods. It can also give lighter penalties to highly complex nucleic acid sequences when the contamination is relatively light, thereby reducing misjudgments.
[0027] 2. Weighted evaluation based on low-quality regions: By identifying continuous low-quality regions and weighting them based on their length and average quality score, this method can accurately reflect the overall quality of nucleic acid sequences. This avoids the misjudgment that may result from traditional methods that simply rely on the uniformity of base distribution, and improves the ability to assess the authenticity and reliability of nucleic acid sequences.
[0028] 3. Based on the calculation results of the comprehensive quality score, the quality of each nucleic acid sequence can be graded, and specific processing strategies can be formulated for nucleic acid sequences of different quality levels, effectively improving the scientificity and practicality of data processing. BRIEF DESCRIPTION OF THE DRAWINGS
[0029] Figure 1 This is a flow chart of a method for sequence quality assessment after magnetic bead-nucleic acid separation according to the present invention. DETAILED DESCRIPTION
[0030] In order to further illustrate the technical means and effects adopted by the present invention to achieve the predetermined purpose of the invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts shall fall within the scope of protection of the present invention.
[0031] Unless defined otherwise, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention belongs.
[0032] The following describes in detail a method for sequence quality assessment after magnetic bead-nucleic acid separation provided by the present invention with reference to the accompanying drawings.
[0033] Refer to the attached Figure 1 , which shows a flow chart of a method for sequence quality assessment after magnetic bead-nucleic acid separation provided by one embodiment of the present invention, the method comprising the following steps:
[0034] S1. Sequencing the nucleic acid sample generated after magnetic bead separation to obtain a nucleic acid sequence set and its corresponding quality score; preprocessing the nucleic acid sequence set to obtain a preprocessed nucleic acid sequence set, and calculating the nucleic acid sequence integrity score based on the preprocessed nucleic acid sequence set using a sequence integrity assessment algorithm;
[0035] The nucleic acid sample generated after magnetic bead separation is sequenced by a sequencer to generate a FASTQ format file. The FASTQ format file contains a nucleic acid sequence set and its corresponding Phred Quality Score (used to measure the sequencing reliability of each base);
[0036] Preprocess the nucleic acid sequence set to obtain a preprocessed nucleic acid sequence set ; The pretreatment method is as follows:
[0037] Extract nucleic acid sequences one by one , and record the length of the nucleic acid sequence, that is, the number of bases in the nucleic acid sequence , use the sequencer to directly read the base symbols in the nucleic acid sequence, and mark the bases with quality scores lower than the quality score threshold (i.e., bases with a higher probability of sequencing errors) as " ”, Indicates unknown bases. The quality score threshold can be set according to the specific implementation scenario and is not limited here. Each nucleic acid sequence is composed of bases. ,support and For sequences containing " "Nucleic acid sequence, calculate the effective length of the nucleic acid sequence , " ” number;
[0038] In order to evaluate the integrity of the nucleic acid sequence obtained after magnetic bead-nucleic acid separation, the nucleic acid sequence integrity score was calculated based on the pre-processed nucleic acid sequence set using the sequence integrity assessment algorithm;
[0039] The sequence integrity assessment algorithm is based on the analysis of the base frequencies of the nucleic acid sequence. The frequency of occurrence of each base will affect the integrity of the nucleic acid sequence. If the base frequency is abnormal, it may indicate contamination or sequencing errors.
[0040] The sequence integrity assessment algorithm calculates the number of " ratio, and the " "The ratio is introduced into the integrity assessment of nucleic acid sequences to quantify the degree of base contamination;
[0041] To further improve the integrity assessment of nucleic acid sequences, the entropy value of nucleic acid sequences was introduced as a measure of nucleic acid sequence complexity. Nucleic acid sequences with high entropy values indicate a more uniform base distribution, that is, a higher complexity of the nucleic acid sequence, which means that the nucleic acid sequence has higher biological reliability. Conversely, nucleic acid sequences with low entropy values indicate a large deviation in base distribution, which may contain systematic errors or contamination. By combining the entropy value of nucleic acid sequences with the contamination ratio, nucleic acid sequences with higher entropy values are made more sensitive to contamination, reflecting that contamination will have a greater impact on nucleic acid sequences.
[0042] The calculation formula for the nucleic acid sequence integrity score is:
[0043] ,
[0044] in, Indicates the The integrity score of the nucleic acid sequence; Indicates the bases in the nucleic acid sequence Perform sum operations on sets; Indicates the bases in a nucleic acid sequence Frequency of occurrence; Indicates the The effective length of a nucleic acid sequence; It represents the pollution penalty term, which combines the pollution ratio with the entropy value and controls the degree of penalty in the form of a power exponent, that is, the penalty for nucleic acid sequences with low pollution and high entropy is lighter, and the penalty for nucleic acid sequences with high pollution and low entropy is stronger; Indicates the The entropy value of the base distribution of a nucleic acid sequence. The entropy value quantifies the uniformity of the distribution by calculating the uncertainty of the base frequency in the nucleic acid sequence. The higher the entropy value, the more uniform the base distribution in the nucleic acid sequence. Conversely, a low entropy value indicates a biased base distribution. The calculation formula is: The method for calculating the entropy value is a technical means well known to those skilled in the art and will not be described in detail here;
[0045] By comprehensively considering the uniformity of base distribution and the contamination ratio, the integrity of the nucleic acid sequence can be accurately assessed;
[0046] S2. Based on the pre-processed nucleic acid sequence set and its corresponding quality scores, a sequence consistency score is calculated using a sequence consistency scoring algorithm; a comprehensive quality score is calculated by integrating the integrity score and consistency score of the nucleic acid sequence, and the nucleic acid sequence is divided into different quality levels according to the comprehensive quality score;
[0047] Based on the pre-processed nucleic acid sequence set and its corresponding quality score, a sequence consistency score is calculated using a sequence consistency scoring algorithm;
[0048] The sequence identity scoring algorithm identifies all continuous low-quality regions based on the quality scores corresponding to the bases in each nucleic acid sequence. The low-quality region refers to a region within a certain length range (3 consecutive bases or more) with a quality score below a set quality score threshold. For each low-quality region, the length and average quality score of the low-quality region are calculated, and the low-quality region length and the average quality score are weighted to calculate the impact value of the low-quality region. For each nucleic acid sequence, the impact values of all low-quality regions are summarized to obtain a total low-quality region weighted score. To avoid the influence of the effective length of the nucleic acid sequence on the nucleic acid sequence identity score, a normalization process is performed so that the total low-quality region weighted score is associated with the effective length of the nucleic acid sequence. A penalty factor is introduced to adjust the degree of influence of the low-quality region in the nucleic acid sequence identity score.
[0049] The calculation formula for the nucleic acid sequence consistency score is:
[0050] ,
[0051] in, Indicates the The consistency score of nucleic acid sequences is used to measure the impact of low-quality regions in nucleic acid sequences; Represents a penalty factor, which is used to adjust the impact of low-quality regions on the nucleic acid sequence consistency score. It can be set according to the specific implementation scenario and is not limited here; Indicates the The number of low-quality regions in a nucleic acid sequence, where a low-quality region is defined as a fragment with a quality score of at least three consecutive bases below a quality score threshold; Indicates the The length weight of the low-quality region is defined as The ratio of the length of the low-quality region to the effective length of the nucleic acid sequence; Indicates the The average quality score of low-quality areas;
[0052] By introducing a weighted assessment of low-quality regions, we avoid misjudging single low-quality bases in traditional methods. Traditional methods tend to oversimplify and rely on the uniformity of base distribution. However, base distribution deviations do not necessarily indicate quality issues in the nucleic acid sequence. The presence of low-quality regions often indicates serious errors in the sequencing process. Through precise assessment methods, the authenticity and reliability of nucleic acid sequences can be effectively reflected.
[0053] The comprehensive quality score is calculated by combining the integrity score and consistency score of the nucleic acid sequence, which takes into account not only the reliability of the nucleic acid sequence itself but also the impact of the contamination level. The range of the comprehensive quality score is arrive The closer the comprehensive quality score is, The higher the quality of the nucleic acid sequence, the lower the comprehensive quality score. It means that the quality of the nucleic acid sequence is poor;
[0054] The formula for calculating the comprehensive quality score is:
[0055] ,
[0056] in, Indicates the Comprehensive quality score of nucleic acid sequences; It represents the maximum value of the integrity score among all nucleic acid sequences and is used to normalize the integrity score;
[0057] After obtaining the comprehensive quality score of each nucleic acid sequence, the nucleic acid sequence is divided into different quality levels according to the comprehensive quality score, such as high quality, medium quality, and low quality. The nucleic acid sequence classification can be set according to the specific implementation scenario and is not limited here. Corresponding processing strategies are formulated for nucleic acid sequences of each quality level:
[0058] For high-quality nucleic acid sequences, it is recommended to use them directly for downstream analysis, such as genome assembly and variant detection, without additional processing;
[0059] For nucleic acid sequences of medium quality, the processing is based on the different performance of the nucleic acid sequence integrity score and consistency score. If the nucleic acid sequence integrity score is lower than the consistency score, it is recommended to optimize the magnetic bead separation process, such as increasing the number of washes or adjusting the buffer concentration, and then re-sequencing. If the nucleic acid sequence consistency score is lower than the integrity score, it is recommended to check the performance of the sequencer, such as calibrating the optical system or replacing the reagents.
[0060] For low-quality nucleic acid sequences, it is recommended to discard the low-quality nucleic acid sequences, re-prepare the nucleic acid samples and re-sequence them to ensure the reliability of subsequent data analysis.
[0061] The quality of each nucleic acid sequence was quantified through a comprehensive quality score, and a specific operation plan was formulated for each quality level through a grading strategy, thereby effectively improving the scientificity and practicality of data processing.
[0062] In summary, a method for sequence quality assessment after magnetic bead-nucleic acid separation was completed.
[0063] The order in which the embodiments of the invention are presented is for illustrative purposes only and does not necessarily represent the superiority or inferiority of the embodiments. The processes depicted in the accompanying drawings do not necessarily require the specific order or sequential order shown to achieve the desired results. In certain embodiments, multitasking and parallel processing are also possible or may be advantageous.
[0064] The various embodiments in this specification are described in a progressive manner, and the same or similar parts between the various embodiments can be referred to each other. Each embodiment focuses on the differences from other embodiments.
[0065] The above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit the same. Although the present invention has been described in detail with reference to the above embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the above embodiments, or make equivalent replacements for some of the technical features therein. These modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be included in the scope of protection of the present invention.
Claims
1. A method for sequence quality assessment after magnetic bead-nucleic acid separation, characterized in that: The following steps are involved: S1. Sequence the nucleic acid sample generated after magnetic bead separation to obtain a nucleic acid sequence set and its corresponding quality score; preprocess the nucleic acid sequence set to obtain a preprocessed nucleic acid sequence set. Based on the preprocessed nucleic acid sequence set, calculate the nucleic acid sequence integrity score using a sequence integrity assessment algorithm. The specific formula is: , in, Indicates the The integrity score of the nucleic acid sequence; Indicates the bases in the nucleic acid sequence Perform sum operations on sets; Indicates the bases in a nucleic acid sequence Frequency of occurrence; Indicates the The effective length of the nucleic acid sequence, ; It represents the pollution penalty term, which combines the pollution ratio with the entropy value and controls the degree of penalty in the form of a power exponent, that is, the penalty for nucleic acid sequences with low pollution and high entropy is lighter, and the penalty for nucleic acid sequences with high pollution and low entropy is stronger; Indicates the The entropy value of the base distribution of a nucleic acid sequence. The entropy value quantifies the uniformity of the distribution by calculating the uncertainty of the base frequency in the nucleic acid sequence. The higher the entropy value, the more uniform the base distribution in the nucleic acid sequence. Conversely, a low entropy value indicates a biased base distribution. The calculation formula is: ; Unknown bases in nucleic acid sequences the number of Indicates the The number of bases in a nucleic acid sequence; S2. Based on the preprocessed nucleic acid sequence set and its corresponding quality score, a sequence consistency score is calculated using a sequence consistency scoring algorithm. The specific formula is: , in, Indicates the The consistency score of nucleic acid sequences is used to measure the impact of low-quality regions in nucleic acid sequences; represents the penalty factor, which is used to adjust the impact of low-quality regions on the nucleic acid sequence consistency score; Indicates the The number of low-quality regions in a nucleic acid sequence, where a low-quality region is defined as a fragment with a quality score of at least three consecutive bases below a quality score threshold; Indicates the The length weight of the low-quality region is defined as The ratio of the length of the low-quality region to the effective length of the nucleic acid sequence; Indicates the The average quality score of low-quality areas; The comprehensive quality score is calculated by integrating the integrity score and consistency score of the nucleic acid sequence, and the nucleic acid sequence is divided into different quality levels according to the comprehensive quality score. The range of the comprehensive quality score is arrive The closer the comprehensive quality score is, The higher the quality of the nucleic acid sequence, the lower the comprehensive quality score. It means that the quality of the nucleic acid sequence is poor; The formula for calculating the comprehensive quality score is: , in, Indicates the Comprehensive quality score of nucleic acid sequences; It represents the maximum value of the integrity score among all nucleic acid sequences and is used to normalize the integrity score.
2. A method for sequence quality assessment after magnetic bead-nucleic acid separation according to claim 1, characterized in that: Said S1 specifically includes: Extract nucleic acid sequences one by one and record the length of the nucleic acid sequence, that is, the number of bases in the nucleic acid sequence. Use a sequencer to read the base symbols in the nucleic acid sequence, and mark the bases with a quality score lower than the quality score threshold as unknown bases. Each nucleic acid sequence is composed of bases. For nucleic acid sequences containing unknown bases, calculate the effective length of the nucleic acid sequence.
3. The method for sequence quality assessment after magnetic bead-nucleic acid separation according to claim 1, characterized in that: Said S1 specifically includes: The sequence integrity assessment algorithm is based on the analysis of the base frequencies of the nucleic acid sequence. The frequency of occurrence of each base will affect the integrity of the nucleic acid sequence. When the base frequency is abnormal, there will be signs of contamination and sequencing errors.
4. A method for sequence quality assessment after magnetic bead-nucleic acid separation according to claim 3, characterized in that: Said S1 specifically includes: In the process of implementing the sequence integrity assessment algorithm, the proportion of unknown bases in each nucleic acid sequence is calculated, and the proportion of unknown bases in the nucleic acid sequence is introduced into the integrity assessment of the nucleic acid sequence to quantify the degree of base contamination.
5. The method for sequence quality assessment after magnetic bead-nucleic acid separation according to claim 4, characterized in that: Said S1 specifically includes: In the implementation of the sequence integrity assessment algorithm, the entropy value of the nucleic acid sequence is introduced as a measure of the complexity of the nucleic acid sequence, and the nucleic acid sequence integrity score is calculated by combining the entropy value of the nucleic acid sequence with the contamination ratio.
6. The method for sequence quality assessment after magnetic bead-nucleic acid separation according to claim 1, characterized in that: Said S2 specifically includes: The sequence consistency scoring algorithm identifies all continuous low-quality regions based on the quality scores corresponding to the bases in each nucleic acid sequence, and for each low-quality region, calculates the length and average quality score of the low-quality region.
7. A method for sequence quality assessment after magnetic bead-nucleic acid separation according to claim 6, characterized in that: Said S2 specifically includes: The sequence consistency scoring algorithm weights the length of the low-quality region and the average quality score to calculate the impact value of the low-quality region, and summarizes the impact values of all low-quality regions to obtain a total low-quality region weighted score.
8. The method for sequence quality assessment after magnetic bead-nucleic acid separation according to claim 7, characterized in that: Said S2 specifically includes: The sequence consistency scoring algorithm performs normalization processing by combining the effective length of the nucleic acid sequence, so that the total weighted score of the low-quality region is associated with the effective length of the nucleic acid sequence, and introduces a penalty factor to adjust the influence of the low-quality region in the nucleic acid sequence consistency score.
9. The method for sequence quality assessment after magnetic bead-nucleic acid separation according to claim 8, characterized in that: Said S2 specifically includes: The comprehensive quality score is calculated by combining the integrity score and consistency score of the nucleic acid sequence. After obtaining the comprehensive quality score of each nucleic acid sequence, the nucleic acid sequence is divided into different quality levels according to the comprehensive quality score, and a corresponding processing strategy is formulated for each quality level of nucleic acid sequence.