Method for processing probe board and electronic equipment

By comparing and screening the sequencing results of the probe plate with the reference sequence, multiple indicators are generated to evaluate the synthesis quality and quantitative uniformity of the probe plate. This solves the problems of high cost, limited detection range and difficulty in guaranteeing uniformity in existing DNA methylation detection methods, and achieves efficient quality control and accurate gene detection.

CN121789778APending Publication Date: 2026-04-03SHANGHAI WEIHE MEDICAL LAB CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-25
Publication Date
2026-04-03

AI Technical Summary

Technical Problem

Existing DNA methylation detection methods such as WGBS, RRBS, and (h)MeDIP-Seq are characterized by high cost, complex data analysis, limited detection range, or large batch-to-batch variations. Furthermore, the uniformity of methylation probe synthesis and quantification is difficult to guarantee, affecting capture efficiency and the accuracy of sequencing results.

Method used

The first alignment result is determined by sequencing results based on the probe plate and the reference sequence, and the second alignment result is generated based on the screened alignment result. Multiple indicators are used to evaluate the synthesis quality and quantitative uniformity of the probe plate, providing a basis for quality control and verification.

Benefits of technology

This enables accurate assessment of the synthesis quality and quantitative uniformity of probe plates, ensuring the reliability and accuracy of gene detection and improving the quality control capabilities of probe plates before use.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121789778A_ABST
    Figure CN121789778A_ABST
Patent Text Reader

Abstract

Embodiments of the present disclosure relate to a method for processing a probe plate. The method includes determining a first comparison result based on a sequencing result of a probe plate and a reference sequence, the reference sequence being a sequence for synthesizing a probe of the probe plate. The method further comprises the step of determining a second comparison result based on screening of the first comparison result. And determining an index for the sequencing result based on the first comparison result and the second comparison result. Through the method, an accurate and comprehensive evaluation basis can be provided for the quality and quantitative uniformity of the synthesized probe plate, and a basis is provided for the quality control of the probe plate.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates generally to the field of bioinformatics, and more specifically to methods and electronic devices for probe plates. Background Technology

[0002] DNA cytosine base methylation is an important epigenetic modification and one of the most stable and abundant types of DNA modification. DNA methylation plays a crucial role in gene expression regulation and chromatin remodeling; methylation modifications in different gene regions or specific sites are closely related to embryonic development, disease occurrence, and progression. The state of DNA methylation can also reflect the long-term effects of the external environment on the organism.

[0003] Therefore, detailed analysis of cytosine methylation status in the genome has significant scientific and clinical application value, especially in the fields of disease mechanism research and early detection and diagnosis of major diseases such as cancer. Currently, several DNA methylation detection methods have been proposed, such as whole genome bisulfite sequencing (WGBS), reduced representation bisulfite sequencing (RRBS), hydroxymethylated DNA immunoprecipitation sequencing (h)MeDIP-Seq, and target bisulfite sequencing (TBS). Summary of the Invention

[0004] According to exemplary embodiments of this disclosure, a method, apparatus, electronic device, computer-readable storage medium, and computer program product for processing probe panels are provided. These provide a basis for quality control and uniformity verification of probe panels (also known as panels) prior to use.

[0005] According to a first aspect of this disclosure, a method for processing a probe plate is provided. The method includes determining a first alignment result based on sequencing results from the probe plate and a reference sequence, the reference sequence being the sequence of probes used to synthesize the probe plate. The method further includes determining a second alignment result based on screening the first alignment result. The method also includes determining an indicator for the sequencing results based on the first and second alignment results.

[0006] According to a second aspect of this disclosure, an apparatus for processing a probe plate is provided. The apparatus includes an alignment unit configured to determine a first alignment result based on sequencing results of the probe plate and a reference sequence, the reference sequence being a sequence of probes used to synthesize the probe plate; a screening unit configured to determine a second alignment result based on screening of the first alignment result; and an indicator determination unit configured to determine an indicator for the sequencing results based on the first and second alignment results.

[0007] In a third aspect of this disclosure, an electronic device is provided, including at least one processor; and a memory for storing at least one program, which, when executed by the at least one processor, causes the at least one processor to implement the method according to a first aspect of this disclosure.

[0008] In a fourth aspect of this disclosure, a computer-readable storage medium is provided having a computer program stored thereon that, when executed by a processor, implements the method according to a first aspect of this disclosure.

[0009] In a fifth aspect of this disclosure, a computer program product is provided. This computer program product includes a computer program that, when executed by a processor, implements the method according to a first aspect of this disclosure.

[0010] It should be understood that the content described in this section is not intended to limit the key or essential features of the embodiments of this disclosure, nor is it intended to restrict the scope of this disclosure. Other features of this disclosure will become readily apparent from the following description. Attached Figure Description

[0011] Figure 1 The illustration shows a schematic diagram of an example environment 100 that can be applied to some embodiments of the present disclosure;

[0012] Figure 2 A schematic flowchart of a method for processing a probe plate according to some embodiments of the present disclosure is shown;

[0013] Figure 3 The illustration shows a schematic diagram of a method for generating a probe quality control report according to some embodiments of the present disclosure;

[0014] Figures 4-7 The illustration shows a schematic diagram of the visualization output results used to evaluate the synthesis quality of the probe plate;

[0015] Figure 8 The illustration shows a schematic block diagram of an apparatus for processing a probe plate according to some embodiments of the present disclosure;

[0016] Figure 9A schematic block diagram of an example device 900 that can be used to implement embodiments of the present disclosure is shown; Detailed Implementation

[0017] It is understood that the data involved in this technical solution (including but not limited to the data itself, the acquisition or use of the data) shall comply with the requirements of relevant laws, regulations and related provisions.

[0018] It is understood that before using the technical solutions disclosed in the various embodiments of this disclosure, users should be informed of the types, scope of use, and usage scenarios of the personal information involved in this disclosure in an appropriate manner in accordance with relevant laws and regulations, and user authorization should be obtained.

[0019] For example, upon receiving a user's proactive request, a prompt message is sent to the user to explicitly inform them that the requested operation will require the acquisition and use of the user's personal information. This allows the user to independently choose whether to provide personal information to the software or hardware, such as the electronic device, application, server, or storage medium performing the operations of this disclosed technical solution, based on the prompt message.

[0020] As an optional but non-limiting implementation, in response to a user's active request, sending a prompt message to the user can be done via a pop-up window, where the prompt message can be presented in text format. Furthermore, the pop-up window can also include a selection control allowing the user to choose "agree" or "disagree" to provide personal information to the electronic device.

[0021] It is understood that the above notification and user authorization process are merely illustrative and do not constitute a limitation on the implementation of this disclosure. Other methods that comply with relevant laws and regulations may also be applied to the implementation of this disclosure.

[0022] Embodiments of this disclosure will now be described in more detail with reference to the accompanying drawings. While some embodiments of this disclosure are shown in the drawings, it should be understood that this disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of this disclosure. It should be understood that the accompanying drawings and embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of protection of this disclosure.

[0023] In the description of embodiments of this disclosure, the term "comprising" and similar terms should be understood as open-ended inclusion, i.e., "including but not limited to". The term "based on" should be understood as "at least partially based on". The term "one embodiment" or "the embodiment" should be understood as "at least one embodiment". The terms "first", "second", etc., may refer to different or the same objects. Other explicit and implicit definitions may also be included below.

[0024] Among DNA methylation detection methods, WGBS technology has a very wide coverage, but the sequencing cost is high and the data analysis is complex, making it unsuitable for large-scale clinical applications; RRBS technology relies on restriction endonucleases, which cannot cover all CpG-enriched regions, resulting in a limited detection range; and (h)MeDIP-Seq technology relies on the selectivity and specificity of antibodies, which is prone to batch-to-batch variations and biases towards highly methylated regions.

[0025] Target-BS utilizes specific oligonucleotide (DNA-oligo) probes to enrich target gene regions or specific CpG sites through hybridization capture, followed by sulfite sequencing or enzymatic conversion sequencing to achieve targeted detection of methylation signals in disease-related regions. In Target-BS technology, the synthesis and quantification uniformity of methylation probes present the following challenges: the overall base complexity of the methylated sequence decreases, easily leading to consecutive identical bases or low-complexity fragments. These sequences are prone to reduced efficiency and quality inhomogeneity during chemical synthesis. Furthermore, gene assemblies (panels) used for methylation capture typically consist of thousands of probes. Quantitative mixing of individual probes is required after synthesis. Due to differences in probe concentration and synthesis efficiency, the uniformity of the molar amount of probes in the final panel is difficult to guarantee, thus affecting capture efficiency and the quantitative accuracy of sequencing results.

[0026] To address the aforementioned problems and other potential issues, embodiments of this disclosure propose a method for processing probe plates. In embodiments of this disclosure, a computing device determines a first alignment result based on the sequencing results of the probe plate and a reference sequence, where the reference sequence is the sequence of probes used to synthesize the probe plate. Next, the computing device can determine a second alignment result based on screening the first alignment result. Then, the computing device determines metrics for the sequencing results based on the first and second alignment results. This method can evaluate the synthesis quality and quantitative uniformity of the probe plate, providing a basis for quality control and uniformity verification before the probe plate is used.

[0027] The embodiments of this disclosure will now be described in further detail with reference to the accompanying drawings. Figure 1 A schematic diagram of an example environment 100 applicable to some embodiments of the present disclosure is shown. In environment 100, computing device 102 can be used to determine metrics 112 for sequencing results 104 against a probe plate.

[0028] Examples of computing device 102 include, but are not limited to, personal computers, server computers, handheld or laptop devices, mobile devices (such as mobile phones, personal digital assistants (PDAs), media players, etc.), multiprocessor systems, consumer electronics, minicomputers, mainframe computers, and distributed computing environments that include any of the above systems or devices.

[0029] like Figure 1 As shown, computing device 102 can acquire sequencing results 104 and reference sequences 106 of the probe plate, where the reference sequence 106 is the sequence of the probes used to synthesize the probe plate. In some embodiments, the synthesized probe plate can be sequenced using next-generation sequencing (NGS) to obtain sequencing results 104. In some embodiments, the reference sequence 106 can be a probe. In some embodiments, the probes are selected from a probe database. In some embodiments, the sequence of each probe is stored as an independent entry in the probe database. In some embodiments, the unique identifier of the probe (Probe ID) can be used as the sequence header. In this way, the sequence of each probe can be used as an independent reference sequence 106 in alignment and statistical analysis.

[0030] In some embodiments of this disclosure, the computing device 102 can perform quality control on the sequencing results 104 of the probe plate. In some embodiments, low-quality reads in the sequencing results 104 of the probe plate can be automatically filtered by software. In some embodiments, specific reads in the sequencing results 104 of the probe plate can be filtered by a preset threshold, such as a preset quality score (Q-score) threshold. For example, low-quality bases with a quality score below Q20 or Q30 can be removed, and adapter sequences can be truncated. In some embodiments, effective reads of appropriate length and high overall quality can be retained. For example, adapter and primer sequences introduced during library construction can be identified and truncated through sequence alignment to prevent them from interfering with subsequent analysis; finally, effective clean reads (clean reads) with lengths within the expected design range (e.g., length not less than 80% of the probe design length), no adapter residues, and high overall base quality are screened and retained as a reliable data basis for subsequent sequence alignment, error rate statistics, diversity assessment, and probe synthesis quality analysis.

[0031] In some embodiments of this disclosure, computing device 102 can determine a first alignment result 108 based on sequencing results 104 of the probe plate and a reference sequence 106, wherein the reference sequence 106 is the sequence of probes used to synthesize the probe plate. In some embodiments, computing device 102 can align quality-controlled reads to a pre-constructed probe sequence database. In some embodiments, computing device 102 can generate a binary alignment / map (BAM) file. In some embodiments, computing device 102 can sort and index the alignment result file.

[0032] In some embodiments of this disclosure, the computing device 102 can filter the first alignment result 108 to obtain a second alignment result 110. In some embodiments, the computing device 102 can filter the first alignment result 108 based on at least one of the mapping quality value (MAPQ), multiple alignment information, number of mismatches, and number of insertion-deletion (InDel) markers to obtain the second alignment result 110. For example, the computing device 102 can use reads in the first alignment result 108 with a MAPQ greater than a preset threshold (e.g., MAPQ>10) as candidate reads in the second alignment result 110.

[0033] In some embodiments, the computing device 102 can remove reads belonging to multiple alignments from the first alignment result 108, thereby filtering the first alignment result 108. Multiple alignments indicate that a single read in the sequencing result has two or more alignment records in the first alignment result 108. For example, if a single read has the same or very high similarity alignment quality value with two or more different reference sequences 106, its unique source cannot be reliably distinguished, resulting in multiple alignment records in the first alignment result 108. In some embodiments, this is especially likely to occur when probe sequences are highly similar, such as when two probes in a probe plate are both 120 bp in length and differ only in 2-3 base positions. For short reads covering these differing regions, the computing device 102 may have difficulty determining which probe it originates from, thus outputting multiple equally preferred alignment results. By filtering out reads from multiple alignments, it can be ensured that in the second alignment result 110, the read can be clearly traced back to a specific probe. In some embodiments, computing device 102 can remove abnormal reads that are obviously mismatched or have an InDel count higher than a threshold.

[0034] In some embodiments, computing device 102 may generate metrics 112 for sequencing results 104 of probe plate based on first alignment result 108 and second alignment result 110, respectively. In some embodiments, metrics 112 may include at least one of probe coverage information (e.g., number of probe reads covered, coverage, average sequencing depth, sequencing depth distribution, insert length distribution, etc.), alignment error information (e.g., number of probe mismatches and number of end mismatches, etc.) and synthesis integrity.

[0035] In some embodiments, the computing device 102 can generate and output visualized statistical results, such as charts, based on the indicator 112. In some embodiments, the computing device 102 can also analyze the indicator 112 to generate results in the form of corresponding analysis reports, evaluation conclusions, quality scores, etc. In other embodiments, the computing device 102 may only output visualized statistical results, while the final analysis work may be completed manually or by other means, and this disclosure does not impose any restrictions on this.

[0036] The above methods can provide accurate and comprehensive evaluation criteria for the synthesis quality and quantitative uniformity of gene combinations, ensuring the uniformity and reliability of gene combinations and providing a good starting point for downstream tasks such as gene detection.

[0037] The above combination Figure 1 A schematic diagram of an example environment 100 that can be applied to some embodiments of the present disclosure is described below, in conjunction with... Figure 2 A schematic flowchart describing a method for processing a probe plate according to some embodiments of the present disclosure. Figure 2 Method 200 in the middle can be derived from Figure 1 The computing device 102 or any suitable device shall execute the command. It should be noted that... Figure 2 The steps shown are merely illustrative and should not be construed as limiting the scope of this disclosure. In different embodiments, the execution order of some steps may be adjusted, or some steps may be omitted or combined, or other additional steps may be introduced.

[0038] Figure 2A schematic flowchart of a method 200 for processing a probe plate according to some embodiments of the present disclosure is shown. At block 202, a first alignment result is determined based on the sequencing results of the probe plate and a reference sequence, the reference sequence being the sequence of probes used to synthesize the probe plate. In some embodiments, a library can be constructed based on the purified probe plate using a next-generation sequencing platform. In some embodiments, the probe plate after library construction can be sequenced to obtain sequencing results containing multiple sequencing reads. In some embodiments, reads in the sequencing results can be screened based on a preset quality threshold. For example, the sequencing results can be preprocessed using software to identify and filter out reads containing an excessively high proportion of low-quality bases or significant contamination. For example, bases with a quality value lower than Q20 (corresponding to an error rate of 1%) or Q30 (corresponding to an error rate of 0.1%) can be removed. In some embodiments, the first alignment result can be determined by comparing the sequences of the probe plate and the probes used to synthesize the probe plate using NGS technology. In some embodiments, only the reads in the selected sequencing results can be compared with the sequences of the probes used to synthesize the probe plate. In this way, the impact of abnormalities in the sequencing process on the first alignment results can be reduced.

[0039] In box 204, a second alignment result is determined based on the filtering of the first alignment result. In some embodiments, quality filtering can be performed on the first alignment result to remove low-quality, non-specific, or significantly mismatched alignment records, thereby determining the second alignment result. In some embodiments, reads in the original sequencing results can be filtered based on one or more preset quality thresholds to generate the second alignment result. For example, quality thresholds may include parameters such as the Alignment Quality Score (MAPQ), upper limit of mismatch count, lower limit of alignment length, and the structural rationality of the Compact Idiosyncratic Gapped Alignment Report (CIGAR). In this way, low-quality or non-specific alignments can be excluded from the second alignment result, ensuring the accuracy of the probe plate synthesis quality assessment.

[0040] In box 206, metrics for the sequencing results are determined based on the first and second alignment results. In some embodiments, the metrics include at least one of probe coverage information, alignment error information, and synthesis integrity. In some embodiments, probe coverage information may include, but is not limited to, at least one of: read coverage, coverage, average sequencing depth, sequencing depth distribution, and insert length distribution. Coverage is defined as the proportion of the number of bases covered by at least one read in each probe, used to assess the synthesis integrity of a single probe; average sequencing depth represents the average number of read coverage layers for all bases on each probe, reflecting the relative abundance of the probe in the synthesis pool; and sequencing depth distribution describes the dispersion of depth between different probes, used to determine the uniformity of probe synthesis. In some embodiments, probe synthesis integrity includes the maximum perfect alignment length of the probe.

[0041] In some embodiments, the alignment error information of the probe includes at least one of the number of mismatches and the number of terminal mismatches. In some embodiments, the number of terminal mismatches can be used to evaluate the synthesis fidelity of the probe terminal region, as terminal synthesis efficiency is generally low and prone to truncation or mismatches, which is an important characterization of probe synthesis process defects. In some embodiments, the alignment error information of the probe can be analyzed by parsing the CIGAR field and NM tag (representing the number of mismatched bases) in the first alignment result and / or the second alignment result. For example, the CIGAR field in the BAM file can be parsed to obtain information on operations such as matching (M), deletion (D), insertion (I), and soft shearing (S). In some embodiments, the length of consecutive mismatch-free fragments in each read can be determined based on the information on operations such as matching (M), deletion (D), insertion (I), and soft shearing (S). In some embodiments, the number of mismatched bases in each read relative to the reference sequence can be obtained by combining the NM tag in the BAM file. In some embodiments, the actual effective length of mismatch-free fragments can be further corrected by combining the CIGAR fragment and the NM value. For example, if the NM value indicates a mismatch, while CIGAR indicates multiple matching intervals, the maximum consecutive mismatch-free length can be recalculated by locating the mismatch position, excluding the interval in which it is located, or by further combining the end alignment quality to determine whether the CIGAR information is incomplete due to end mismatch, thereby achieving a more refined fragment validity assessment and correction.

[0042] In some embodiments, a first indicator for sequencing results can be determined based on a first alignment result. This first indicator is used to evaluate at least one of the overall probe synthesis status, the quality of the probe database, and the synthesis quality of the probe plate. In some embodiments, a second indicator can also be determined based on a second alignment result to evaluate the synthesis quality of the probe plate. In some embodiments, by comparing the differences between the first and second indicators, errors introduced in the probe processing flow (such as at least one of probe synthesis errors and inadequate purification), biases introduced during probe database establishment, and biases introduced during probe plate sequencing can be identified.

[0043] The above methods can provide accurate and comprehensive evaluation criteria for the quality and quantitative uniformity of the synthesized probe plate, thus providing a foundation for the quality control of the probe plate.

[0044] Figure 3 The illustration shows a schematic diagram of a method 300 for generating a probe quality control report according to some embodiments of the present disclosure. For example... Figure 3 As shown, method 300 can receive probe sequence 302 and panel sequencing data 304 as input. In some embodiments, probe sequence 302 is the raw sequence information compiled from a designed set of target probes, and each probe has a unique identifier (ProbeID). In some embodiments, paired-end sequencing files R1.fastq.gz and R2.fastq.gz obtained by directly constructing libraries from probe synthesis products using a next-generation sequencing platform can be used as panel sequencing data 304.

[0045] In some embodiments, method 300 may establish a probe sequence FASTA database 306 based on probe sequence 302. For example, the sequence of each probe is stored as an independent entry in a FASTA format file, and its ProbeID is set as the sequence header (pseudochromosome) to achieve precise anchoring at the probe level during alignment.

[0046] Method 300 can perform sequence alignment 308 between panel sequencing data 304 and a reference sequence. In some embodiments, the reference sequence is a probe selected from the probe sequence FASTA database 306. In some embodiments, the panel sequencing data 304 can be screened before sequence alignment 308. In some embodiments, the screened reads in the panel sequencing data 304 can be aligned to the probe sequence FASTA database 306 to generate a complete alignment result (Raw) 310. In some embodiments, the unscreened RawBAM file output by the alignment tool can be used as the complete alignment result (Raw) 310. In some embodiments, the complete alignment result (Raw) 310 includes all alignment events, including low-quality, multiple alignments, and anomalous reads, thereby preserving complete original alignment information. In some embodiments, the complete alignment result (Raw) 310 can be sorted and indexed for subsequent statistical analysis.

[0047] Method 300 can extract high-quality alignment results (HQ) 312 from the complete alignment results (Raw) 310. This can be achieved, for example, by setting an alignment quality value (MAPQ) threshold (e.g., MAPQ ≥ 10), removing reads from multiple or secondary alignments, and eliminating abnormal reads with obvious mismatches or excessive indels. In this way, low-quality alignments caused by synthesis defects, incomplete purification, or sequencing noise can be filtered out from the high-quality alignment results (HQ) 312, retaining high-confidence matches and thus improving the accuracy of subsequent quantification and quality assessment. In some embodiments, the high-quality alignment results (HQ) 312, representing ideal synthesis performance, are used for comparative analysis with the complete alignment results (Raw) 310.

[0048] In some embodiments, method 300 may perform alignment result statistics 320 on the complete alignment result (Raw) 310 and the high-quality alignment result (HQ) 312 respectively. In some embodiments, method 300 may extract at least one of the following indicators from the complete alignment result (Raw) 310 and the high-quality alignment result (HQ) 312: probe uniformity 322, probe coverage 324, NGS insert length distribution 326, and NGS read depth 328. Specifically, probe coverage 324 reflects the proportion of bases in the reference sequence of each probe that are covered by the sequencing read, and this indicator can reflect whether there are gaps or incomplete amplifications during the synthesis and library construction of a probe; probe uniformity 322 describes the uniformity of sequencing depth among probes and assesses the consistency of the synthesis process; NGS insert length distribution 326 is used to infer the actual synthesized length of the probe and fragment integrity; and NGS read depth 328 reflects the distribution characteristics of the overall sequencing depth and can be used to assist in judging sample concentration and library construction efficiency. In some embodiments, the proportion of bases covered by the sequencing reads can be calculated on a per-probe reference sequence basis. In some embodiments, coverage can be calculated using the following formula:

[0049] Coverage = Number of bases in the probe covered by at least one read / Total probe length

[0050] In some embodiments, method 300 can perform alignment file read length marker parsing 314 on both the complete alignment result (Raw) 310 and the high-quality alignment result (HQ) 312. For example, method 300 can parse information such as CIGAR words, NM tags (in BAM files, NM tags indicate the edit distance, i.e., the minimum number of nucleotide edits required to convert the read sequence into a reference sequence), and MD tags (in BAM files, MD tags show the specific number of mismatches) in the Raw BAM and HQBAM files respectively, generating indicators such as base mismatch distribution 316 and mismatch-free full length 318. Among them, the base mismatch distribution 316 counts the total number of mismatches and the number of terminal mismatches for each read, used to identify synthesis truncation or erroneous incorporation problems; the mismatch-free full length 318 represents the maximum consecutive perfect alignment length, directly reflecting the integrity of probe synthesis. In some embodiments, specific scripts can be used to batch parse the various indicators in the complete alignment result (Raw) 310 and / or the high-quality alignment result (HQ) 312. In some embodiments, if the differences between the metrics corresponding to the complete alignment result (Raw) 310 and the high-quality alignment result (HQ) 312 are small, it indicates that the overall synthesis quality is high; if the differences are significant, it suggests that there are a large number of low-quality alignments, requiring resynthesis or process optimization. In some embodiments, specific scripts can be designed to merge various metrics to the probe level, forming a synthesis quality report for each probe in the Panel.

[0051] In some embodiments, method 300 can perform statistical testing and visualization operations on one or more of the generated indicators. For example, method 300 can generate probe coverage distribution maps, read count histograms, mismatch rate heatmaps, insert fragment length curves, and synthesis integrity tables to visually present the overall quality status of the panel. In some embodiments, method 300 can generate corresponding visualized statistical results, such as charts, based on one or more indicators corresponding to the complete alignment results (Raw) 310. In some embodiments, method 300 can also generate corresponding visualized statistical results based on one or more indicators of the high-quality alignment results (HQ) 312. In some embodiments, the effectiveness of the screening strategy can be verified by comparing the statistical results of the complete alignment results (Raw) 310 and the high-quality alignment results (HQ) 312.

[0052] In some embodiments, method 300 can output a final probe quality control report 332. In some embodiments, all analytical results can be integrated into the probe quality control report 332 to provide a basis for quality control and uniformity verification before panel use. The probe quality control report 332 can not only be used to determine whether the current batch of probes is qualified, but also guide subsequent probe screening and resynthesis strategies: for example, probes with poor synthesis quality can be eliminated or resynthesized separately in the next round of synthesis; while probes with excellent performance can be retained to improve detection performance.

[0053] Using the methods described above, a complete quantitative and quality control analysis workflow for hybridization capture probes was constructed. By processing the Raw and HQ alignment results in parallel and combining multi-dimensional indicators such as coverage, depth, mismatch, and integrity, a refined quality assessment of probes and panels was achieved. This method not only enables a systematic evaluation of probe synthesis quality but also provides a basis for optimizing panel design and improving detection accuracy.

[0054] Figures 4-7 The illustration shows a schematic diagram of the visualization output used to evaluate the synthesis quality of the probe plate. In, for example... Figures 4-7 In the exemplary visualization output shown, the total number of probes N=254174, and the data used comes from high-quality comparison results after quality filtering. It should be understood that the statistical indicators, data statistics results, and visualization forms shown in the figure are only examples, and this disclosure does not limit the specific content and form of the display.

[0055] Figure 4 An example of a probe coverage distribution histogram is shown. This graph is generated by statistically analyzing the percentage of bases covered by sequencing reads in the reference sequence of each probe. The coverage formula is: Coverage = Number of bases covered by at least one read / Total probe length. The horizontal axis represents the percentage of probe length covered by sequencing reads, and the vertical axis represents the number of probes with the corresponding coverage (Count). This graph can visually reflect whether there are gaps or incomplete amplifications during probe synthesis and library construction. For example, in... Figure 4 The probe coverage distribution histogram shown, based on high-quality alignment results, shows that the vast majority of probes (approximately 240,000) achieved 100% coverage, with only a very small number of probes showing partial missing or uncovered conditions.

[0056] Figure 5This example illustrates the distribution of average sequencing depth for each probe. The graph is plotted by statistically analyzing the average sequencing depth of all bases on each probe. The horizontal axis represents the average read depth per probe, and the vertical axis represents the frequency of probes at that depth. (The text then repeats itself, so the translation only includes the first instance.) Figure 5 The average sequencing depth distribution curve shown can identify whether there are probes with consistently low or high synthesis efficiency. In, for example... Figure 5 In the exemplary probe average sequencing depth distribution shown, the depth of most probes is concentrated in the range of 100 to 200 reads, while a small number of probes have lower or higher depths. This intuitively reflects the non-uniformity in actual synthesis. It should be understood that the distribution curves and specific values ​​here are merely examples, and this disclosure does not limit the specific depth distribution pattern. In some embodiments, probe average sequencing depth distribution maps can also be generated for the original alignment results and the selected high-quality alignment results. In this way, it is possible to distinguish between the deviations introduced by the sequencing or library preparation process and the depth differences caused by defects in the synthesis quality of the probes themselves.

[0057] Figure 6 This shows a histogram illustrating the distribution of the maximum consecutive mismatch-free lengths of all sequencing reads at the individual probe level. (Example) Figure 6 As shown, the horizontal axis represents the maximum mismatch-free length (Length), and the vertical axis represents the number of reads of the corresponding length (Count). In some embodiments, this can be achieved by extracting the alignment interval of each read on the target probe and counting the maximum consecutive mismatch-free length within it. At the individual probe level, the maximum consecutive mismatch-free length of all sequencing reads reflects the longest high-fidelity alignment fragment of a single read on the probe, which can sensitively capture local quality problems caused by probe synthesis defects (such as incorrect base incorporation) or sequencing errors. For example, for a probe with ideal synthesis quality, the distribution of its reads should be significantly concentrated in a longer mismatch-free length interval.

[0058] Figure 7 This displays a global histogram showing the overall distribution of the maximum mismatch-free length for all probes. (Example:) Figure 7 As shown, by summarizing the mismatch-free length data of all probes (e.g., N=254174 in this example), the synthesis quality and sequencing stability of the entire probe population can be evaluated from a macroscopic perspective. Figure 7In the exemplary embodiment shown, the horizontal axis divides the maximum mismatch-free length into four intervals: <80 bp, 80–100 bp, 100–120 bp, and ≥120 bp, while the vertical axis represents the read counts within each interval, in units of 1e7. This global perspective helps to form an intuitive judgment of the overall performance of the probe library; for example, if a large number of reads are concentrated in long fragment intervals (e.g., ≥120 bp), it indicates that the overall synthesis fidelity of the probe population is high. In some embodiments, Figure 6 and Figure 7 Together, they provide a complete quality view from local to global. It should be understood that the interval divisions and specific distributions shown in the figures are merely examples, and this disclosure does not limit the specific presentation format.

[0059] The above methods enable the evaluation of probe plate synthesis quality from multiple perspectives, including probe integrity, synthesis uniformity, and sequence fidelity, through various metrics such as coverage, sequencing depth, and mismatch-free length, overcoming the limitations of single-metric evaluation. Furthermore, multi-level analysis of individual probe sequences and the entire probe sequence library allows for assessment of individual anomalous probes and the overall performance of the probe database. These methods provide data support for optimizing probe plate synthesis processes and screening high-quality probe libraries, thereby ensuring the reliability of downstream applications.

[0060] It should be understood that in the embodiments of this disclosure, "first," "second," etc., are only used to indicate that multiple objects may be different, but at the same time, it does not exclude that two objects are the same, and should not be interpreted as any limitation on the embodiments of this disclosure.

[0061] It should also be understood that the manner, situation, category, and division of embodiments in the present disclosure are for the convenience of description only and should not constitute a special limitation. Various manners, categories, situations, and features in the embodiments can be combined with each other where logically consistent.

[0062] It should also be understood that the foregoing is merely to help those skilled in the art better understand the embodiments of this disclosure, and is not intended to limit the scope of the embodiments of this disclosure. Those skilled in the art can make various modifications, variations, or combinations based on the foregoing. Such modifications, variations, or combinations are also within the scope of the embodiments of this disclosure.

[0063] It should also be understood that the above description focuses on highlighting the differences between the various embodiments. Similarities or commonalities can be referenced or learned from each other, and for the sake of brevity, they will not be repeated here.

[0064] Figure 8 A schematic block diagram of an example device 800 according to some embodiments of the present disclosure is shown. Device 800 can be implemented by software, hardware, or a combination of both. Figure 8 As shown, the device 800 includes an alignment unit 802, a screening unit 804, and an indicator determination unit 806. The alignment unit 802 is configured to determine a first alignment result based on the sequencing results of the probe plate and a reference sequence, wherein the reference sequence is the sequence of probes used to synthesize the probe plate; the screening unit 804 is configured to determine a second alignment result based on screening the first alignment result; and the indicator determination unit 806 is configured to determine an indicator for the sequencing results based on the first and second alignment results.

[0065] In some embodiments, the screening unit 804 is further configured to screen the first alignment result based on at least one of the alignment quality value, multiple alignment information, number of mismatches, and number of insertion / deletion markers of the first alignment result to obtain a second alignment result, wherein the multiple alignment information indicates that a single read in the sequencing result has two or more alignment records in the first alignment result.

[0066] In some embodiments, the probes are selected from a probe database, which includes unique identifiers and sequences of the probes.

[0067] In some embodiments, the indicator determination unit 806 includes a first indicator determination unit configured to determine a first indicator for the sequencing results based on a first alignment result, the first indicator being used to determine at least one of the overall probe synthesis status, the quality of the probe database, and the synthesis quality of the probe plate; and a second indicator determination unit configured to determine a second indicator for the sequencing results based on a second alignment result, the second indicator being used to determine the synthesis quality of the probe plate.

[0068] In some embodiments, the difference between the first and second indicators is used to determine anomalies caused by factors other than the synthesis quality of the probe plate, wherein the anomalies include at least one of the following: errors introduced in the probe processing flow, including at least one of probe synthesis errors and inadequate purification; biases introduced in the process of establishing the probe database; and biases introduced in the sequencing process of the probe plate.

[0069] In some embodiments, the apparatus 800 further includes a first generation unit configured to generate a first visualization chart corresponding to a first comparison result based on a first indicator; a second generation unit configured to generate a second visualization chart corresponding to a second comparison result based on a second indicator; and an output unit configured to output the first visualization chart and the second visualization chart. In some embodiments, the metrics include at least one of probe coverage information, alignment error information, and synthesis integrity.

[0070] In some embodiments, the probe coverage information includes at least one of the following: probe read coverage number, coverage, average sequencing depth, sequencing depth distribution, and insert length distribution.

[0071] In some embodiments, the sequencing results include multiple reads, and the device 800 further includes a read filtering unit configured to filter reads in the sequencing results based on a preset quality threshold.

[0072] In some embodiments, the index includes the length of consecutive mismatch-free segments in each read segment, and the index determination unit 806 further includes: an alignment information extraction unit configured to extract alignment information for each read segment based on a first alignment result and a second alignment result, the alignment information including at least one of matching, missing, insertion, and soft pruning of the read segment; and a length determination unit configured to determine the length of consecutive mismatch-free segments in each read segment based on the alignment information.

[0073] In some embodiments, the apparatus 800 further includes a mismatch base number determination unit configured to determine the number of mismatch bases in each read relative to a reference sequence based on a first alignment result and a second alignment result; and a correction unit configured to correct the length of consecutive mismatch-free segments in each read based on the number of mismatch bases.

[0074] In some embodiments, the index includes the length of consecutive mismatch-free segments in each read segment, and the index determination unit 806 further includes: a maximum length determination unit configured to determine the length of the largest consecutive mismatch-free segment in each read segment; and a length distribution determination unit configured to determine the length distribution of consecutive mismatch-free segments in multiple read segments.

[0075] In some embodiments, the probe plate is used for methylation detection.

[0076] Figure 8 The device 800 can be used to achieve the above-mentioned combination. Figures 1 to 3 For the sake of brevity, the process described will not be repeated here.

[0077] The division of modules or units in the embodiments of this disclosure is illustrative and only represents one logical functional division. In actual implementation, there may be other division methods. Furthermore, the functional units in the disclosed embodiments may be integrated into one unit, exist as separate physical entities, or two or more units may be integrated into one unit. The integrated unit described above can be implemented in hardware or as a software functional unit.

[0078] Figure 9 A schematic block diagram of an example device 900 that can be used to implement embodiments of the present disclosure is shown. Figure 1The computing device 102 can be implemented using device 900. As shown, device 900 includes a central processing unit (CPU) 901, which can perform various appropriate actions and processes according to computer program instructions stored in read-only memory (ROM) 902 or loaded from storage unit 908 into random access memory (RAM) 903. RAM 903 can also store various programs and data required for the operation of device 900. CPU 901, ROM 902, and RAM 903 are interconnected via bus 904. Input / output (I / O) interface 909 is also connected to bus 904.

[0079] Multiple components in device 900 are connected to I / O interface 909, including: input unit 906, such as keyboard, mouse, etc.; output unit 907, such as various types of monitors, speakers, etc.; storage unit 908, such as disk, optical disk, etc.; and communication unit 909, such as network card, modem, wireless transceiver, etc. Communication unit 909 allows device 900 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.

[0080] The various processes and handling described above, such as method 200, can be executed by processing unit 901. For example, in some embodiments, method 200 can be implemented as a computer software program tangibly contained in a machine-readable medium, such as storage unit 908. In some embodiments, part or all of the computer program can be loaded and / or installed on device 900 via ROM 902 and / or communication unit 909. When the computer program is loaded into RAM 903 and executed by CPU 901, one or more actions of the example method 200 described above can be performed.

[0081] According to an exemplary implementation of this disclosure, a computer-readable storage medium is provided that stores computer-executable instructions thereon, wherein the computer-executable instructions are executed by a processor to implement the methods described above. According to an exemplary implementation of this disclosure, a computer program product is also provided, which is tangibly stored on a non-transitory computer-readable medium and includes computer-executable instructions, which are executed by a processor to implement the methods described above. According to an exemplary implementation of this disclosure, a computer program product is provided that stores a computer program thereon, which, when executed by a processor, implements the methods described above.

[0082] Various aspects of this disclosure are described herein with reference to flowchart illustrations and / or block diagrams of methods, apparatuses, devices, and computer program products implemented according to this disclosure. It should be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer-readable program instructions.

[0083] These computer-readable program instructions can be provided to a processing unit of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus to produce a machine such that, when executed by the processing unit of the computer or other programmable data processing apparatus, they create means for implementing the functions / actions specified in one or more blocks of the flowchart and / or block diagram. These computer-readable program instructions can also be stored in a computer-readable storage medium that causes a computer, programmable data processing apparatus, and / or other device to operate in a particular manner. Thus, the computer-readable medium storing the instructions comprises an article of manufacture that includes instructions for implementing aspects of the functions / actions specified in one or more blocks of the flowchart and / or block diagram.

[0084] Computer-readable program instructions can be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable data processing apparatus, or other device to produce a computer-implemented process, thereby causing the instructions that execute on the computer, other programmable data processing apparatus, or other device to perform the functions / actions specified in one or more boxes of a flowchart and / or block diagram.

[0085] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of an instruction, which contains one or more executable instructions for implementing the specified logical function. In some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutive blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, may be implemented using a dedicated hardware-based system that performs the specified function or action, or using a combination of dedicated hardware and computer instructions.

[0086] Various implementations of this disclosure have been described above. These descriptions are exemplary and not exhaustive, nor are they limited to the disclosed implementations. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the described implementations. The terminology used herein is chosen to best explain the principles, practical applications, or improvements to technology in the market, or to enable others skilled in the art to understand the various implementations disclosed herein.

Claims

1. A method for processing a probe plate, comprising: Based on the sequencing results of the probe plate and the reference sequence, a first alignment result is determined, wherein the reference sequence is the sequence used to synthesize the probes of the probe plate; Based on the filtering of the first comparison results, the second comparison result is determined; as well as Based on the first alignment result and the second alignment result, indicators for the sequencing results are determined.

2. The method according to claim 1, wherein filtering the first comparison result to obtain the second comparison result includes: Based on at least one of the alignment quality value, multiple alignment information, number of mismatches, and number of insertion / deletion markers of the first alignment result, the first alignment result is filtered to obtain the second alignment result, wherein the multiple alignment information indicates that a single read in the sequencing result has two or more alignment records in the first alignment result.

3. The method of claim 1, wherein the probe is selected from a probe database, the probe database including the unique identifier and sequence of the probe.

4. The method according to claim 3, wherein determining the indicators for the sequencing results based on the first alignment result and the second alignment result includes: Based on the first alignment result, a first indicator is determined for the sequencing result. The first indicator is used to determine at least one of the following: the overall synthesis status of the probe, the quality of the probe database, and the synthesis quality of the probe plate. as well as Based on the second alignment result, a second indicator is determined for the sequencing result, the second indicator being used to determine the synthesis quality of the probe plate.

5. The method of claim 4, wherein the difference between the first indicator and the second indicator is used to determine anomalies caused by factors other than the synthesis quality of the probe plate, wherein the anomalies include at least one of the following: Errors introduced in the probe processing procedure include at least one of the probe synthesis error and inadequate purification; Bias introduced during the establishment of the probe database; and Bias introduced during the sequencing process of the probe plate.

6. The method according to claim 4, further comprising: Based on the first indicator, a first visualization chart corresponding to the first comparison result is generated; Based on the second indicator, a second visualization chart corresponding to the second comparison result is generated; as well as Output the first visualization chart and the second visualization chart.

7. The method of claim 1, wherein the metrics include at least one of probe coverage information, alignment error information, and synthesis integrity.

8. The method according to claim 7, wherein the probe coverage information includes at least one of the probe read coverage number, coverage, average sequencing depth, sequencing depth distribution, and insert length distribution.

9. The method of claim 7, wherein the probe alignment error information includes at least one of the number of probe mismatches and the number of end mismatches.

10. The method according to claim 1, wherein the sequencing result comprises a plurality of reads, and further comprises: Based on a preset quality threshold, reads in the sequencing results are filtered.

11. The method of claim 10, wherein the metric includes the length of consecutive mismatch-free segments in each read, and determining the metric for the sequencing result based on the first alignment result and the second alignment result includes: Based on the first comparison result and the second comparison result, the comparison information of each read segment is extracted, and the comparison information includes at least one of the following: matching, missing, insertion, and soft pruning of the read segment; as well as Based on the comparison information, the length of the consecutive mismatch-free segments in each read segment is determined.

12. The method of claim 11, further comprising: Based on the first alignment result and the second alignment result, the number of mismatched bases in each read segment compared to the reference sequence is determined; as well as Based on the number of mismatched bases, the length of the consecutive mismatch-free segments in each read is corrected.

13. The method of claim 10, wherein the metric includes the length of consecutive mismatch-free segments in each read, and determining the metric for the sequencing result based on the first alignment result and the second alignment result includes: Determine the length of the largest consecutive mismatch-free segment in each of the read segments; as well as Determine the length distribution of the consecutive mismatch-free segments among the multiple read segments.

14. The method of claim 1, wherein the probe plate is used for methylation detection.

15. An electronic device comprising: At least one processing unit; At least one memory coupled to the at least one processing unit and storing instructions for execution by the at least one processing unit, the instructions causing the electronic device to perform the method according to any one of claims 1 to 14 when executed by the at least one processing unit.

16. An apparatus for processing a probe plate, comprising: The alignment unit is configured to determine a first alignment result based on the sequencing results of the probe plate and a reference sequence, wherein the reference sequence is the sequence used to synthesize the probes of the probe plate; The filtering unit is configured to determine the second comparison result based on the filtering of the first comparison result; as well as The indicator determination unit is configured to determine an indicator for the sequencing result based on the first alignment result and the second alignment result.

17. A computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the method according to any one of claims 1 to 14.