Method, device, equipment and product for determining miss probability of probe sequence
By constructing reference genomes with different methylation states and performing high-sensitivity alignment, the challenges of probe sequence design and non-specific binding in methylation sequencing were addressed, improving detection accuracy and sensitivity while ensuring the specificity of probe sequence design.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-25
- Publication Date
- 2026-03-27
AI Technical Summary
During methylation sequencing, the difficulty of probe sequence design increases, the risk of non-specific binding is high, leading to a decrease in detection accuracy and sensitivity.
Reference genomes of hypomethylated positive strand, hypomethylated negative strand, hypermethylated positive strand, and hypermethylated negative strand were constructed. The off-target probability of probe sequences was assessed by homologous sequence alignment. High-sensitivity alignment was performed using the blastn tool. Soft masking and low-complexity region filtering were disabled. Multi-threaded parallel processing was used to improve efficiency.
It improves the accuracy and sensitivity of probe sequence off-target detection, reduces false matches and missed detections, and provides specific control over probe sequence design.
Smart Images

Figure CN121747700A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] Embodiments of the present disclosure generally relate to the field of bioinformatics, and in particular, to methods, apparatuses, devices and products for determining off-target probability of a probe sequence. BACKGROUND
[0002] In the field of disease detection, target region methylation sequencing method has become a commonly used detection means in actual business due to the advantages of reducing sequencing cost and improving sequencing depth of specific regions. The basic principle is: using sequence-specific oligonucleotide (DNA-oligo) probe sequence, enriching target gene region or specific CpG site through hybridization capture, and then combining bisulfite sequencing or enzyme conversion sequencing to realize directional detection of disease-related region methylation signal.
[0003] In the process of methylation sequencing, this chemical or enzymatic treatment will convert all cytosine (Cytosine, C) except 5-methylcytosine (5mC) into thymine (Thymine, T). This significantly reduces the base complexity of the genome sequence, making a large number of originally distinguishable sequences similar, introducing a large number of non-specific binding (off-target) risks, and significantly increasing the difficulty of probe sequence design. SUMMARY
[0004] Embodiments of the present disclosure provide a method, apparatus, device, medium and product for determining off-target probability of a probe sequence.
[0005] According to a first aspect of the present disclosure, a method for determining off-target probability of a probe sequence is provided. The method comprises constructing a first genome, the first genome comprising at least one of a low methylation positive strand, a low methylation negative strand, a high methylation positive strand and a high methylation negative strand. The method further comprises determining the off-target probability of the probe sequence based on the constructed first genome. According to a second aspect of the present disclosure, an apparatus for determining off-target probability of a probe sequence is provided. The apparatus comprises a first genome construction module configured to construct a first genome, the first genome comprising at least one of a low methylation positive strand, a low methylation negative strand, a high methylation positive strand and a high methylation negative strand; and an off-target probability determination module configured to determine the off-target probability of the probe sequence based on the constructed first genome.
[0006] In a third aspect of the present disclosure, an electronic device is provided, comprising at least one processor; and a memory for storing at least one program, when the at least one program is executed by the at least one processor, causing the at least one processor to implement the method according to the first aspect of the present disclosure.
[0007] In a fourth aspect of this disclosure, a computer-readable storage medium is provided having a computer program stored thereon that, when executed by a processor, implements the method according to a first aspect of this disclosure.
[0008] In a fifth aspect of this disclosure, a computer program product is provided. This computer program product includes a computer program that, when executed by a processor, implements the method according to a first aspect of this disclosure.
[0009] It should be understood that the content described in this section is not intended to limit the key or essential features of the embodiments of this disclosure, nor is it intended to restrict the scope of this disclosure. Other features of this disclosure will become readily apparent from the following description. Attached Figure Description
[0010] The above and other objects, features and advantages of this disclosure will become more apparent from the accompanying drawings, in which like reference numerals generally denote like parts.
[0011] Figure 1 The illustration shows a schematic diagram of an example environment 100 that can be applied to some embodiments of the present disclosure;
[0012] Figure 2 The illustration shows a flowchart of a method 200 for determining the off-target probability of a probe sequence according to some embodiments of the present disclosure;
[0013] Figure 3 The illustration shows a schematic diagram of an example flow 300 of a probe sequence off-target risk prediction method according to some embodiments of the present disclosure.
[0014] Figure 4 The illustration shows a schematic diagram of a method 400 for constructing a reference genome according to some embodiments of the present disclosure;
[0015] Figure 5 The illustration shows a schematic block diagram of an apparatus 500 for determining the off-target probability of a probe sequence according to some embodiments of the present disclosure.
[0016] Figure 6 A schematic block diagram of an example device 600 that can be used to implement embodiments of the present disclosure is shown. Detailed Implementation
[0017] It is understood that the data involved in this technical solution (including but not limited to the data itself, the acquisition or use of the data) shall comply with the requirements of relevant laws, regulations and related provisions.
[0018] Embodiments of this disclosure will now be described in more detail with reference to the accompanying drawings. While some embodiments of this disclosure are shown in the drawings, it should be understood that this disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of this disclosure. It should be understood that the accompanying drawings and embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of protection of this disclosure.
[0019] In the description of embodiments of this disclosure, the term "comprising" and similar terms should be understood as open-ended inclusion, i.e., "including but not limited to". The term "based on" should be understood as "at least partially based on". The term "one embodiment" or "the embodiment" should be understood as "at least one embodiment". The terms "first", "second", etc., may refer to different or the same objects. Other explicit and implicit definitions may also be included below.
[0020] During DNA methylation sequencing, many previously differentiated fragments in the transformed genome tend to become homogeneous, making non-specific hybridization more likely between probe sequences and sequences from different regions. Simultaneous binding of probe sequences to multiple similar fragments in the genome can lead to signal confusion or increased noise, thus reducing the accuracy of methylation detection. During sample processing, DNA is divided into numerous short fragments after enzymatic or chemical treatment. These short fragments, with their reduced complexity, are more likely to exhibit "local similarity binding" with probe sequences, further increasing the risk of off-target effects. Furthermore, after methylation transformation, the original genome's double strands (OT / OB) no longer maintain strict complementarity, leading to additional non-specific binding of probe sequences on one strand, thereby disrupting the specific capture of signals from the target region. In regions enriched by cytosine-phosphate-guanine (CpG) sites, the combination of methylation states at multiple sites further affects sequence diversity. Transformed local sequences often exhibit high homology, increasing the chance of probe sequences binding to non-target regions.
[0021] To address the aforementioned problems and other potential issues, embodiments of this disclosure propose a method for determining the off-target probability of a probe sequence. In embodiments of this disclosure, a computing device can construct a first genome, which includes at least one of a hypomethylated positive strand, a hypomethylated negative strand, a hypermethylated positive strand, and a hypermethylated negative strand. Next, the computing device can determine the off-target probability of the probe sequence based on the constructed first genome. This method, by assessing the off-target risk of methylation-capture probe sequences on the methylated genome, solves the problems of decreased base complexity and inaccurate alignment caused by sequence shifts due to methylation conversion, thereby improving the accuracy and sensitivity of probe sequence off-target detection, effectively avoiding false matches or missed detections, and providing a foundation for specific control of probe sequence design.
[0022] The embodiments of this disclosure will now be described in further detail with reference to the accompanying drawings. Figure 1 A schematic diagram of an example environment 100 applicable to some embodiments of the present disclosure is shown. In environment 100, computing device 102 can be used to determine the off-target probability of a probe sequence.
[0023] Examples of computing device 102 include, but are not limited to, personal computers, server computers, handheld or laptop devices, mobile devices (such as mobile phones, personal digital assistants (PDAs), media players, etc.), multiprocessor systems, consumer electronics, minicomputers, mainframe computers, and distributed computing environments that include any of the above systems or devices.
[0024] like Figure 1As shown, computing device 102 can construct a reference genome 106 based on genome 104. In some embodiments, reference genome 106 can be used to determine the off-target probability 112 of probe sequences. In some embodiments, reference genome 106 may include various methylation-converted genomes, such as Forward_unmethylated (FU) genome, Forward_methylated (FM) genome, Reverse_unmethylated (RU) genome, and Reverse_methylated (RM) genome, which correspond to simulated sequences of hypomethylated positive strand, hypermethylated positive strand, hypomethylated negative strand, and hypermethylated negative strand, respectively. In some embodiments, the FU genome is constructed by converting all cytosine (C) bases at the cytosine-any-base (non-G)-any-base (non-G) (Cytosine-HH, CHH) sites, cytosine-any-base (non-G)-guanine (Cytosine-H-Guanine, CHG) sites, and CpG sites on the positive strand of the original genome 104 to thymine (T) bases, to simulate the sequence of the positive strand in a hypomethylated state. In some embodiments, the FM genome converts only the C at the non-CpG sites (i.e., CHH and CHG) on the original genome 104 to T, while leaving the C at the CpG sites unchanged, to simulate the case where the CpG sites on the positive strand are protected and do not undergo conversion in a hypermethylated state. In some embodiments, the RU genome is constructed by converting all C-complementary guanine (G) bases on the negative strand to adenine (A) bases, to simulate the sequence of the negative strand in a hypomethylated state. In some embodiments, the RM genome converts G at non-CpG positions in the negative strand to A, while leaving G at CpG positions unchanged, to reflect the conversion characteristics of the negative strand under hypermethylation conditions.
[0025] The computing device 102 can determine the off-target probability 112 of the probe sequence based on the constructed reference genome 108. In some embodiments, the computing device 102 can determine the sequence similarity between the probe sequence 110 and the constructed reference genome 108 by performing homologous sequence alignment, and determine the off-target probability 112 of the probe sequence based on the similarity. The similarity and the off-target probability 112 of the probe sequence are positively correlated. In one embodiment, the blastn tool can be used for local alignment. In some embodiments, by setting high sensitivity parameters and disabling soft masking of the constructed reference genome 108, while turning off low-complexity region filtering in the alignment algorithm, it can be ensured that all bases in the constructed reference genome 108 and the entire sequence of the probe sequence 110 participate in the alignment process. In this way, potential off-target sites can be missed, thereby improving the sensitivity and completeness of detection. In some embodiments, the multi-threading functionality supported by alignment tools such as blastn can be utilized to fully utilize multi-core CPU resources within a single node. In some embodiments, the probe sequence set is divided into multiple subtasks using customized analysis scripts and distributed to multiple computing nodes or processes for parallel processing, achieving cross-node task-level parallelism. This approach alleviates the performance bottleneck caused by high sensitivity parameters, significantly improving overall alignment efficiency while ensuring detection accuracy, making it suitable for the design and screening of large-scale methylation probe sequences.
[0026] The above method can assess the off-target risk of methylation capture probe sequences on the genome after methylation transformation, solve the problem of inaccurate alignment caused by the decrease in base complexity and sequence shift due to methylation transformation, thereby improving the accuracy and sensitivity of probe sequence off-target detection, effectively avoiding false matches or missed detections, and providing a basis for the specific control of probe sequence design.
[0027] The above combination Figure 1 A schematic diagram of an example environment 100 that can be applied to some embodiments of the present disclosure is described below, in conjunction with... Figure 2 A flowchart describing a method for determining the off-target probability of a probe sequence according to some embodiments of the present disclosure. Figure 2 Method 200 in the middle can be derived from Figure 1 The computing device 102 or any suitable device shall execute the command. It should be noted that... Figure 2 The steps shown are merely illustrative and should not be construed as limiting the scope of this disclosure. In different embodiments, the execution order of some steps may be adjusted, or some steps may be omitted or combined, or other additional steps may be introduced.
[0028] like Figure 2As shown, in block 202 of method 200, a first genome is constructed, comprising at least one of a hypomethylated positive strand, a hypomethylated negative strand, a hypermethylated positive strand, and a hypermethylated negative strand. In some embodiments, the first genome can be constructed by altering bases at specific sites in the genome. For example, a first genome comprising a hypomethylated positive strand is constructed by converting cytosine at the CHH, CHG, and CpG sites on the positive strand of a second genome to thymine. Similarly, a first genome comprising a hypomethylated negative strand is constructed by converting guanine at the CHH, CHG, and CpG sites on the negative strand of a second genome to adenine. Likewise, a first genome comprising a hypermethylated positive strand is constructed by converting cytosine at the CHH and CHG sites on the positive strand of a second genome to thymine. And also, a first genome comprising a hypermethylated negative strand is constructed by converting guanine at the CHH and CHG sites on the negative strand of a second genome to adenine. In some embodiments, before altering bases at specific sites in the second genome, the location of CpG sites on the second genome can be determined, and then, based on the determined location, the region containing the CpG sites on the second genome can be masked. In this way, cytosine on the positive strand of the genome in unmasked regions can be converted to thymine during the construction of a hypermethylated positive strand, and guanine on the negative strand of the genome in unmasked regions can be converted to adenine during the construction of a hypermethylated negative strand.
[0029] In box 204, based on the constructed first genome, the off-target probability of the probe sequence is determined. In some embodiments, the off-target risk can be quantified by performing sequence alignment of the probe sequence with the four types of first genomes (i.e., the first genome containing a low-methylated positive strand, the first genome containing a low-methylated negative strand, the first genome containing a high-methylated positive strand, and the first genome containing a high-methylated negative strand) respectively, and systematically evaluating the potential non-specific binding sites of the probe sequence under different methylation states (i.e., low-methylation or high-methylation) and different strand orientations (positive strand or negative strand).
[0030] In other embodiments, the analysis can be performed on the genome under specific methylation states. For example, the probe sequence can be aligned with hypomethylated positive and hypomethylated negative strands to assess the potential off-target effects of the probe sequence in an overall hypomethylated background; or, hypermethylated positive and hypermethylated negative strands can be used to assess the binding specificity of the probe sequence in a hypermethylated environment.
[0031] In other embodiments, the directionality and specificity of probe sequence design can be ensured by comparing the probe sequence with a first genome sequence exhibiting a specific methylation state and strand orientation. For example, the specificity of a designed low-methylation orientation probe sequence on the positive strand can be checked by examining the low-methylation positive strand. Similarly, the specificity of a designed low-methylation orientation probe sequence on the reverse strand can be checked by examining the low-methylation negative strand. Likewise, the specificity of a designed high-methylation orientation probe sequence on the positive strand can be checked by examining the high-methylation positive strand. And again, the specificity of a designed high-methylation orientation probe sequence on the reverse strand can be checked by examining the high-methylation negative strand.
[0032] The above method can assess the off-target risk of methylation detection probe sequences on the genome after methylation transformation, solve the problem of inaccurate alignment caused by the decrease in base complexity and sequence shift due to methylation transformation, thereby improving the accuracy and sensitivity of probe sequence off-target detection, effectively avoiding false matches or missed detections, and providing a basis for the specific control of probe sequence design.
[0033] The above combination Figure 2 A flowchart describing a method for training an image generation model according to some embodiments of this disclosure is provided below. Figure 3 A schematic diagram of an example flow 300 describing a probe sequence off-target risk prediction method according to some embodiments of the present disclosure. Figure 3 Example method 300 in the example can be derived from Figure 1 The computing device 102 or any suitable device in the process is used for processing. It is understood that process 300 is merely illustrative and not intended to limit the scope of embodiments of this disclosure. For example, in Figure 3 The steps shown may be interchanged, some steps may be combined, omitted or modified, or additional steps not shown may be included.
[0034] like Figure 3 As shown, method 300 starts with the human reference genome 304 and obtains a fully methyl reference genome (positive strand FM) 306 by converting all C bases at the CHH / CHG sites to T bases. In some embodiments, a fully non-methyl reference genome (positive strand FU) 310 can be obtained by converting all C bases at the CpG sites in the fully methyl reference genome (positive strand FM) 306 to thymine T bases. In some embodiments, method 300 can start with the human reference genome 304 and obtain a fully methyl reference genome (negative strand FM) 308 by converting complementary G bases at the CHH / CHG sites to A bases. In some cases, method 300 can convert all G bases at the CpG sites in the fully methyl reference genome (negative strand RM) 308 to A bases to obtain a fully non-methyl reference genome (negative strand RU) 312.
[0035] In some embodiments, the probe sequence includes the original top strand (OT) and the complementary strand to the original top strand (ctOT). In some embodiments, the probe sequence includes the original revers strand (also known as the original bottom strand, OB) and the complementary strand to the original bottom strand (ctOB).
[0036] In some embodiments, method 300 can establish a one-to-one alignment rule between probe sequences and reference genomes: for example, aligning a highly methylated OT / ctOT probe sequence 314 with a fully methylated reference genome (positive strand FM) 306; aligning a low-methylated OT / ctOT probe sequence 316 with a fully non-methylated reference genome (positive strand FU) 310; aligning a highly methylated OB / ctOB probe sequence 318 with a fully methylated reference genome (negative strand RM) 308; and aligning a low-methylated reverse strand probe sequence (OB / ctOB) 320 with a fully non-methylated reference genome (negative strand RU) 312. Through this mapping relationship, it is possible to accurately predict whether a probe sequence will non-specifically bind to other regions of the reference genome under simulated methylation backgrounds.
[0037] In some embodiments, homology comparison can be used to determine homology sequence similarity 322, homology sequence alignment position 324, alignment mismatch 326, and alignment score and statistical significance 328, thereby outputting a single probe sequence off-target risk assessment table 330. In embodiments of this disclosure, the BLASTn algorithm can be used to perform homology comparison on candidate probe sequences. BLAST (Basic Local Alignment Search Tool) is a commonly used sequence alignment algorithm for quickly finding potential homologous sequences of nucleic acid or protein sequences in a database. Table 1 shows exemplary parameter settings for the blastn tool according to some embodiments of this disclosure.
[0038] Table 1
[0039] In embodiments of this disclosure, a global alignment operation can be performed on the probe sequence and the reference genome. As shown in Table 1, in some embodiments of this disclosure, low-complexity region filtering in the alignment algorithm can be disabled by setting `dust` to `no`. In some embodiments, soft masking can also be disabled by setting `soft_masking` to `False`. In this way, it can be ensured that all bases in the reference genome and the entire probe sequence participate in the alignment operation. In some embodiments, a standard nucleotide alignment mode can be used by setting `task` to `blastn`. In some embodiments, tabular results can be output by setting `outfmt` to 6. In some embodiments, the tabular results include alignment position, similarity, mismatch, alignment score, and statistical significance. In some embodiments, parallel computation can be accelerated by setting `num_threads` to 12, improving the alignment efficiency of large-scale probe sequence sets.
[0040] In some embodiments, probe sequences can be numbered based on at least one of the following: genomic location of the probe sequence, orientation of the strand to which the probe sequence belongs, target region of the probe sequence, and methylation state of the probe sequence, thereby facilitating the querying and indexing of specific probe sequences. For example, for a probe sequence numbered 1:3022066-3022387_sliding:1-120_OT_Methyl, in this numbering method, the first field (e.g., 1:3022066-3022387) indicates the genomic coordinates of the probe sequence. 1:3022066-3022387 indicates that the probe sequence corresponds to the genomic location on human chromosome 1 (chr1), with a start site of 3,022,066 and an end site of 3,022,387, and a interval length of 322 bp. The second field (e.g., sliding:1-120) indicates the relative position of the probe sequence within the sliding window. sliding:1-120 indicates that the probe sequence is one of the sub-fragments generated using the sliding window strategy within the target interval, and is located at position 1 to 120. The third field (e.g., OT) indicates the orientation of the strand to which the probe sequence belongs. OT represents that the probe sequence is designed on the top strand (top). The fourth field (e.g., Methyl) indicates the methylation state of the probe sequence, where Methyl indicates that the probe sequence is in a hypermethylated state.
[0041] Table 2 shows exemplary alignment results according to some embodiments of the present disclosure. Taking the candidate probe sequence “1:3022066-3022387_sliding:1-120_OT_Methyl” as an example, the BLAST homology alignment method was used to search for its potential hitting sequences in the methylated transformation reference genome (including the forward OT strand and the reverse OB strand). The alignment results show that the probe sequence has multiple hitting sites in the reference genome. By statistically analyzing the alignment score (bit score), the percentage similarity of aligned bases (%), the length of the aligned region, and its location information in the genome for each hitting entry, the compiled results shown in Table 2 can be obtained.
[0042] Table 2
[0043] The results shown in Table 2 include the hit sequence ID, the start and end points of the alignment on the chromosome, the length of the alignment region, base similarity, number of mismatches, number of gaps, and the final alignment score. Taking the results in Table 2 as an example, the first hit record shows that the query sequence completely matches the sequence at positions 3022067 to 3022186 on chr1_CT_converted, with an alignment length of 120 bp, a similarity of 100%, and no mismatches or gaps, indicating that this position is its original source site and has the highest binding specificity. The second hit record shows that the query sequence can also partially match the 55061876–55061944 region of chr1_CT_converted, which is a potential off-target site; the two sequences are aligned within a 71 bp range, with a similarity of 81.69%, containing 11 mismatched bases and 2 gaps, suggesting a certain degree of sequence homology but low binding stability. The remaining hit records follow the same pattern, corresponding to other possible cross-reaction sites in the genome.
[0044] In some embodiments, the comprehensive alignment score output by the alignment algorithm (such as the bit score of BLAST) can be used as a unified quantification standard. In some embodiments, an alignment score threshold (e.g., 60 points) can be set, and hit results with alignment scores higher than this threshold are considered as potential off-target sites with significant binding risk. In some embodiments, method 300 counts the number of high-risk hits corresponding to each query sequence (e.g., the number of hits with alignment scores > 60) and generates a probe sequence specificity score accordingly. For example, if the number of high-risk hits exceeds a preset upper limit (e.g., 20), the probe sequence is determined to have a high risk of non-specific binding, and optimization or removal is recommended. This method realizes the transformation from multi-dimensional alignment parameters to a single risk score, significantly improving the efficiency and reliability of probe sequence design.
[0045] The above method enables the efficient construction of a reference genome suitable for probe sequence off-target risk assessment, reducing computational resource consumption and maintenance costs. This method achieves accurate identification and quantitative assessment of potential off-target sites in probe sequences, effectively reducing the risk of false matches or missed detections due to unreasonable parameter weights, and significantly improving the accuracy and sensitivity of probe sequence off-target detection.
[0046] The above combination Figure 3 A flowchart illustrating an example process for predicting the off-target risk of probe sequences according to some embodiments of this disclosure is provided below. Figure 4 A schematic diagram illustrating a method 400 for constructing a reference genome according to some embodiments of the present disclosure. Figure 4 Example method 400 in the example can be derived from Figure 1 The computational device 102 or any suitable device may be used for processing. This example method 400 describes the process of constructing a construction reference genome according to embodiments of this disclosure.
[0047] At box 402, the standard human genome sequence file 402 is read. In embodiments of this disclosure, to simulate the base conversion process after bisulfite treatment in a fully hypomethylated state, all C bases in the genome sequence can be converted to T bases 404. The resulting sequence represents the positive strand in a fully demethylated state, i.e., a positive strand fully hypomethylated sequence. At box 406, the positive strand fully hypomethylated sequence file is output. In some embodiments, all G bases in the genome sequence can be converted to A bases 408, thereby simulating the result of bisulfite treatment on the negative strand, i.e., a negative strand fully hypomethylated sequence, which can be used as a reverse matching reference. At box 410, the negative strand fully hypomethylated sequence file is output. In some embodiments, the same genome can be processed sequentially through the steps shown in boxes 404 and 408 (the order of boxes 404 and 408 is not limited), and positive / negative hypomethylated reference sequence files are output.
[0048] At box 412, scan the genome to locate the base coordinates of all CpG dinucleotides. At box 414, Site marking and temporary masking can be applied to the regions containing CpG sites to prevent erroneous modification during subsequent conversions. For example, the corresponding regions can be replaced with "N" or "#" symbols. In box 416, for all non-CpG regions outside the mask (simulating CpG in a methylated protected state), C bases are converted to T bases. In box 418, the original CpG sequence of the masked region is restored (i.e., C and G remain unchanged). In box 420, the positive reference sequence retaining the original methylation state only in the CpG regions is output, i.e., the positive-strand full-hypermethylated sequence file. In box 422, the genome is scanned, and for non-CpG regions outside the mask (simulating CpG in a methylated protected state), G bases are converted to A bases. In box 418, the original CpG complementary sequence of the masked region is restored (i.e., G and C remain unchanged), resulting in the converted negative-strand hypermethylated sequence. In box 420, the negative-strand hypermethylated sequence file is output. This sequence can be used as the full-hypermethylated reference strand for the antisense direction. In some embodiments, the same reference sequence can be processed sequentially through the steps shown in boxes 416 and 422 (the order of boxes 416 and 422 is not limited), and positive / negative hypermethylated reference sequence files are output.
[0049] Taking the genome sequence chr1:3022066-3022387 as an example, the standard human reference genome sequence and the transformed genome sequences with four different orientations and methylation states are shown in Table 3:
[0050] Table 3
[0051] Using the above methods, four reference genomes were constructed, including low-methylated positive strand, low-methylated negative strand, high-methylated positive strand, and high-methylated negative strand. This improved the construction efficiency of the reference genomes, reduced the construction cost, and provided a good starting point for assessing the off-target risk of probe sequences on the methylated genome sequences.
[0052] Figure 5 The illustration shows a schematic block diagram of an apparatus for determining the off-target probability of a probe sequence according to some embodiments of the present disclosure. Figure 5 As shown, device 500 can be implemented through software, hardware, or a combination of both. For example... Figure 5 As shown, the device 500 includes a first genome construction module 502 and an off-target probability determination module 504.
[0053] The first genome construction module 502 is configured to construct a first genome, which includes at least one of hypomethylated positive strand, hypomethylated negative strand, hypermethylated positive strand, and hypermethylated negative strand. The off-target probability determination module 504 is configured to determine the off-target probability of the probe sequence based on the constructed first genome.
[0054] In some embodiments, the probe sequence includes at least one of a first probe sequence, a second probe sequence, a third probe sequence, and a fourth probe sequence, wherein the first probe sequence is used to detect hypomethylated positive chains, the second probe sequence is used to detect hypomethylated negative chains, the third probe sequence is used to detect hypermethylated positive chains, and the fourth probe sequence is used to detect hypermethylated negative chains. The off-target probability determination module 504 is configured to determine the off-target risk of the first probe sequence based on the hypomethylated positive chain; determine the off-target risk of the second probe sequence based on the hypomethylated negative chain; determine the off-target risk of the third probe sequence based on the hypermethylated positive chain; or determine the off-target risk of the fourth probe sequence based on the hypermethylated negative chain.
[0055] In some embodiments, the first genome construction module 502 further includes a hypomethylated positive strand construction module, configured to construct a first genome comprising a hypomethylated positive strand by converting cytosine at CHH sites, CHG sites, and CpG sites on the positive strand of the second genome to thymine; and a hypomethylated negative strand construction module, configured to construct a first genome comprising a hypomethylated negative strand by converting guanine at CHH sites, CHG sites, and CpG sites on the negative strand of the second genome to adenine.
[0056] In some embodiments, the first genome construction module 502 further includes a hypermethylated positive strand construction module, configured to construct a first genome comprising a hypermethylated positive strand by converting cytosine at CHH sites and CHG sites on the positive strand of the second genome to thymine; and a hypermethylated negative strand construction module, configured to construct a first genome comprising a hypermethylated negative strand by converting guanine at CHH sites and CHG sites on the negative strand of the second genome to adenine.
[0057] In some embodiments, the first genome construction module 502 further includes a location determination module configured to determine the location of CpG sites on the second genome; and a masking module configured to mask the region where the CpG sites on the second genome are located based on the location.
[0058] In some embodiments, the hypermethylated positive strand building module further includes a third conversion module configured to convert cytosine on the positive strand of an unmasked region of the genome to thymine. In some embodiments, the hypermethylated negative strand building module further includes a fourth conversion module configured to convert guanine on the negative strand of an unmasked region of the genome to adenine.
[0059] In some embodiments, the off-target probability determination module 504 further includes an alignment module configured to perform an alignment operation on the first genome and the probe sequence; a similarity determination module configured to determine the similarity between the probe sequence and the first genome based on the alignment results of the alignment operation; and a probability determination module configured to determine the off-target probability of the probe sequence based on the similarity.
[0060] In some embodiments, the alignment operation is a homologous sequence alignment operation, the alignment result includes an alignment score for the hit sequence, and the off-target probability is positively correlated with the alignment score.
[0061] In some embodiments, the alignment module includes a global alignment module configured to perform a global alignment of the probe sequence and the first genome.
[0062] In some embodiments, the alignment module is further configured to ensure that all bases in the first genome and the full sequence of the probe sequence participate in the alignment operation by disabling the soft mask of the first genome and turning off the low-complexity region filtering in the alignment algorithm.
[0063] In some embodiments, the off-target probability determination module 504 includes a high-risk sequence determination module, configured to determine the hit sequence as a high-risk hit sequence in response to the alignment score of the hit sequence being greater than a preset alignment score threshold; and a high-risk probability module, configured to determine the off-target probability of the probe sequence based on the high-risk hit sequence.
[0064] In some embodiments, the off-target probability determination module 504 includes a high-risk sequence quantity module, configured to determine the off-target probability of a probe sequence based on the number of high-risk hit sequences.
[0065] In some embodiments, the device 500 further includes a probe sequence numbering module configured to number the probe sequence based on at least one of the genomic location of the probe sequence, the orientation of the strand to which the probe sequence belongs, the target region of the probe sequence, and the methylation state of the probe sequence, the numbering being used to obtain the probe sequence.
[0066] Figure 6 A block diagram of a device 600 capable of implementing various embodiments of the present disclosure is shown. The device 600 can be implemented to perform the methods described above. Figure 6As shown, device 600 includes a central processing unit (CPU) and / or a graphics processing unit (GPU) 601, which can perform various appropriate actions and processes according to computer program instructions stored in read-only memory (ROM) 602 or loaded from storage unit 608 into random access memory (RAM) 603. Various programs and data required for the operation of device 600 can also be stored in RAM 603. CPU / GPU 601, ROM 602, and RAM 603 are interconnected via bus 604. Input / output (I / O) interface 605 is also connected to bus 604. Although not shown in... Figure 6 As shown, device 600 may also include a coprocessor.
[0067] Multiple components in device 600 are connected to I / O interface 605, including: input unit 606, such as keyboard, mouse, etc.; output unit 607, such as various types of monitors, speakers, etc.; storage unit 608, such as disk, optical disk, etc.; and communication unit 609, such as network card, modem, wireless transceiver, etc. Communication unit 609 allows device 600 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.
[0068] The various methods or processes described above can be executed by CPU / GPU 601. For example, in some embodiments, the methods can be implemented as computer software programs tangibly contained in a machine-readable medium, such as storage unit 608. In some embodiments, part or all of the computer program can be loaded and / or installed on device 600 via ROM 602 and / or communication unit 609. When the computer program is loaded into RAM 603 and executed by CPU / GPU 601, one or more steps or actions in the methods or processes described above can be performed.
[0069] In some embodiments, the methods and processes described above can be implemented as a computer program product. The computer program product may include a computer-readable storage medium having computer-readable program instructions loaded thereon for performing various aspects of this disclosure.
[0070] Computer-readable storage media can be tangible devices capable of holding and storing instructions for use by an instruction execution device. Computer-readable storage media can be, for example, but not limited to, electrical storage devices, magnetic storage devices, optical storage devices, electromagnetic storage devices, semiconductor storage devices, or any suitable combination thereof. More specific examples (a non-exhaustive list) of computer-readable storage media include: portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), static random access memory (SRAM), portable compact disc read-only memory (CD-ROM), digital multifunction disc (DVD), memory sticks, floppy disks, mechanical encoding devices, such as punch cards or recessed protrusions storing instructions thereon, and any suitable combination thereof. The computer-readable storage media used herein are not to be construed as transient signals themselves, such as radio waves or other freely propagating electromagnetic waves, electromagnetic waves propagating through waveguides or other transmission media (e.g., light pulses through fiber optic cables), or electrical signals transmitted through wires.
[0071] The computer-readable program instructions described herein can be downloaded from computer-readable storage media to various computing / processing devices, or downloaded via a network, such as the Internet, a local area network (LAN), a wide area network (WAN), and / or a wireless network, to an external computer or external storage device. The network may include copper cables, fiber optic cables, wireless transmission, routers, firewalls, switches, gateway computers, and / or edge servers. A network adapter card or network interface in each computing / processing device receives the computer-readable program instructions from the network and forwards them to the computer-readable storage media in the respective computing / processing device.
[0072] Computer program instructions used to perform the operations of this disclosure may be assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine-dependent instructions, microcode, firmware instructions, status setting data, or source code or object code written in any combination of one or more programming languages, including object-oriented programming languages and conventional procedural programming languages. The computer-readable program instructions may execute entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving a remote computer, the remote computer may be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or may be connected to an external computer (e.g., via the Internet using an Internet service provider). In some embodiments, electronic circuitry, such as programmable logic circuitry, field-programmable gate arrays (FPGAs), or programmable logic arrays (PLAs), is personalized by utilizing the status information of the computer-readable program instructions to execute the computer-readable program instructions, thereby implementing various aspects of this disclosure.
[0073] These computer-readable program instructions can be provided to a processing unit of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus to produce a machine such that, when executed by the processing unit of the computer or other programmable data processing apparatus, they create means for implementing the functions / actions specified in one or more blocks of the flowchart and / or block diagram. These computer-readable program instructions can also be stored in a computer-readable storage medium that causes a computer, programmable data processing apparatus, and / or other device to operate in a particular manner. Thus, the computer-readable medium storing the instructions comprises an article of manufacture that includes instructions for implementing aspects of the functions / actions specified in one or more blocks of the flowchart and / or block diagram.
[0074] Computer-readable program instructions may also be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable data processing apparatus, or other device to produce a computer-implemented process, thereby causing the instructions executed on the computer, other programmable data processing apparatus, or other device to perform the functions / actions specified in one or more boxes of a flowchart and / or block diagram.
[0075] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of devices, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of an instruction containing one or more executable instructions for implementing a specified logical function. In some alternative implementations, the functions marked in the blocks may occur in a different order than those marked in the drawings. For example, two consecutive blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, may be implemented using a dedicated hardware-based system that performs the specified function or action, or using a combination of dedicated hardware and computer instructions.
[0076] The various embodiments of this disclosure have been described above. These descriptions are exemplary and not exhaustive, nor are they limited to the disclosed embodiments. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the described embodiments. The terminology used herein is chosen to best explain the principles, practical applications, or technical improvements to the technology in the market, or to enable others skilled in the art to understand the embodiments disclosed herein.
Claims
1. A method for determining the off-target probability of a probe sequence, comprising: Construct a first genome, the first genome comprising at least one of a hypomethylated positive strand, a hypomethylated negative strand, a hypermethylated positive strand, and a hypermethylated negative strand; as well as Based on the constructed first genome, the off-target probability of the probe sequence is determined.
2. The method according to claim 1, wherein the probe sequence comprises at least one of a first probe sequence, a second probe sequence, a third probe sequence, and a fourth probe sequence, wherein the first probe sequence is used to detect the hypomethylated positive chain, the second probe sequence is used to detect the hypomethylated negative chain, the third probe sequence is used to detect the hypermethylated positive chain, and the fourth probe sequence is used to detect the hypermethylated negative chain, and determining the off-target risk of the probe sequence includes at least one of the following: Based on the hypomethylated positive strand, the off-target risk of the first probe sequence is determined; Based on the hypomethylated negative chain, the off-target risk of the second probe sequence is determined; Based on the highly methylated positive strand, the off-target risk of the third probe sequence is determined; or Based on the highly methylated negative chain, the off-target risk of the fourth probe sequence is determined.
3. The method of claim 1, wherein constructing the first genome comprises: The first genome comprising the hypomethylated positive strand was constructed by converting cytosine at the CHH, CHG, and CpG sites on the positive strand of the second genome to thymine. as well as The first genome comprising the hypomethylated negative strand is constructed by converting guanine at the CHH, CHG, and CpG sites on the negative strand of the second genome to adenine.
4. The method of claim 1, wherein constructing the first genome comprises: The first genome comprising the highly methylated positive strand is constructed by converting cytosine at the CHH and CHG sites on the positive strand of the second genome to thymine; as well as The first genome comprising the highly methylated negative strand is constructed by converting guanine at the CHH and CHG sites on the negative strand of the second genome to adenine.
5. The method of claim 4, wherein constructing the first genome comprises: Determine the location of the CpG site on the second genome; as well as Based on the location, the region where the CpG site is located on the second genome is masked.
6. The method of claim 5, wherein constructing the hypermethylated positive strand comprises converting the cytosine on the positive strand in an unmasked region of the genome to the thymine, and constructing the hypermethylated negative strand comprises converting the guanine on the negative strand in an unmasked region of the genome to the adenine.
7. The method of claim 1, wherein determining the off-target probability of the probe sequence based on the constructed first genome and the probe sequence comprises: The first genome and the probe sequence are compared. Based on the alignment results of the alignment operation, the similarity between the probe sequence and the first genome is determined; as well as Based on the similarity, the off-target probability of the probe sequence is determined.
8. The method of claim 7, wherein the alignment operation is a homologous sequence alignment operation, the alignment result includes an alignment score for the hit sequence, and the off-target probability is positively correlated with the alignment score.
9. The method according to claim 8, wherein the comparison operation comprises: A global alignment was performed between the probe sequence and the first genome.
10. The method of claim 9, wherein the comparison operation comprises: By disabling the soft mask of the first genome and turning off the low-complexity region filtering in the alignment algorithm, it is ensured that all bases in the first genome and the full sequence of the probe sequence participate in the alignment operation.
11. The method of claim 8, wherein determining the off-target probability of the probe sequence comprises: In response to the alignment score of the hit sequence being greater than a preset alignment score threshold, the hit sequence is identified as a high-risk hit sequence; as well as Based on the high-risk hit sequence, the off-target probability of the probe sequence is determined.
12. The method of claim 11, wherein determining the off-target probability of the probe sequence based on the number of high-risk hit sequences comprises: The off-target probability of the probe sequence is determined based on the number of high-risk hit sequences.
13. The method of claim 7, further comprising: The probe sequence is numbered based on at least one of the following: the genomic location of the probe sequence, the orientation of the strand to which the probe sequence belongs, the target region of the probe sequence, and the methylation state of the probe sequence. The numbering is used to obtain the probe sequence.
14. An apparatus for determining the off-target probability of a probe sequence, comprising: A first genome construction module is configured to construct a first genome, the first genome comprising at least one of a hypomethylated positive strand, a hypomethylated negative strand, a hypermethylated positive strand, and a hypermethylated negative strand; and An off-target probability determination module is configured to determine the off-target probability of the probe sequence based on the constructed first genome.
15. An electronic device comprising: At least one processor; as well as A memory for storing at least one program, which, when executed by the at least one processor, causes the at least one processor to implement the method according to any one of claims 1-13.