Method, computing device and computer storage medium for determining positive breakpoints

By acquiring and analyzing the mismatch sequence consistency supporting the read length, calculating the slope and number, and automatically determining the positive breakpoints of the genome, the problems of time-consuming and labor-intensive methods and high false positive rates in traditional methods are solved, achieving efficient and accurate breakpoint judgment.

CN114496087BActive Publication Date: 2025-09-30SHANGHAI ORIGIMED CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202210073366.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-01-21
Publication Date
2025-09-30
Estimated Expiration
2042-01-21

AI Technical Summary

Technical Problem

Traditional methods are unable to automatically and accurately determine the authenticity of positive breakpoints in the genome, resulting in many false positive judgments and time-consuming and labor-intensive results, making it difficult to meet large-scale clinical needs.

Method used

By obtaining supporting read lengths across the same breakpoint, determining the read length that meets the mismatch sequence consistency condition, recording the mismatch sequence length, calculating the slope, and automatically judging the positive breakpoint based on the slope and the number of read lengths.

Benefits of technology

It realizes automatic and accurate judgment of the truth or falsity of breakpoints, improves judgment efficiency and accuracy, and is suitable for large-scale clinical needs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114496087B_ABST
    Figure CN114496087B_ABST
Patent Text Reader

Abstract

The present disclosure relates to a method, computing device and computer storage medium for determining positive breakpoints. The method includes: obtaining all supporting reads across the same breakpoint, one end of which is aligned with the first genome and the other end is not aligned with the second genome, so as to determine the read length that meets the mismatch sequence consistency condition in the supporting read length; recording the mismatch sequence length of each read length in the read length that meets the predetermined mismatch sequence consistency condition, so as to determine the read length that matches the predetermined mismatch sequence, and the predetermined mismatch sequence is determined based on the recorded mismatch sequence length; based on the mismatch sequence length, sorting the read lengths that match the predetermined mismatch sequence to calculate the slope about the same breakpoint; and determining the positive breakpoint based on the calculated slope and the number of read lengths that match the predetermined mismatch sequence. The present disclosure helps to quickly and correctly identify false positives and can automatically and accurately judge the true and false of breakpoints.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates generally to biological information processing, and in particular, to methods, computing devices, and computer storage media for determining positive breakpoints. Background Art

[0002] Traditional methods for identifying positive breakpoints include two main approaches. The first involves using rearrangement calling software to calculate the statistical distribution of bases near the breakpoint to confirm suspected breakpoints. The second involves manually using IGV software to identify mismatches in support reads.

[0003] In the first scheme, the breakpoints are confirmed mainly based on the fact that the bases at the same genomic coordinates of the mismatched parts of the reads across the breakpoints are the same and are uniquely aligned to the breakpoint position of another recombinant gene, without making a true or false judgment on the breakpoints. Therefore, many false positive judgment results about the breakpoints will be generated.

[0004] In the second approach, breakpoints are determined by sequentially staggering the start and end points of reads of different lengths after sequencing the genome. This requires manual interpretation using IGV software, which is time-consuming and labor-intensive, resulting in low throughput and making it difficult to meet large-scale clinical needs.

[0005] In summary, the traditional solution for determining positive breakpoints has the disadvantage that it is difficult to automatically and accurately determine whether the breakpoint is true or false. Summary of the Invention

[0006] The present disclosure provides a method, computing device, and computer storage medium for determining a positive breakpoint, which can automatically and accurately determine whether the breakpoint is true or false.

[0007] According to the first aspect of the present disclosure, a method for determining a positive breakpoint is provided. The method comprises: obtaining all supporting reads across the same breakpoint, one end of which is aligned with the first genome and the other end is not aligned with the second genome, so as to determine the read length that meets the mismatch sequence consistency condition in the supporting read length; recording the mismatch sequence length of each read length in the read length that meets the predetermined mismatch sequence consistency condition, so as to determine the read length that matches the predetermined mismatch sequence, and the predetermined mismatch sequence is determined based on the recorded mismatch sequence length; based on the mismatch sequence length, sorting the read lengths that match the predetermined mismatch sequence to calculate the slope about the same breakpoint; and determining the positive breakpoint based on the calculated slope and the number of read lengths that match the predetermined mismatch sequence.

[0008] According to a second aspect of the present invention, a computing device is further provided, comprising: a memory configured to store one or more computer programs; and a processor coupled to the memory and configured to execute the one or more programs so that the device performs the method of the first aspect of the present disclosure.

[0009] According to a third aspect of the present disclosure, a non-transitory computer-readable storage medium is further provided, wherein the non-transitory computer-readable storage medium stores machine-executable instructions, which, when executed, enable a machine to perform the method of the first aspect of the present disclosure.

[0010] In some embodiments, determining a read length that meets the mismatch sequence consistency condition in the supported read lengths includes: in response to determining that the mismatch sequences of the supported read lengths have the same bases at the same genomic coordinate position, determining that the current supported read length meets the mismatch sequence consistency condition.

[0011] In some embodiments, determining a positive breakpoint based on the calculated slope and the number of reads that match the predetermined mismatch sequence includes: determining whether the slope is zero; and in response to determining that the slope is zero, determining that the same breakpoint associated with the slope is a false positive breakpoint.

[0012] In some embodiments, determining a positive breakpoint based on the calculated slope and the number of reads matching the predetermined mismatch sequence includes: in response to determining that the slope is not zero, determining whether the number of breakpoints associated with the same breakpoint is greater than or equal to a first predetermined number threshold; in response to determining that the number of breakpoints is greater than or equal to the first predetermined number threshold, determining whether the number of reads matching the predetermined mismatch sequence is less than a second predetermined number threshold; and in response to the number of reads matching the predetermined mismatch sequence being less than the second predetermined number threshold, determining that the same breakpoint associated with the slope is a false positive breakpoint.

[0013] In some embodiments, determining a positive breakpoint based on the calculated slope and the number of reads matching a predetermined mismatch sequence includes: in response to determining that the slope is not zero and the number of breakpoints associated with the non-zero slope is between a third predetermined number threshold and a first number threshold, determining whether the number of reads matching the predetermined mismatch sequence for each breakpoint is greater than or equal to a fourth predetermined number threshold, the first number threshold being greater than the third predetermined number threshold; in response to determining that the number of reads matching the predetermined mismatch sequence for each breakpoint is greater than or equal to the fourth predetermined number threshold, determining whether the mismatch sequence is a base-imbalanced repeat unit sequence; and in response to determining that the mismatch sequence is not a base-imbalanced repeat unit sequence, determining the same breakpoint associated with the slope as a true positive breakpoint.

[0014] In some embodiments, determining a positive breakpoint based on the calculated slope and the number of reads matching a predetermined mismatch sequence includes: determining the same breakpoint associated with the slope as a true positive breakpoint in response to determining that the slope is not and any of the following conditions is met: the slope is within a first predetermined slope range, and the number of reads matching the predetermined mismatch sequence is less than or equal to a predetermined read number threshold; and the slope is greater than or equal to a second predetermined slope threshold, and the number of reads matching the predetermined mismatch sequence is less than or equal to a predetermined read number threshold.

[0015] In some embodiments, sorting the reads that match a predetermined mismatch sequence based on the mismatch sequence length so as to calculate the slope about the same breakpoint includes: sorting the reads that match the predetermined mismatch sequence in ascending order according to increasing mismatch sequence length; and calculating the slope about the same breakpoint based on the sorted mismatch sequence lengths and sorting order data of the reads that match the predetermined mismatch sequence.

[0016] In some embodiments, calculating the slope about the same breakpoint based on the sorted mismatch sequence lengths and sorting order data of the read lengths that match the predetermined mismatch sequence includes: calculating the slope about the same breakpoint based on the sorted mismatch sequence lengths and sorting order data of the read lengths that match the predetermined mismatch sequence via a linear regression function.

[0017] In some embodiments, obtaining all supporting reads across the same breakpoint, one end of which is aligned to the first genome and the other end of which is not aligned to the second genome, includes: based on the comparison result data of the sequencing sequence of the sample to be tested and the whole genome reference sequence, extracting supporting reads that meet the following two conditions to form a sub-alignment sequence: the lengths of the genome clusters after the rearrangement of the two gene intervals are all within a predetermined range; and the rearranged two gene intervals have a pairwise relationship with imbalanced reads; based on the sub-alignment sequence, determining all supporting reads across the same breakpoint, one end of which is aligned to the first genome and the other end of which is not aligned to the second genome. In some embodiments, the predetermined mismatch sequence is determined based on the recorded mismatch sequence lengths, including any of the following: the predetermined mismatch sequence is determined based on the minimum mismatch sequence length among the recorded mismatch sequence lengths; the predetermined mismatch sequence is determined based on the maximum mismatch sequence length among the recorded mismatch sequence lengths; or the predetermined mismatch sequence is determined based on the mismatch sequence length with the highest frequency among the recorded mismatch sequence lengths.

[0018] This summary is provided to introduce a selection of concepts in a simplified form that are further described below in the detailed description. This summary is not intended to identify key features or essential features of the disclosure, nor is it intended to limit the scope of the disclosure. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] Figure 1 A schematic diagram of a system for implementing a method for determining a positive breakpoint according to an embodiment of the present disclosure is shown.

[0020] Figure 2 A flowchart of a method for determining a positive breakpoint according to an embodiment of the present disclosure is shown.

[0021] Figure 3 A flowchart of a method for determining a positive breakpoint according to an embodiment of the present disclosure is shown.

[0022] Figure 4 A flowchart of a method for calculating a slope about a same breakpoint according to an embodiment of the present disclosure is shown.

[0023] Figure 5 A schematic diagram showing a method for calculating the slope about the same breakpoint.

[0024] Figure 6 A schematic diagram showing a case where the slope is zero about the same breakpoint.

[0025] Figure 7 A schematic diagram showing a method for calculating the slope about the same breakpoint;

[0026] Figure 8 A block diagram schematically illustrates an electronic device suitable for implementing the embodiments of the present disclosure; and

[0027] In the various drawings, the same or corresponding reference numerals denote the same or corresponding parts. DETAILED DESCRIPTION

[0028] The preferred embodiments of the present disclosure will be described in more detail below with reference to the accompanying drawings. Although preferred embodiments of the present disclosure are shown in the accompanying drawings, it should be understood that the present disclosure can be implemented in various forms and should not be limited by the embodiments set forth herein. Rather, these embodiments are provided to make the present disclosure more thorough and complete, and to fully convey the scope of the present disclosure to those skilled in the art.

[0029] Regarding the source of alignment information, in the following embodiments, prior to proceeding, paired-end sequencing is performed on each sequencing fragment obtained by probe capture of the test sample to obtain paired-end data, which includes pairs of paired read lengths; the obtained paired-end data is then aligned to the reference genome to obtain alignment information.

[0030] As used herein, the term "including" and its variations represent open inclusion, i.e., "including but not limited to." Unless otherwise stated, the term "or" means "and / or." The term "based on" means "based at least in part on." The terms "one example embodiment" and "an embodiment" mean "at least one example embodiment." The term "another embodiment" means "at least one additional embodiment." The terms "first," "second," etc. may refer to different or identical objects.

[0031] In addition, the term "sequencing fragment" used in this article is generally an RNA library constructed by using a library construction process adapted to a sequencing platform from a sample to be tested from a target individual, and its composition is a random RNA fragment of a certain length.

[0032] The term "read length" used in this article refers to the sequencing sequence obtained by sequencing the ends of the sequencing fragment. The term "paired read length" means: two sequencing sequences from both ends of the same sequencing fragment obtained by double-end sequencing. The paired read lengths can be divided into different types of paired read lengths according to the results of the alignment of the paired read lengths on the whole genome reference sequence. The term "misaligned read length pair" or "paired misaligned read length" means: the results of the above alignment show that the paired read lengths cannot be aligned to the whole genome reference sequence according to the normal mapping spacing and direction, or one or both of the read lengths cannot be completely aligned to the whole genome reference sequence at one position, or the two paired read lengths cannot be aligned to the same chromosome, usually referring to PE discordant reads and clipped reads pairs.

[0033] The term "breakpoint" refers to the site where consecutive bases that match and mismatch the reference sequence are located within a read. The term "same breakpoint" refers to all breakpoints that map to the same position on the whole-genome reference sequence.

[0034] As described above, for the traditional scheme of using rearrangement calling software to calculate the statistical distribution type of bases near the breakpoint to confirm suspected breakpoints, it mainly confirms the breakpoint based on the same bases at the same genomic coordinate position of the mismatched parts of the support read lengths across the breakpoints, and uniquely aligns to the breakpoint position of another recombinant gene, without making a judgment on the true or false breakpoints. Therefore, many false positive judgment results about the breakpoints will be generated. For the traditional scheme that mainly uses IGV software to manually interpret the mismatch of the support read lengths, it is necessary to manually confirm the breakpoint by determining the sequential and orderly staggering of the start or end points of different read lengths. Therefore, the judgment is time-consuming and labor-intensive, with low throughput, and it is difficult to meet large-scale clinical needs. Therefore, the traditional scheme for determining positive breakpoints is difficult to automatically and accurately judge the true or false of the breakpoints.

[0035] In order to at least partially solve the above-mentioned problems and one or more of other potential problems, an example embodiment of the present disclosure proposes a scheme for determining positive breakpoints. The scheme includes: determining the read lengths that meet the mismatch sequence consistency conditions in the supporting read lengths across the same breakpoint, and determining the read lengths that match the predetermined mismatch sequence. The present disclosure can obtain a plurality of random events consisting of a plurality of reads that meet the mismatch sequence consistency conditions, representing support for the same rearrangement. In addition, by sorting the read lengths determined to match the predetermined mismatch sequence in order to calculate the slope about the same breakpoint, and determining the positive breakpoint based on the calculated slope and the number of read lengths that match the predetermined mismatch sequence, the present disclosure can automatically and efficiently obtain data for indicating the consistency and discrepancy of the mismatch sequence read lengths (mismatch reads, or softclip reads) on one side of the breakpoint, and then automatically and accurately judge the true or false of the breakpoint.

[0036] Figure 1 FIG. 1 is a schematic diagram of a system 100 for implementing a method for determining a positive breakpoint according to an embodiment of the present disclosure. Figure 1 As shown, the system 100 includes, for example, a computing device 110, a sequencing device 130, a bioinformatics server 140, and a network 150. The computing device 110 can exchange data with the sequencing device 130 and the bioinformatics server 140 via the network 150 in a wired or wireless manner.

[0037] The sequencing device 130 is used, for example, to sequence a sample to be tested from a target individual. For example, paired-end sequencing is performed on each sequencing fragment obtained by probe capture of the sample to be tested to obtain paired-end sequencing data. The sequencing device 130 is also used to send the sequencing sequence (e.g., paired-end data) of the sample to be tested to the computing device 110. In some embodiments, the sequencing sequence of the sample to be tested comes from the bioinformatics server 140.

[0038] Regarding the computing device 110, it is used, for example, to obtain all supporting reads across the same breakpoint, so as to determine the reads that meet the mismatch sequence consistency condition among the supporting reads; record the mismatch sequence length of each read in the reads that meet the predetermined mismatch sequence consistency condition, so as to determine the read that matches the predetermined mismatch sequence. The computing device 110 is also used to sort the reads that match the predetermined mismatch sequence based on the mismatch sequence length, so as to calculate the slope about the same breakpoint; and determine the positive breakpoint based on the calculated slope and the number of reads that match the predetermined mismatch sequence. In some embodiments, the computing device 110 may have one or more processing units, including dedicated processing units such as GPUs, FPGAs, and ASICs, and general-purpose processing units such as CPUs. In addition, one or more virtual machines may also be running on each computing device. The computing device 110 includes, for example: a read length determination unit 112 that meets the mismatch sequence consistency condition, a read length determination unit 114 that matches the predetermined mismatch sequence, a slope calculation unit 116 about the same breakpoint, and a positive breakpoint determination unit 118. The above-mentioned read length determination unit 112 that meets the mismatch sequence consistency condition, the read length unit 114 that matches the predetermined mismatch sequence, the slope calculation unit 116 for the same breakpoint, and the positive breakpoint determination unit 118 can be configured on one or more computing devices 110.

[0039] Regarding the read length determination unit 112 that meets the mismatch sequence consistency condition, it is used to obtain all supporting read lengths across the same breakpoint, one end of which is aligned with the first genome and the other end is not aligned with the second genome, so as to determine the read length that meets the mismatch sequence consistency condition among the supporting read lengths.

[0040] Regarding the read length determination unit 114 that matches the predetermined mismatch sequence, it is used to record the mismatch sequence length of each read length in the read length that meets the predetermined mismatch sequence consistency condition, so as to determine the read length that matches the predetermined mismatch sequence, and the predetermined mismatch sequence is determined based on the recorded mismatch sequence length.

[0041] The slope calculation unit 116 for the same breakpoint is configured to sort the read lengths matching the predetermined mismatch sequence based on the mismatch sequence length, so as to calculate the slope for the same breakpoint.

[0042] The positive breakpoint determining unit 118 is configured to determine a positive breakpoint based on the calculated slope and the number of reads matching the predetermined mismatch sequence.

[0043] The following will be combined Figure 2 A method 200 for determining a positive breakpoint according to an embodiment of the present disclosure is described. Figure 2 FIG. 2 is a flow chart of a method 200 for determining a positive breakpoint according to an embodiment of the present disclosure. It should be understood that the method 200 may be used, for example, in Figure 8 The electronic device 800 is described. Figure 1 The method 200 is executed at the described computing device 110. It should be understood that the method 200 may also include additional actions not shown and / or may omit actions shown, and the scope of the present disclosure is not limited in this respect.

[0044] At step 202, the computing device 110 obtains all supporting reads across the same breakpoint that are aligned to the first genome at one end and not aligned to the second genome at the other end, so as to determine the reads that meet the mismatch sequence consistency condition among the supporting reads. Figure 5 As shown, the vertical arrow 522 indicates the same breakpoint spanned by multiple different supporting reads. The left side of the same breakpoint indicates the gene segment aligned to the first genome in the different supporting reads. The right side of the same breakpoint indicates the mismatch sequence that is not aligned to the second genome in the different supporting reads.

[0045] Regarding the method for obtaining all supported read lengths, it includes, for example: first, the computing device 110 obtains the alignment result data of the double-end sequencing data and the whole genome reference sequence (the alignment result data is, for example, an input bam file). Then, based on the alignment result data, the computing device 110 extracts the supporting read lengths that meet the following two conditions to form a sub-alignment sequence: the lengths of the genomes after clustering of the two rearranged gene intervals are all within a predetermined range; and the rearranged two gene intervals have a pairwise relationship of imbalanced read lengths. The sub-alignment sequence is, for example, a subbam file for the rearranged gene. Afterwards, the computing device 110 determines all supported read lengths across the same breakpoint, one end of which is aligned to the first genome and the other end is not aligned to the second genome, based on the sub-alignment sequence. For example, soft clipreads are calculated based on the subbam file for the rearranged gene, that is, all supported read lengths across the same breakpoint, one end of which is aligned to the first genome and the other end is not aligned to the second genome, are obtained. Regarding the predetermined range, it is, for example, but not limited to, 1000bp.

[0046] Regarding the method for determining the read length that meets the mismatch sequence consistency condition, it includes, for example: the computing device 110 determines whether the mismatch sequence supporting the read length has the same base at the same genomic coordinate position; if the computing device 110 determines that the mismatch sequence supporting the read length has the same base at the same genomic coordinate position, determining that the current supported read length meets the mismatch sequence consistency condition. Figure 5 As shown in FIG, in the mismatch sequences of different soft clip reads, the bases at each vertically identical genomic coordinate position are the same, and the different soft clip reads are read lengths that meet the mismatch sequence consistency condition.

[0047] At step 204 , the computing device 110 records the mismatch sequence length of each read among the reads that meet the predetermined mismatch sequence consistency condition to determine a read that matches a predetermined mismatch sequence, where the predetermined mismatch sequence is determined based on the recorded mismatch sequence lengths.

[0048] In some embodiments, the predetermined mismatch sequence is determined based on the minimum mismatch sequence length in the recorded mismatch sequence length. For example, the mismatch sequence with the minimum mismatch sequence length is ATC, and the predetermined mismatch sequence is determined to be ATC. The computing device 110 uses the predetermined mismatch sequence ATC as a reference object to match the mismatch sequence of each read length in the read length that meets the predetermined mismatch sequence consistency condition, so as to determine the read length matched with the predetermined mismatch sequence. For example, the read lengths where the mismatch sequences ATCG and ATCT are respectively located are all determined to be read lengths that match the predetermined mismatch sequence ATC. By adopting the above-mentioned means, that is, by using the mismatch sequence with the minimum mismatch sequence length as the predetermined mismatch sequence to select the read length matched, the present disclosure can match more mismatch sequences based on relatively loose conditions, avoid the situation where individual breakpoints are not matched due to reasons such as comparison quality, and thus help avoid missed detection of breakpoints.

[0049] In some embodiments, the predetermined mismatch sequence is determined based on the maximum mismatch sequence length among the recorded mismatch sequence lengths. For example, the mismatch sequence with the maximum mismatch sequence length is ATGCTGA. The read length where the mismatch sequence ATGCT is located is not determined as the read length that matches the predetermined mismatch sequence ATGCTGA, while the read length where the mismatch sequence ATGCTGAC is located is determined to be the read length that matches the predetermined mismatch sequence ATGCTGA. By adopting the above-mentioned means, that is, by taking the mismatch sequence with the maximum mismatch sequence length as the predetermined mismatch sequence to select the matching read length, the present disclosure can match more accurate mismatch sequences based on more stringent conditions, which is conducive to improving the accuracy of the judgment result.

[0050] It should be understood that, in some embodiments, the predetermined mismatch sequence is determined based on the mismatch sequence length with the highest frequency among the mismatch sequence lengths of each read length.

[0051] Regarding the method for determining the read length that matches a predetermined mismatch sequence, it includes, for example: the computing device 110 compares the breakpoints obtained from multiple reads that meet the predetermined mismatch sequence consistency conditions, so that a data set is generated based on different reads across the same breakpoint, and the length of the mismatch sequence extracted in the data set is calculated (the mismatch sequence length is, for example, characterized by the number of mismatch bases), thereby obtaining the mismatch sequence length of different reads across the same breakpoint; the mismatch sequence corresponding to the read with the smallest mismatch sequence length is determined as the string for the predetermined mismatch sequence, and the mismatch sequences of the remaining reads are matched with the string to record the length of the mismatch sequence that matches the predetermined mismatch sequence. The length of the mismatch sequence is, for example, the number of bases in the mismatch sequence that match the string.

[0052] At step 206 , the computing device 110 sorts the reads matching the predetermined mismatch sequence based on the mismatch sequence lengths to calculate a slope about the same breakpoint.

[0053] Regarding the method for calculating the slope about the same breakpoint, for example, it includes: the computing device 110 sorts the read lengths matching the predetermined mismatch sequence in ascending order according to the order of increasing lengths of the mismatch sequence; and calculating the slope about the same breakpoint based on the lengths of the sorted mismatch sequences and the sorting order data. Figure 4 and Figure 5 The method for calculating the slope about the same breakpoint has been described and will not be repeated here.

[0054] At step 208 , the computing device 110 determines positive breakpoints based on the calculated slope and the number of reads that match the predetermined mismatch sequence.

[0055] Regarding the method for determining a positive breakpoint, for example, the method includes: the computing device 110 determines whether the slope is zero; if the computing device 110 determines that the slope is zero, determining that the same breakpoint associated with the slope is a false positive breakpoint. For example, if the computing device 110 determines that the slope is zero for the fusion gene AB, it indicates that the breakpoints of the fusion gene AB are flush. Figure 6 A schematic diagram showing the case where the slope is zero about the same breakpoint is shown, as Figure 6 As shown, this indicates that the same breakpoint associated with the zero slope is a false positive breakpoint.

[0056] If the computing device 110 determines that the slope is not zero, it determines whether the number of breakpoints associated with the same breakpoint is greater than or equal to a first predetermined number threshold (the first predetermined number threshold is, for example, but not limited to, 4); if it is determined that the number of breakpoints is greater than or equal to the first predetermined number threshold, it determines whether the number of reads matching the predetermined mismatch sequence is less than a second predetermined number threshold (the second predetermined number threshold is, for example, but not limited to, 30); if the number of reads matching the predetermined mismatch sequence is less than the second predetermined number threshold, it determines that the same breakpoint associated with the slope is a false positive breakpoint. For example, if the computing device 110 determines that the number of breakpoints with a non-zero slope for the fusion gene AB is greater than or equal to 4, and the number of reads matching the predetermined mismatch sequence is less than 30, it indicates that the breakpoints in the reads matching the predetermined mismatch sequence are chaotic, and the same breakpoint associated with the slope is determined to be a false positive breakpoint.

[0057] In some embodiments, if the computing device 110 determines that the slope is not zero and the number of breakpoints associated with the non-zero slope is between a third predetermined number threshold (the third predetermined number threshold is, for example, but not limited to, 2) and a first number threshold (the first predetermined number threshold is, for example, but not limited to, 4), determining whether the number of reads matching the predetermined mismatch sequence for each breakpoint is greater than or equal to a fourth predetermined number threshold, where the first number threshold is greater than the third predetermined number threshold; if it is determined that the number of reads matching the predetermined mismatch sequence for each breakpoint is greater than or equal to the fourth predetermined number threshold, determining whether the mismatch sequence is a base-imbalanced repeat unit sequence (the fourth predetermined number threshold is, for example, but not limited to, 5), determining whether the mismatch sequence is a base-imbalanced repeat unit sequence; and if the computing device 110 determines that the mismatch sequence is not a base-imbalanced repeat unit sequence, determining that the same breakpoint associated with the slope is a true positive breakpoint. For example, if the computing device 110 determines that the number of breakpoints with a non-zero slope for the fusion gene AB is between 2 and 4, and the number of read lengths matching the predetermined mismatch sequence for each breakpoint is greater than or equal to 5, and the mismatch sequence is not a base-unbalanced repeat unit sequence, then the same breakpoint associated with the slope is determined to be a true positive breakpoint.

[0058] In the above scheme, by determining the read lengths that meet the mismatch sequence consistency conditions among the supporting read lengths across the same breakpoint, and determining the read lengths that match the predetermined mismatch sequence, the present disclosure can obtain a plurality of random events consisting of a plurality of reads that meet the mismatch sequence consistency conditions, representing support for the same rearrangement. In addition, by sorting the read lengths that match the predetermined mismatch sequence in order to calculate the slope about the same breakpoint, and determining the positive breakpoint based on the calculated slope and the number of read lengths that match the predetermined mismatch sequence, the present disclosure can automatically and efficiently obtain data for indicating the consistency and discrepancy of the mismatch reads (or soft clip reads) on one side of the breakpoint, and then automatically and accurately judge the authenticity of the breakpoint. Experimental data show that for the fusion gene AB, based on the slope of the corresponding breakpoint and the number of read lengths of the mismatch sequence that supports the slope, the accuracy of the present disclosure in determining the positive breakpoint can reach 95%.

[0059] The following will be combined Figure 3 A method 300 for determining a positive breakpoint according to an embodiment of the present disclosure is described. Figure 3 FIG. 3 is a flow chart showing a method 300 for determining a positive breakpoint according to an embodiment of the present disclosure. It should be understood that the method 300 may be used, for example, in Figure 8 The electronic device 800 is described. Figure 1 The method 300 is executed at the described computing device 110. It should be understood that the method 300 may also include additional actions not shown and / or may omit actions shown, and the scope of the present disclosure is not limited in this respect.

[0060] At step 302 , computing device 110 determines whether the slope is zero. If computing device 110 determines that the slope is zero, the process proceeds to step 308 to determine that the same breakpoint associated with the slope is a false positive breakpoint.

[0061] At step 304, if the computing device 110 determines that the slope is not zero, it determines whether a first predetermined condition is satisfied, the first predetermined condition including any one of the following: the slope is within a first predetermined slope range, and the number of reads that match the predetermined mismatch sequence is less than or equal to a predetermined read number threshold; and the slope is greater than or equal to a second predetermined slope threshold, and the number of reads that match the predetermined mismatch sequence is less than or equal to a predetermined read number threshold.

[0062] At step 306 , if the computing device 110 determines that the first predetermined condition is satisfied, the same breakpoint associated with the slope is determined to be a true positive breakpoint.

[0063] It should be understood that when the slope is moderate (e.g., the slope is within a first predetermined slope range) and the number of reads determined to match the predetermined mismatch sequence is small (e.g., less than or equal to a predetermined read number threshold), the longest length of the mismatch sequence is moderate, the distance between the lengths of the mismatch sequence is large and relatively even, and thus the probability of a true positive breakpoint is high. When the calculated slope is large and the number of reads determined to match the predetermined mismatch sequence is small, the longest length of the mismatch sequence is large, the distance between the lengths of the mismatch sequence is large, and thus the probability of a true positive breakpoint is high.

[0064] If the computing device 110 determines that the slope is small and the number of reads determined to match the predetermined mismatch sequence is large, it means that the longest length of the mismatch sequence is short and the length difference between the lengths of different mismatch sequences is small, and the false positive probability is large. If the computing device 110 determines that the slope is moderate and the number of reads determined to match the predetermined mismatch sequence is large, it means that the longest length of the mismatch sequence is moderate and the length difference between the lengths of different mismatch sequences is small, and the false positive probability is large. If the computing device 110 determines that the slope is large and the number of reads determined to match the predetermined mismatch sequence is large, it means that the longest length of the mismatch sequence is large and the length difference between the lengths of different mismatch sequences is small, and the false positive probability is large. If the computing device 110 determines that the slope is small and the number of reads determined to match the predetermined mismatch sequence is small, it means that the longest length of the mismatch sequence is short, and the false positive probability is large.

[0065] The following will be combined Figure 4 、 Figure 5 and Figure 7 A method 400 for determining a positive breakpoint according to an embodiment of the present disclosure is described. Figure 4 A flow chart of a method 400 for determining a positive breakpoint according to an embodiment of the present disclosure is shown. Figure 7 Schematic diagram of a method for calculating the slope about the same breakpoint is shown. It should be understood that the method 400 can be used, for example, in Figure 8 The electronic device 800 is described. Figure 1 The method 40 is executed at the described computing device 110. It should be understood that the method 40 may also include additional actions not shown and / or may omit actions shown, and the scope of the present disclosure is not limited in this respect.

[0066] At step 402, the computing device 110 sorts the reads that match the predetermined mismatch sequence in ascending order of length. For example, the computing device 110 sorts the reads that match the predetermined mismatch sequence in ascending order of length of the mismatch sequence, thereby obtaining mismatch sequence length information for each read that spans the same breakpoint and has the same mismatch sequence. Figure 5It shows the reads spanning the same breakpoint, which are sorted in ascending order according to the length of the mismatch sequence of the reads matching the predetermined mismatch sequence.

[0067] At step 404, the computing device 110 calculates a slope for the same breakpoint based on the lengths of the sorted mismatch sequences and the sorting order data of the reads that match the predetermined mismatch sequences. For example, a best-fit straight line is determined using a linear regression function based on the lengths of the sorted mismatch sequences and the corresponding sorting order data to obtain coefficients of the corresponding linear regression function; the slope for the same breakpoint is determined based on the coefficients.

[0068] The following is an example of a method for calculating the slope about the same breakpoint via general linear regression using formula (1).

[0069]

[0070] In the above formula (1), Represents the length of the mismatch sequence. x represents the sorting order data of the corresponding reads. p represents the pth sorting order data. w=[w1,w2…w p ] is a coefficient. w0 represents a constant.

[0071] For example, Figure 7 As shown, the computing device 110 uses the length of the mismatch sequence as the y variable and the sorting order data of the corresponding reads as the x variable. Figure 7 Each point in the indicates the length of the mismatch sequence corresponding to the order in which the mismatch sequences of the corresponding reads are sorted. For example, mark 702 indicates that the mismatch sequence length of the read ranked first in ascending order of mismatch sequence length is 3, mark 704 indicates that the mismatch sequence length of the read ranked second in ascending order of mismatch sequence length is 4, and so on. Mark 706 indicates that the mismatch sequence length of the read ranked Nth in ascending order of mismatch sequence length is M. Using the linear regression function, the best fit line 708 is determined to obtain the coefficient w of the corresponding linear regression function; and based on the coefficient w, the slope about the same breakpoint is determined.

[0072] In the above scheme, the present disclosure conveniently obtains information that can abstractly indicate the discrepancy of mismatched sequences supporting read lengths through the calculated slope size, so as to facilitate the system to automatically identify the discrepancy of mismatched sequences supporting read lengths based at least on the slope.

[0073] Figure 8 The block diagram of an electronic device 800 suitable for implementing the embodiment of the present disclosure is schematically shown. The device 800 may be a device for implementing Figures 2 to 4The apparatus of methods 200 to 400 is shown. Figure 7 As shown, the device 800 includes a central processing unit (CPU) 801, which can perform various appropriate actions and processes according to computer program instructions stored in a read-only memory (ROM) 802 or computer program instructions loaded from a storage unit 808 into a random access memory (RAM) 803. Various programs and data required for the operation of the device 800 can also be stored in the RAM 803. The CPU 801, ROM 802, and RAM 803 are connected to each other via a bus 804. An input / output (I / O) interface 805 is also connected to the bus 804.

[0074] Multiple components in device 800 are connected to I / O interface 805, including an input unit 806, an output unit 807, and a storage unit 808. Processing unit 801 performs the various methods and processes described above, such as methods 200 through 400. For example, in some embodiments, methods 200 through 400 may be implemented as a computer software program stored on a machine-readable medium, such as storage unit 808. In some embodiments, part or all of the computer program may be loaded and / or installed onto device 800 via ROM 802 and / or communication unit 809. When the computer program is loaded into RAM 803 and executed by CPU 801, one or more operations of methods 200-500, 800, 900, and 1200 described above may be performed. Alternatively, in other embodiments, CPU 801 may be configured to perform one or more actions of methods 200 through 400 in any other suitable manner (e.g., via firmware).

[0075] It should be further noted that the present disclosure may be a method, an apparatus, a system and / or a computer program product. The computer program product may include a computer-readable storage medium carrying computer-readable program instructions for executing various aspects of the present disclosure.

[0076] Computer-readable storage medium can be a tangible device that can keep and store the instructions used by the instruction execution device.Computer-readable storage medium can be, for example, but not limited to, an electrical storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device or any suitable combination thereof.More specific examples (non-exhaustive list) of computer-readable storage medium include: a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), a static random access memory (SRAM), a portable compact disc read-only memory (CD-ROM), a digital versatile disk (DVD), a memory stick, a floppy disk, a mechanical encoding device, for example, a punch card or a convex structure in a groove having instructions stored thereon, and any suitable combination thereof.Computer-readable storage medium used herein is not interpreted as a transient signal itself, such as radio waves or other freely propagating electromagnetic waves, electromagnetic waves propagated by waveguides or other transmission media (for example, light pulses by fiber optic cables), or electrical signals transmitted by wires.

[0077] The computer-readable program instructions described herein can be downloaded from a computer-readable storage medium to each computing / processing device, or downloaded to an external computer or external storage device via a network, such as the Internet, a local area network, a wide area network, and / or a wireless network. The network can include copper transmission cables, fiber optic transmission, wireless transmission, routers, firewalls, switches, gateway computers, and / or edge servers. The network adapter card or network interface in each computing / processing device receives the computer-readable program instructions from the network and forwards the computer-readable program instructions to be stored in the computer-readable storage medium in each computing / processing device.

[0078] The computer program instructions for performing the operations of the present disclosure may be assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine-dependent instructions, microcode, firmware instructions, state setting data, or source code or object code written in any combination of one or more programming languages, including object-oriented programming languages ​​such as Smalltalk, C++, and conventional procedural programming languages ​​such as "C" language or similar programming languages. Computer-readable program instructions may be executed entirely on a user's computer, partially on a user's computer, as an independent software package, partially on a user's computer, partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer via any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., utilizing an Internet service provider to connect via the Internet). In some embodiments, an electronic circuit, such as a programmable logic circuit, a field programmable gate array (FPGA), or a programmable logic array (PLA), may be personalized by utilizing the state information of the computer-readable program instructions. The electronic circuit may execute the computer-readable program instructions, thereby realizing various aspects of the present disclosure.

[0079] Various aspects of the present disclosure are described herein with reference to flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to embodiments of the present disclosure. It should be understood that each block of the flowcharts and / or block diagrams, and combinations of blocks in the flowcharts and / or block diagrams, can be implemented by computer-readable program instructions.

[0080] These computer-readable program instructions can be provided to a processor in a voice interaction device, a general-purpose computer, a special-purpose computer, or a processing unit of another programmable data processing device, thereby producing a machine such that when these instructions are executed by the processing unit of the computer or other programmable data processing device, a device is generated that implements the functions / actions specified in one or more blocks in the flowchart and / or block diagram. These computer-readable program instructions can also be stored in a computer-readable storage medium, where these instructions cause the computer, programmable data processing device, and / or other device to operate in a specific manner. Thus, the computer-readable medium storing the instructions comprises an article of manufacture that includes instructions for implementing various aspects of the functions / actions specified in one or more blocks in the flowchart and / or block diagram.

[0081] Computer-readable program instructions may also be loaded onto a computer, other programmable data processing apparatus, or other device so that a sequence of operating steps are performed on the computer, other programmable data processing apparatus, or other device to produce a computer-implemented process, thereby causing the instructions executed on the computer, other programmable data processing apparatus, or other device to implement the functions / actions specified in one or more blocks in the flowchart and / or block diagram.

[0082] The flow charts and block diagrams in the accompanying drawings show the possible architecture, functions and operations of the devices, methods and computer program products according to multiple embodiments of the present disclosure. In this regard, each box in the flow chart or block diagram can represent a part of a module, program segment or instruction, and the part of this module, program segment or instruction contains one or more executable instructions for realizing the prescribed logical function. In some alternative implementations, the functions marked in the box can also occur in a sequence different from that marked in the accompanying drawings. For example, two consecutive boxes can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flow chart, and the combination of the boxes in the block diagram and / or flow chart can be implemented with a special hardware-based system that performs the prescribed function or action, or can be implemented with a combination of special hardware and computer instructions.

[0083] While various embodiments of the present disclosure have been described above, the foregoing description is intended to be illustrative, non-exhaustive, and not limited to the disclosed embodiments. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the described embodiments. The terminology used herein is selected to best explain the principles of the embodiments, their practical applications, or technological improvements in the marketplace, or to enable others skilled in the art to understand the embodiments disclosed herein.

[0084] The above are merely optional embodiments of the present disclosure and are not intended to limit the present disclosure. Those skilled in the art will readily appreciate that the present disclosure may be modified and varied in various ways. Any modifications, equivalent replacements, improvements, and the like made within the spirit and principles of the present disclosure shall be included within the scope of protection of the present disclosure.

Claims

1. A method for determining a positive breakpoint, comprising: Obtaining all supporting reads spanning the same breakpoint, one end of which is aligned to the first genome and the other end of which is not aligned to the second genome, so as to determine a read that meets the mismatch sequence consistency condition among the supporting reads, including determining that the current supporting read meets the mismatch sequence consistency condition in response to determining that the mismatch sequences of the supporting reads have the same base at the same genomic coordinate position; Recording the mismatch sequence length of each read length among the read lengths that meet the predetermined mismatch sequence consistency condition, so as to determine the read length that matches the predetermined mismatch sequence, wherein the predetermined mismatch sequence is determined based on the recorded mismatch sequence lengths; sorting the reads matching the predetermined mismatch sequence based on the mismatch sequence lengths to calculate a slope about the same breakpoint, wherein the slope about the same breakpoint is calculated via a linear regression function based on the mismatch sequence lengths and the sorted order data of the sorted reads matching the predetermined mismatch sequence; and Positive breakpoints were determined based on the calculated slope and the number of reads that matched the predetermined mismatch sequence.

2. The method of claim 1 , wherein determining a positive breakpoint based on the calculated slope and the number of reads matching a predetermined mismatch sequence comprises: Determine whether the slope is zero; as well as In response to determining that the slope is zero, the same breakpoint associated with the slope is determined to be a false positive breakpoint.

3. The method of claim 2 , wherein determining a positive breakpoint based on the calculated slope and the number of reads matching a predetermined mismatch sequence comprises: In response to determining that the slope is not zero, determining whether a number of breakpoints of the same breakpoint associated with the slope is greater than or equal to a first predetermined number threshold; In response to determining that the number of breakpoints is greater than or equal to a first predetermined number threshold, determining whether the number of reads matching a predetermined mismatch sequence is less than a second predetermined number threshold; as well as In response to the number of reads matching the predetermined mismatch sequence being less than a second predetermined number threshold, determining the same breakpoint associated with the slope as a false positive breakpoint.

4. The method of claim 2, wherein determining a positive breakpoint based on the calculated slope and the number of reads matching a predetermined mismatch sequence comprises: In response to determining that the slope is non-zero and the number of breakpoints associated with the non-zero slope is between a third predetermined number threshold and the first number threshold, determining whether the number of reads that match the predetermined mismatch sequence for each breakpoint is greater than or equal to a fourth predetermined number threshold, the first number threshold being greater than the third predetermined number threshold; In response to determining that the number of reads matching the predetermined mismatch sequence for each breakpoint is greater than or equal to a fourth predetermined number threshold, determining whether the mismatch sequence is a base-unbalanced repeat unit sequence; as well as In response to determining that the mismatched sequence is not a base-imbalanced repeat unit sequence, the same breakpoint associated with the slope is determined to be a true positive breakpoint.

5. The method of claim 2, wherein determining a positive breakpoint based on the calculated slope and the number of reads matching a predetermined mismatch sequence comprises: In response to determining that the slope is not and satisfying any one of the following conditions, determining the same breakpoint associated with the slope as a true positive breakpoint: The slope is within a first predetermined slope range, and the number of reads matching the predetermined mismatch sequence is less than or equal to a predetermined read number threshold; as well as The slope is greater than or equal to a second predetermined slope threshold, and the number of reads matching the predetermined mismatch sequence is less than or equal to a predetermined read number threshold.

6. The method according to claim 1 , wherein sorting the reads matching a predetermined mismatch sequence based on the mismatch sequence length so as to calculate the slope about the same breakpoint comprises: Sorting the read lengths matching the predetermined mismatch sequence in ascending order according to the order of increasing mismatch sequence length; as well as The slope about the same breakpoint is calculated based on the mismatch sequence lengths and the ranking order data of the sorted reads that match the predetermined mismatch sequence.

7. The method of claim 1 , wherein obtaining all supporting reads spanning the same breakpoint that align to the first genome at one end and do not align to the second genome at the other end comprises: Based on the comparison results of the sequencing sequence of the sample to be tested and the whole genome reference sequence, supporting reads that meet the following two conditions are extracted to form sub-alignment sequences: The lengths of the two rearranged gene intervals after the respective genome clusters are both within a predetermined range; and Rearrange the pairwise relationship between two genomic intervals with misaligned read lengths; Based on the sub-aligned sequences, all supporting reads spanning the same breakpoint that align to the first genome on one end and do not align to the second genome on the other end are determined.

8. The method according to claim 1, wherein the predetermined mismatch sequence is determined based on the recorded mismatch sequence lengths and comprises any one of the following: The predetermined mismatch sequence is determined based on the minimum mismatch sequence length among the recorded mismatch sequence lengths; The predetermined mismatch sequence is determined based on the maximum mismatch sequence length among the recorded mismatch sequence lengths; or The predetermined mismatch sequence is determined based on the mismatch sequence length with the highest occurrence frequency among the recorded mismatch sequence lengths.

9. A computing device comprising: at least one processing unit; At least one memory, the at least one memory being coupled to the at least one processing unit and storing instructions for execution by the at least one processing unit, the instructions, when executed by the at least one processing unit, causing the apparatus to perform the steps of the method according to any one of claims 1 to 8.

10. A computer-readable storage medium having a computer program stored thereon, wherein the computer program implements the method according to any one of claims 1 to 8 when executed by a machine.