Method, device, electronic device and program product for tandem repeat sequence analysis

By comparing and detecting the number of repetitions in gene fragments, computer equipment accurately identifies tandem repeats in gene sequences, solving the problem of low sensitivity in traditional methods and achieving efficient tandem repeat identification.

CN122135782APending Publication Date: 2026-06-02TIANJIN KINGMED CENT FOR CLINICAL CO LTD

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
TIANJIN KINGMED CENT FOR CLINICAL CO LTD
Filing Date
2026-02-27
Publication Date
2026-06-02

AI Technical Summary

Technical Problem

Traditional bioinformatics analysis processes have low sensitivity in identifying tandem repeats of genes, making them prone to missed detections.

Method used

The first read sequence to be analyzed and the reference gene sequence are obtained by computer equipment. Multiple detection gene fragments are compared and extracted, and their repetition counts in the reference and target sequences are detected. The repetition count threshold is used to determine whether there is tandem duplication.

Benefits of technology

It improves the sensitivity and accuracy of tandem repeat identification, reduces the risk of missed detection, and can more comprehensively capture repeat segments in the target sequence.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122135782A_ABST
    Figure CN122135782A_ABST
Patent Text Reader

Abstract

This application discloses a method, apparatus, electronic device, and program product for sequence analysis of tandem repeats. The method includes: acquiring a first read sequence and a reference gene sequence using a computer device; comparing the first read sequence and the reference gene sequence to identify one or more read sequence fragments in the first read sequence that do not match the reference gene sequence; extracting multiple different detection gene fragments from the reference gene sequence; detecting the first repeat count of each detection gene fragment in the reference gene sequence and the second repeat count in the target read sequence fragment; if the second repeat count of any detection gene fragment is not less than a threshold and the second repeat count is greater than the corresponding first repeat count, then it is determined that tandem repeats exist in the first read sequence. This application can accurately identify whether tandem repeats exist in the first read sequence, improving the sensitivity and accuracy of tandem repeat identification and reducing the risk of missed detection of tandem repeats.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of data processing technology, specifically to a method, apparatus, electronic device, and program product for analyzing tandem repeat sequences. Background Technology

[0002] Tandem repeats refer to the repeated occurrence of short sequence fragments composed of multiple base combinations within a gene's nucleotide sequence. These repeats can affect the expression, modification, and related physiological functions of the gene. Because tandem repeats in some genes exhibit varying repeat lengths and breakpoint locations, traditional bioinformatics analysis procedures often have low sensitivity in identifying tandem repeats, leading to frequent false negatives. Therefore, accurately identifying tandem repeats in gene sequences is a pressing technical challenge that needs to be addressed. Summary of the Invention

[0003] This application discloses a sequence analysis method, apparatus, electronic device, and program product for tandem repeats, which can accurately identify whether tandem repeats exist in the first read sequence, improve the sensitivity and accuracy of tandem repeat identification, and effectively reduce the risk of missed detection of tandem repeats.

[0004] The first aspect of this application discloses a method for analyzing tandem repeat sequences, the method comprising: The computer device acquires a first read sequence to be analyzed and a reference gene sequence. The first read sequence is a gene sequence fragment read from the original gene data, and the reference gene sequence is a gene sequence of the first exon region on the target chromosome. The computer device compares the first read sequence with the reference gene sequence to determine one or more read sequence fragments in the first read sequence that do not match the reference gene sequence; The computer device extracts multiple different detection gene fragments from the reference gene sequence, and detects a first repeat of each detection gene fragment in the reference gene sequence, and detects a second repeat of each detection gene fragment in a target read sequence fragment, wherein the target read sequence fragment is any one of the one or more read sequence fragments; If any detected gene fragment has a second repeat number not less than the number threshold, and the corresponding second repeat number is greater than the corresponding first repeat number, the computer device determines that there is a tandem repeat in the first read sequence. In some possible embodiments, the computer device compares the first read sequence with the reference gene sequence to determine one or more read sequence fragments in the first read sequence that do not match the reference gene sequence, including: The computer device compares the first read sequence with the reference gene sequence to determine the differential bases in the first read sequence that do not match the reference gene sequence. The computer device connects multiple differential bases with consecutive base positions or multiple differential bases with intervals between base positions not greater than a first interval threshold to obtain one or more read sequence fragments.

[0005] In some possible embodiments, after obtaining one or more read sequence fragments, the method further includes: The computer device determines a reference sequence fragment in the reference gene sequence corresponding to the target read sequence fragment based on the start and end base positions of the target read sequence fragment. The start base position of the reference sequence fragment is less than or equal to the start base position of the target read sequence fragment, and the end base position of the reference sequence fragment is greater than or equal to the end base position of the target read sequence fragment. The computer device extracts multiple different detection gene fragments from the reference gene sequence and detects the first repeat number of each detection gene fragment in the reference gene sequence, including: The computer device extracts multiple different detection gene fragments from the reference sequence fragment and detects the first repeat number of each detection gene fragment in the reference sequence fragment.

[0006] In some possible embodiments, after determining one or more read sequence fragments in the first read sequence that do not match the reference gene sequence, the method further includes: The sequence fragments corresponding to the first read sequence are compared with the sequence fragments corresponding to the second read sequence to obtain the similarity between the sequence fragments of the first read sequence and the sequence fragments of the second read sequence, wherein the second read sequence is another gene sequence fragment read from the original gene data. If the similarity between the first read segment of the first read segment sequence and the second read segment of the second read segment sequence is greater than a similarity threshold, the computer device will classify the first read segment sequence and the second read segment sequence into a target read segment group. After the computer device determines that there is a cascaded repetition in the first read segment sequence, the method further includes: The computer device determines that each read segment sequence contained in the target read segment group has cascaded repetition.

[0007] In some possible embodiments, the computer device extracts multiple different detection gene fragments from the reference gene sequence, including: The computer device traverses the reference gene sequence and extracts multiple detection gene fragments of length K. The interval between the starting base positions of two adjacent detection gene fragments is N, where N is a positive integer and is less than or equal to K. K is a positive integer greater than or equal to M, and M is a preset minimum base length.

[0008] In some possible embodiments, the computer device traverses the reference gene sequence to extract multiple detection gene fragments of length K, including: The computer device uses M as the current K to traverse the reference gene sequence and extract multiple detection gene fragments with a base length of K. If the current K is less than or equal to the target length threshold, then the current K is updated to K+1, and the step of traversing the reference gene sequence and extracting multiple detection gene fragments of length K is re-executed until the current K is greater than the target length threshold.

[0009] In some possible embodiments, before the computer device acquires the first read sequence to be analyzed and the reference gene sequence, the method further includes: The computer device acquires multiple initial read sequences corresponding to the original gene data; The computer device preprocesses the multiple initial read segments to obtain multiple target read segments that meet the first quality condition; The computer device aligns the plurality of target read sequences with the reference genome. If the alignment results determine that the plurality of target read sequences meet the second quality condition, then the step of comparing the first read sequence with the reference genome sequence is executed; the first read sequence is any one of the plurality of target read sequences. The first quality condition includes one or more of the following conditions: The read sequence does not contain a connector sequence fragment; The read sequence does not contain polyadenylation or polyguanylate signals; The proportion of bases in the read sequence whose base information cannot be determined is lower than the first proportion threshold. The proportion of bases in the read sequence whose base quality value is greater than or equal to the target base quality value is higher than the second proportion threshold. The proportions of guanine G and cytosine C bases in the total number of bases in the read sequence are within the target range; The second quality condition includes one or more of the following conditions: Among the multiple target read sequences, the proportion of read sequences that are aligned to the reference genome is higher than the third proportion threshold. Among the multiple target read sequences, the proportion of read sequences aligned to the exon regions of the reference genome is higher than the fourth proportion threshold. Among the multiple target read sequences, the proportion of read sequences that are aligned to the ribosomal region of the reference genome is lower than the fifth proportion threshold.

[0010] A second aspect of this application discloses a sequence analysis apparatus for tandem repeats, the apparatus comprising: The sequence acquisition module is used to acquire the first read sequence to be analyzed and the reference gene sequence. The first read sequence is a gene sequence fragment read from the original gene data, and the reference gene sequence is the gene sequence of the first exon region on the target chromosome. The sequence comparison module is used to compare the first read sequence with the reference gene sequence to determine one or more read sequence fragments in the first read sequence that do not match the reference gene sequence. The fragment alignment module is used to extract multiple different detection gene fragments from the reference gene sequence, and to detect the first repeat number of each detection gene fragment in the reference gene sequence, and to detect the second repeat number of each detection gene fragment in the target read sequence fragment, wherein the target read sequence fragment is any one of the one or more read sequence fragments; The fragment detection module is used to determine that there is a tandem repeat in the first read sequence if there is a second repeat number corresponding to any detection gene fragment that is not less than the number threshold and the corresponding second repeat number is greater than the corresponding first repeat number.

[0011] A third aspect of this application discloses an electronic device, including a memory and a processor, wherein the memory stores a computer program, and when the computer program is executed by the processor, the processor causes the processor to implement the method described in any of the above embodiments.

[0012] A fourth aspect of this application discloses a computer program product comprising a computer program that, when executed by a processor, causes the processor to perform the method described in any of the above embodiments.

[0013] In this embodiment, a computer device acquires a first read sequence to be analyzed and a reference gene sequence, compares the first read sequence and the reference gene sequence, identifies one or more read sequence fragments in the first read sequence that do not match the reference gene sequence, and extracts multiple different detection gene fragments from the reference gene sequence. The computer device detects the first repetition count of each detection gene fragment in the reference gene sequence and the second repetition count of each detection gene fragment in the target read sequence fragment. The first repetition count reflects the normal inherent repetition level of each detection gene fragment in the reference gene sequence, while the second repetition count reflects the actual repetition level of each detection gene fragment in the target read sequence fragment. Therefore, if the second repetition count corresponding to any detection gene fragment is not less than a threshold and the corresponding second repetition count is greater than the corresponding first repetition count, the computer device can determine that the detection gene fragment has undergone more repetitions in the target read sequence fragment. This accurately determines the presence of tandem repetitions in the first read sequence, improving the sensitivity and accuracy of tandem repetition identification. Furthermore, since the detection is performed separately based on multiple different detection gene fragments extracted from the reference gene sequence, it can more comprehensively capture repetitive fragments in the target read sequence fragment, effectively reducing the risk of missed detection of tandem repetitions. Attached Figure Description

[0014] To more clearly illustrate the technical solutions in the embodiments of this application, the drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0015] Figure 1 This is a schematic diagram illustrating cascaded repetition in a read segment sequence provided in an embodiment of this application. Figure 2 A flowchart illustrating a method for analyzing tandem repeat sequences provided in this application embodiment; Figure 3 A flowchart of data preprocessing provided for embodiments of this application; Figure 4 This is a schematic diagram of detecting tandem repeats provided in an embodiment of this application; Figure 5 A flowchart for determining read sequence fragments based on differential bases provided in this application embodiment; Figure 6 A schematic diagram illustrating the determination of read sequence fragments based on differential bases, provided for embodiments of this application; Figure 7 A flowchart for extracting and detecting gene fragments from a reference sequence fragment is provided as an embodiment of this application; Figure 8 A schematic diagram illustrating the determination of a reference sequence segment corresponding to a target read segment, provided in an embodiment of this application; Figure 9 A flowchart for performing cascaded repetition analysis on read groups provided in this application embodiment; Figure 10 This is a schematic diagram illustrating the division of multiple read segment sequences into read segment groups, as provided in the embodiments of this application. Figure 11 A flowchart for extracting multiple different detection gene fragments provided in this application embodiment; Figure 12 A schematic diagram illustrating the extraction of multiple different detection gene fragments provided in an embodiment of this application; Figure 13 A schematic diagram illustrating a sequence analysis of tandem repeats provided in an embodiment of this application; Figure 14 A structural block diagram of a tandem repeat sequence analysis device provided in an embodiment of this application; Figure 15 This is a structural block diagram of an electronic device provided in an embodiment of this application. Detailed Implementation

[0016] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0017] It should be noted that the terms "comprising" and "having," and any variations thereof, in the embodiments and accompanying drawings of this application are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that includes a series of steps or units is not limited to the steps or units listed, but may optionally include steps or units not listed, or may optionally include other steps or units inherent to these processes, methods, products, or devices.

[0018] The sequence analysis method for tandem repeats provided in this application can be applied to computer devices, including but not limited to personal computers, tablets, laptops, and cloud servers.

[0019] Figure 1 This is a schematic diagram illustrating cascaded repetition in a read segment sequence provided in an embodiment of this application. For example... Figure 1As shown, read sequence 100 contains a short sequence 110 "AGGCCG" consisting of multiple bases. This short sequence 110 is repeated multiple times in read sequence 100 in a continuous, uninterrupted, and end-to-end manner. A computer device can determine that the gene to which read sequence 100 belongs has tandem repeats if tandem repeats are found in read sequence 100.

[0020] Read sequence 100 can refer to the base sequence obtained after sequencing a deoxyribonucleic acid sample or a ribonucleic acid sample.

[0021] In some embodiments, high-throughput sequencing technology can be used to sequence deoxyribonucleic acid (DNA) or ribonucleic acid (RNA) samples. High-throughput sequencing technology refers to the technology that simultaneously sequences a large number of sequences in a DNA or ribonucleic acid sample. High-throughput sequencing technologies may include sequencing-by-synthesis, semiconductor sequencing, and single-molecule real-time sequencing. The read sequence 100 can be the basic data unit produced by high-throughput sequencing.

[0022] In some embodiments, the gene to which read sequence 100 belongs may include the FMS-associated receptor tyrosine kinase 3 (FLT3) gene or the upstream binding transcription factor (UBTF) gene.

[0023] It should be noted that tandem repeats in the FLT3 gene, primarily manifested as internal tandem repeats (FLT3-ITD), are one of the most common driver mutations in acute myeloid leukemia (AML). Clinically, AML patients carrying FLT3-ITD often exhibit a high rate of minimal residual disease and an increased risk of relapse. Meanwhile, tandem repeat mutations in the UBTF gene (UBTF-TD) are novel genetic abnormalities discovered in recent years in AML and myelodysplastic syndromes, with a particularly high detection rate in childhood AML. Therefore, using computer equipment to detect the presence of tandem repeats in the corresponding read sequences of the FLT3 or UBTF genes can aid in targeted therapy and prognostic assessment of AML.

[0024] It should be noted that the sequence analysis results obtained by the computer device on the read sequence 100, which show tandem repeats, are only intermediate data and cannot be directly used for medical diagnosis. Medical professionals need to combine the sequence analysis results with other data for analysis and diagnosis.

[0025] In some embodiments, a computer device can detect whether there is a tandem repeat in the read sequence 100 by comparing the read sequence 100 with the reference gene sequence corresponding to the gene to which the read sequence 100 belongs.

[0026] The reference gene sequence may include the base sequence of the gene to which read sequence 100 in the human genome reference set belongs.

[0027] Optionally, if the gene to which read sequence 100 belongs is the UBTF gene, the reference gene sequence may include the base sequence corresponding to the region of exon 13 or exon 9 of chromosome 17 in the human genome reference set; if the gene to which read sequence 100 belongs is the FLT3 gene, the reference gene sequence may include the base sequence corresponding to the region of exon 14 or exon 15 of chromosome 13 in the human genome reference set.

[0028] In this embodiment, a computer device acquires a read sequence 100 to be analyzed and a reference gene sequence. The read sequence 100 and the reference gene sequence are compared to identify one or more read sequence fragments in the read sequence 100 that do not match the reference gene sequence. The computer device extracts multiple different detection gene fragments from the reference gene sequence, detects the first repeat count of each detection gene fragment in the reference gene sequence, and detects the second repeat count of each detection gene fragment in the target read sequence fragment. If the second repeat count corresponding to any detection gene fragment is not less than a threshold value and the corresponding second repeat count is greater than the corresponding first repeat count, the computer device can determine that the detection gene fragment has been repeated more times in the target read sequence fragment. This indicates the presence of tandem repeats in the read sequence 100, improving the sensitivity and accuracy of tandem repeat identification. Furthermore, since the detection is performed separately based on multiple different detection gene fragments extracted from the reference gene sequence, it can more comprehensively capture repeat fragments in the target read sequence fragment, effectively reducing the risk of missed detection of tandem repeats.

[0029] Figure 2 This is a flowchart illustrating a method for analyzing tandem repeat sequences provided in an embodiment of this application. Figure 2 As shown, the method may include the following steps: Step 202: The computer device acquires the first read sequence to be analyzed and the reference gene sequence.

[0030] The first read sequence is a gene sequence fragment read from the original gene data, and the reference gene sequence is the gene sequence of the first exon region on the target chromosome. Optionally, the reference gene sequence may be the gene sequence of the first exon region on the target chromosome in a reference genome, and the reference genome may be the human genome reference set, for example, the human genome reference set hg38.

[0031] Raw genetic data refers to the collection of gene sequences obtained by sequencing the nucleic acid molecules of biological samples such as blood, tissues, or cells using gene sequencing equipment.

[0032] The target chromosome can be the chromosome containing the gene that needs to be detected for tandem duplication. For example, if a computer device needs to detect whether the UBTF gene has tandem duplication, the target chromosome can be chromosome 17, where the UBTF gene is located; if a computer device needs to detect whether the FLT3 gene has tandem duplication, the target chromosome can be chromosome 13, where the FLT3 gene is located.

[0033] The first exon region can be any one of multiple exon regions on the target chromosome. When the target chromosome is chromosome 13, the first exon region can be any one of multiple exon regions on chromosome 13; when the target chromosome is chromosome 17, the first exon region can be any one of multiple exon regions on chromosome 17.

[0034] In some embodiments, a computer device may align a first read sequence with a reference genome to determine the reference gene sequence corresponding to the first read sequence.

[0035] Computer equipment can use STAR software to align the first read sequence into the reference genome, generate first alignment information, and then determine the reference gene sequence corresponding to the first read sequence in the reference genome based on the first alignment information.

[0036] For example, if a computer device uses STAR software to align the first read sequence to the region of exon 13 on chromosome 17 in the reference genome, the computer device can determine, based on the corresponding first alignment information, that the reference gene sequence corresponding to the first read sequence is the base sequence corresponding to the region of exon 13 on chromosome 17 in the reference genome.

[0037] In some embodiments, where the original gene data is the data directly output by the gene sequencing device after sequencing, the computer device can perform data preprocessing on multiple initial read sequences corresponding to the original gene data to obtain the first read sequence to be analyzed. Figure 3 A flowchart illustrating data preprocessing provided for embodiments of this application. For example... Figure 3 As shown, the method may further include the following steps: Step 301: The computer device acquires multiple initial read sequences corresponding to the original gene data.

[0038] Step 303: The computer device preprocesses multiple initial read segments to obtain multiple target read segments that meet the first quality condition.

[0039] The first quality condition includes one or more of the following conditions: Condition 1: The read sequence does not contain a connector sequence fragment.

[0040] If a read contains a linker sequence fragment, it indicates that the read carries a non-human-derived external sequence introduced during the sequencing process. This can interfere with subsequent alignment analysis with the human reference gene sequence, thus affecting the accuracy of the alignment results. Therefore, it is necessary to remove reads containing linker sequence fragments.

[0041] Condition 2: The read sequence does not contain polyadenylated or polyguanylic acid signals.

[0042] Polyadenylation signal, also known as polyA, refers to a base sequence fragment formed by the tandem of multiple adenine A atoms, while polyguanine signal, also known as polyG, refers to a base sequence fragment formed by the tandem of multiple guanine G atoms.

[0043] Understandably, if a read contains polyadenylated or polyguanylic acid (PGA) signals, it indicates that the read may originate from ribonucleic acid or a non-biological sequence introduced during sequencing, which will also affect the accuracy of subsequent alignment results. Therefore, it is necessary to remove reads containing PGA or PGA signals.

[0044] Condition 3: The proportion of bases in the read sequence whose base information cannot be determined is lower than the first proportion threshold.

[0045] Optionally, the first proportion threshold can be 5%. If the proportion of bases in a read sequence whose base information cannot be determined is higher than 5% of the total number of bases in the read sequence, the read sequence can be considered not to meet the first quality condition.

[0046] Condition 4: The proportion of bases in the read sequence whose base quality value is greater than or equal to the target base quality value to the total number of bases in the read sequence is higher than the second proportion threshold.

[0047] The target base quality value may include one or more of 10, 20, and 30. The second proportion threshold may include the proportion threshold corresponding to the target base quality value.

[0048] For example, when the target base mass value is 10, the second proportion threshold can be 80%; when the target base mass value is 20, the second proportion threshold can be 90%; and when the target base mass value is 30, the second proportion threshold can be 80%.

[0049] Condition 5: The proportion of guanine G and cytosine C bases in the total number of bases in the read sequence is within the target range.

[0050] For example, the target range can be 40% to 60%. When the proportion of guanine G and cytosine C bases in the total number of bases in the read sequence is within 40% to 60%, the read sequence can be considered to meet the first quality condition.

[0051] Optionally, the first quality condition may also include: the effective sequencing quantity of the entire read sequence is greater than a preset minimum sequencing quantity. The preset minimum sequencing quantity can be set to 10G, or 10 billion base pairs.

[0052] Step 305: The computer device aligns multiple target read sequences with the reference genome. If the alignment results indicate that multiple target read sequences meet the second quality condition, then the step of comparing the first read sequence with the reference genome sequence is executed.

[0053] The first read sequence is any one of multiple target read sequences.

[0054] The second quality condition includes one or more of the following conditions: Condition 1: Among multiple target read sequences, the proportion of read sequences aligned to the reference genome is higher than the third proportion threshold.

[0055] If the proportion of reads that align to the reference genome in multiple target read sequences is lower than the third proportion threshold, it indicates that there are many non-human-derived sequences in the multiple target read sequences, and they are not of alignment value.

[0056] Optionally, the third proportion threshold can be 80%. If the proportion of reads aligned to the reference genome in multiple target read sequences is higher than 80%, then the multiple target read sequences are determined to meet the second quality condition.

[0057] Condition 2: Among multiple target read sequences, the proportion of read sequences aligned to exon regions in the reference genome is higher than the fourth proportion threshold.

[0058] Since tandem repeats in genes to be detected, such as UBTF or FLT3, typically occur in exon regions, a high sequencing coverage of exon regions can be ensured when the proportion of reads aligned to exon regions in the reference genome is higher than the fourth proportion threshold, thereby improving the accuracy of tandem repeat detection in the genes to be detected.

[0059] Optionally, the fourth proportion threshold can be 40%. If the proportion of reads aligned to exon regions in the reference genome among multiple target read sequences is higher than 40%, then the multiple target read sequences are determined to meet the second quality condition.

[0060] Condition 3: Among multiple target read sequences, the proportion of read sequences aligned to the ribosomal region in the reference genome is lower than the fifth proportion threshold.

[0061] Optionally, the fifth proportion threshold can be 10%. If the proportion of reads aligned to ribosomal regions in the reference genome among multiple target read sequences is less than 10%, then the multiple target read sequences are determined to meet the second quality condition.

[0062] By using the above method, multiple initial read sequences in the original gene data are preprocessed before comparing the first read sequence with the reference gene sequence to obtain target read sequences that meet the first and second quality conditions. This not only improves the matching degree between the read sequence and the reference genome and ensures the high quality of the first read sequence, but also helps to improve the efficiency of the tandem repeat sequence analysis process.

[0063] Step 204: The computer device compares the first read sequence with the reference gene sequence to identify one or more read sequence fragments in the first read sequence that do not match the reference gene sequence.

[0064] A read sequence fragment may include multiple bases in the first read sequence that do not match the reference gene sequence.

[0065] In some embodiments, the computer device may compare the first read sequence with a reference gene sequence, establish a mapping relationship between the positions of each base in the reference gene sequence and the positions of each base in the first read sequence, identify multiple bases in the first read sequence that do not match the reference gene sequence, and determine one or more read sequence fragments based on the multiple bases.

[0066] Computer equipment can use STAR software to establish a mapping relationship between the positions of each base in the reference gene sequence and the positions of each base in the first read sequence, so as to identify multiple bases in the first read sequence that do not match the reference gene sequence, and determine one or more read sequence fragments based on multiple bases.

[0067] Step 206: The computer device extracts multiple different detection gene fragments from the reference gene sequence and detects the first repeat number of each detection gene fragment in the reference gene sequence and the second repeat number of each detection gene fragment in the target read sequence fragment.

[0068] The target read sequence segment is any one of one or more read sequence segments.

[0069] Understandably, the first number of repetitions reflects the normal inherent repetition level of the gene fragment being detected in the reference gene sequence, while the second number of repetitions reflects the actual repetition level of the gene fragment being detected in the target read sequence. By comparing the first number of repetitions and the second number of repetitions, it can be determined whether the gene fragment being detected is tandemly repeated.

[0070] In some embodiments, a computer device may extract multiple detection gene fragments with the same base length but different constituent bases from a reference gene sequence, and / or multiple detection gene fragments with the same core base sequence but different base lengths.

[0071] Step 208: If there exists a second repeat number corresponding to any detected gene fragment that is not less than the number threshold, and the corresponding second repeat number is greater than the corresponding first repeat number, the computer device determines that there is a tandem repeat in the first read sequence. If any detected gene fragment has a second repeat number not less than the number threshold, and the corresponding second repeat number is greater than the corresponding first repeat number, the computer device can determine that the detected gene fragment is a tandem repeat sequence in the target read sequence fragment, and that tandem repeats exist in the first read sequence.

[0072] The number of attempts threshold can be set according to actual needs; for example, the number of attempts threshold can be set to two or three.

[0073] Optionally, the number of times threshold can be set according to the base length corresponding to the gene fragment being detected. For example, if the base length corresponding to the gene fragment being detected is short, the number of times threshold can be set to three times, and if the base length corresponding to the gene fragment being detected is long, the number of times threshold can be set to four times.

[0074] It is understandable that if the second repetition number is greater than the corresponding first repetition number, it can be determined that the detected gene fragment is repeated multiple times in the target read sequence fragment. The specific number of repetitions can be determined based on the difference between the second repetition number and the first repetition number. If the second repetition number is not less than a preset threshold number, it can be determined that the repetition of the detected gene fragment is a tandem repetition, rather than a repetition of the detected gene fragment caused by mutations such as insertion.

[0075] For example, Figure 4 This is a schematic diagram illustrating the detection of tandem repeats provided in an embodiment of this application. Figure 4As shown, after acquiring the target read sequence 401 and the reference gene sequence 402, the computer device compares the target read sequence 401 and the reference gene sequence 402 to identify a read sequence fragment 403 in the target read sequence 401 that does not match the reference gene sequence 402. The start base position of this read sequence fragment 403 is 13 bp, and the end base position is 42 bp. The computer device can extract the target detection gene fragment 404 from the reference gene sequence 402 and detect that the first repeat of the target detection gene fragment 404 in the reference gene sequence 402 is once, and the second repeat of the target detection gene fragment 404 in the read sequence fragment 403 is twice. Therefore, the computer device determines that the second repeat of the target detection gene fragment 404 is not less than the threshold of two, and the corresponding second repeat is greater than the corresponding first repeat. The computer device can then determine that tandem repeats exist in the target read sequence 401. Furthermore, the computer equipment can identify the longest base-length gene fragment that meets the above conditions as the tandem repeat sequence in the target read sequence. For example... Figure 4 As shown, the second repeat number is the same for both the detection gene fragment "AGGCCGC" and the detection gene fragment "AGGCCGCTCCTCCT". Therefore, the computer device can determine that the longer detection gene fragment "AGGCCGCTCCTCCT" is a tandem repeat sequence in the target read sequence fragment.

[0076] In some embodiments, when there are multiple read sequence fragments in the first read sequence, the computer device may, after determining that the detection gene fragment is a tandem repeat sequence in the target read sequence fragment, detect the second repeat number of each detection gene fragment in the next read sequence fragment in the first read sequence based on the multiple different detection gene fragments extracted, so as to determine whether tandem repeats exist in each read sequence fragment of the first read sequence.

[0077] By detecting whether tandem repeats exist in each read segment of the first read sequence, the mutation frequency of the first read sequence can be calculated by statistically analyzing the proportion of read segments with tandem repeats to the total number of read segments in the first read sequence. This provides richer and more accurate data for subsequent analysis.

[0078] In some embodiments, if a reference gene sequence is not determined based on the first read sequence, the computer device may, after determining that there is no tandem repeat in the first read sequence, set the gene sequence of the second exon region on the target chromosome as the new reference gene sequence and re-execute step 204.

[0079] For example, tandem repeats in the UBTF gene typically occur in exon 13 or exon 9 of chromosome 17. Therefore, if a computer device uses the base sequence corresponding to exon 13 as a reference gene sequence and compares the first read sequence with the reference gene sequence to determine that there is no tandem repeat sequence in the first read sequence, the computer device can use the base sequence corresponding to exon 9 as a new reference gene sequence and compare the first read sequence with the new reference gene sequence again to determine whether there is a tandem repeat sequence in the first read sequence.

[0080] In this embodiment, a computer device acquires a first read sequence to be analyzed and a reference gene sequence, compares the first read sequence and the reference gene sequence, identifies one or more read sequence fragments in the first read sequence that do not match the reference gene sequence, and extracts multiple different detection gene fragments from the reference gene sequence. The computer device detects the first repetition count of each detection gene fragment in the reference gene sequence and the second repetition count of each detection gene fragment in the target read sequence fragment. The first repetition count reflects the normal inherent repetition level of each detection gene fragment in the reference gene sequence, while the second repetition count reflects the actual repetition level of each detection gene fragment in the target read sequence fragment. Therefore, if the second repetition count corresponding to any detection gene fragment is not less than a threshold and the corresponding second repetition count is greater than the corresponding first repetition count, the computer device can determine that the detection gene fragment has undergone more repetitions in the target read sequence fragment. This accurately determines the presence of tandem repetitions in the first read sequence, improving the sensitivity and accuracy of tandem repetition identification. Furthermore, since the detection is performed separately based on multiple different detection gene fragments extracted from the reference gene sequence, it can more comprehensively capture repetitive fragments in the target read sequence fragment, effectively reducing the risk of missed detection of tandem repetitions.

[0081] In some embodiments, such as Figure 5 As shown, the step of comparing the first read sequence with the reference gene sequence to identify one or more read sequence fragments in the first read sequence that do not match the reference gene sequence may include the following steps: Step 501: The computer device compares the first read sequence with the reference gene sequence to identify the differential bases in the first read sequence that do not match the reference gene sequence.

[0082] Computer equipment can use STAR software to compare the first read sequence with the reference gene sequence to identify the differentially expressed bases in the first read sequence that do not match the reference gene sequence.

[0083] Step 503: The computer device connects multiple differential bases with consecutive base positions or multiple differential bases with intervals between base positions not greater than a first interval threshold to obtain one or more read sequence fragments.

[0084] The first interval threshold can be set according to actual needs. For example, the first interval threshold can be set to 3 bp. If the interval between the base positions of multiple different bases is no greater than 3 bp, multiple different bases can be connected to obtain one or more read sequence fragments.

[0085] It is understandable that there may be multiple matching bases between two differing bases. However, these multiple matching bases may not be normal bases, but rather mutations caused by base identification errors during sequencing or other reasons. Furthermore, the smaller the interval between the base positions of two differing bases, the greater the probability that the multiple matching bases within the interval are mutations. If these differing bases separated by only a few matching bases are divided into corresponding read sequences, it will not only increase the number of read sequences and reduce sequence analysis efficiency, but also cause a long tandem repeat sequence to be split into multiple shorter repeat sequences, failing to accurately reflect the tandem repeat of the first read sequence. Therefore, it is necessary to connect multiple differing bases with intervals no greater than a first interval threshold to form one or more read sequences.

[0086] In one implementation, a computer device can first connect multiple different bases with consecutive base positions to form multiple short segments, and then connect multiple short segments with an interval between base positions not greater than a first interval threshold to obtain one or more read sequence segments.

[0087] Figure 6 This is a schematic diagram illustrating the determination of read sequence fragments based on differential bases, provided as an embodiment of this application. For example... Figure 6 As shown, the first interval threshold can be set to 3 bp. The computer device compares the first read sequence 600 with the reference gene sequence to determine the differential bases in the first read sequence that do not match the reference gene sequence. Figure 6 (The black portion in the middle) and the continuous differential bases are connected to form the first short fragment 601, the second short fragment 602, the third short fragment 603 and the fourth short fragment 604. The interval between the first short fragment 601 and the second short fragment 602 is 2 bp, the interval between the second short fragment 602 and the third short fragment 603 is 1 bp, and the interval between the third short fragment 603 and the fourth short fragment 604 is 10 bp. The computer device can connect the first short fragment 601, the second short fragment 602 and the third short fragment 603 with an interval of no more than 3 bp to form the read sequence 605. The fourth short fragment 604 is used as another read sequence and is not connected.

[0088] In this embodiment, the computer device connects multiple differential bases with consecutive base positions or multiple differential bases with intervals no greater than a first interval threshold to obtain one or more read sequence fragments. This not only effectively filters base matching caused by sequencing errors and avoids splitting multiple differential bases in the same mutation region into different read sequences, but also avoids fragmentation of tandem repeat sequences, improving the efficiency and accuracy of sequence analysis. Furthermore, it helps to accurately reconstruct the true mismatched read sequences of the target read sequence, improving the accuracy of tandem repeat detection of read sequences.

[0089] In some embodiments, since the target read sequence usually corresponds to only a part of the reference gene sequence rather than the entire reference gene sequence, it is necessary to extract the detection gene fragment from a part of the corresponding reference gene sequence to avoid extracting too many detection gene fragments from the reference gene sequence, which would lead to low data processing efficiency.

[0090] Figure 7 This is a flowchart illustrating the extraction and detection of gene fragments from a reference sequence fragment, provided as an embodiment of this application. Figure 7 As shown, after obtaining one or more read segment sequences, the method further includes the following steps: Step 701: The computer device determines the reference sequence fragment corresponding to the target read sequence fragment in the reference gene sequence based on the start and end base positions corresponding to the target read sequence fragment.

[0091] The start base position of the reference sequence fragment is less than or equal to the start base position of the target read sequence fragment, and the end base position of the reference sequence fragment is greater than or equal to the end base position of the target read sequence fragment.

[0092] By employing the above method, the reference sequence fragment can not only cover the target read sequence fragment, but also some normal base regions upstream and downstream of the target read sequence fragment. This allows for the extraction of sufficient detection gene fragments, ensuring the comprehensiveness and accuracy of subsequent sequence analysis and guaranteeing that the reliability of the final sequence analysis results is not affected by the limitation of the sequence fragment range.

[0093] In one implementation, a computer device can determine a matching sequence fragment with the same start and end base positions in a reference gene sequence based on the start and end base positions corresponding to the target read sequence fragment. Then, based on the matching sequence fragment and flanking sequence fragments with a preset base length on both sides of the matching sequence fragment, the reference sequence fragment corresponding to the target read sequence fragment can be determined.

[0094] The base lengths of the flanking sequence fragments on both sides can be set according to actual needs. For example, the base lengths of the flanking sequence fragments on both sides can be set to 50 bp.

[0095] Optionally, the base lengths of the flanking sequence fragments on both sides of the matching sequence fragment can also be different. For example, the base length of the flanking sequence fragment on the left side of the matching sequence fragment can be 50 bp, and the base length of the flanking sequence fragment on the right side of the matching sequence fragment can be 20 bp.

[0096] Optionally, since tandem repeat sequences are usually repeated in a continuous, uninterrupted, and end-to-end manner, that is, the repeat unit can be detected in the sequence fragment before the start base position corresponding to the target read sequence fragment, the computer device can use the sequence fragment between the start base position of the reference gene sequence and the end base position corresponding to the target read sequence fragment as the reference sequence fragment.

[0097] Figure 8 This is a schematic diagram illustrating the determination of a reference sequence segment corresponding to the target read segment, as provided in an embodiment of this application. Figure 8 As shown, the computer device can determine a matching sequence fragment 803 with the same start and end base positions in the reference gene sequence 800 based on the start and end base positions corresponding to the target read sequence fragment 802 of the first read sequence 801. Then, based on the matching sequence fragment 803, the flanking sequence fragment with a length of 50 bp to the left of the matching sequence fragment 803, and the flanking sequence fragment with a length of 20 bp to the right of the matching sequence fragment 803, the reference sequence fragment 804 corresponding to the target read sequence fragment 802 is determined.

[0098] Step 703: The computer device extracts multiple different detection gene fragments from the reference sequence fragment and detects the first repeat number of each detection gene fragment in the reference sequence fragment and the second repeat number of each detection gene fragment in the target read sequence fragment.

[0099] In this embodiment, the computer device determines the reference sequence fragment corresponding to the target read sequence fragment in the reference gene sequence based on the start and end base positions of the target read sequence fragment, and extracts multiple different detection gene fragments from the reference sequence fragment. The first repetition number of each detection gene fragment in the reference sequence fragment is detected, which can effectively reduce the number of detection gene fragments that are not related to the target read sequence fragment, thereby reducing the computational load of sequence alignment analysis and improving the efficiency of sequence analysis.

[0100] In some embodiments, since a single sequencing operation typically generates multiple target read sequences, and these multiple target read sequences may all have the same or similar read sequence fragments, when comparing multiple target read sequences with a reference gene sequence, multiple target read sequences with the same or similar read sequence fragments can be classified into the same read group. If the computer device determines that a target read sequence has tandem repeats, it can be determined that each read sequence in the same read group has tandem repeats.

[0101] Figure 9 This is a flowchart illustrating the cascade repetition analysis of read segments provided in an embodiment of this application. For example... Figure 9 As shown, after identifying one or more read sequences in the first read sequence that do not match the reference gene sequence, the method further includes the following steps: Step 901: The computer device compares each segment of the first read sequence with each segment of the second read sequence to obtain the similarity between each segment of the first read sequence and each segment of the second read sequence.

[0102] The second read sequence is another gene sequence fragment read from the original gene data.

[0103] The second read sequence segment can be a read sequence segment whose start and end base positions are the same as those of the first read sequence segment of the first read sequence.

[0104] In one implementation, a computer device can determine the similarity between the first read sequence fragment of the first read sequence and the second read sequence fragment of the second read sequence by counting the number of matching bases between the first read sequence fragment of the first read sequence and the second read sequence fragment of the second read sequence.

[0105] The similarity threshold can be set according to actual needs.

[0106] As another implementation, a computer device can determine the similarity between the first read sequence fragment of the first read sequence and the second read sequence fragment of the second read sequence by counting the number of different bases between the first read sequence fragment of the first read sequence and the second read sequence fragment of the second read sequence. The fewer the number of different bases, the higher the similarity.

[0107] Step 903: If the similarity between the first read segment of the first read sequence and the second read segment of the second read sequence is greater than the similarity threshold, the computer device classifies the first read sequence and the second read sequence into the target read group.

[0108] Each read sequence contained in the target read group has read sequence segments with the same start base position and end base position.

[0109] It should be noted that bases in the tandem repeat sequences of UBTF-TD are prone to mutation. For example, the third base G in the tandem repeat sequence "AGGCCGCTCCTCCTG" may mutate to C, forming "AGCCCGCTCCTCCTG". This causes the original tandem repeat sequence to be split into multiple shorter repeat sequences, making it difficult to accurately determine the repetition count of the gene fragment in the corresponding read sequence of the UBTF gene. Therefore, when the similarity between the first read sequence and the second read sequence exceeds the similarity threshold, the first and second read sequences need to be grouped into the target read group. This allows for sequence analysis of read sequences within the same group, which not only avoids failing to detect mutated tandem repeat sequences in the UBTF gene but also helps reduce the number of times the same read sequence is analyzed, further improving the efficiency of sequence analysis.

[0110] In some embodiments, a computer device may count the number of different bases between the first read sequence fragment of the first read sequence and the second read sequence fragment of the second read sequence, and if the number of different bases between the first read sequence fragment of the first read sequence and the second read sequence fragment of the second read sequence is no more than 3, determine that the similarity between the first read sequence fragment of the first read sequence and the second read sequence fragment of the second read sequence is greater than a similarity threshold, thereby classifying the first read sequence and the second read sequence into the target read group.

[0111] Figure 10 This is a schematic diagram illustrating the division of multiple read segment sequences into read segment groups, as provided in the embodiments of this application. Figure 10 As shown, the computer device can acquire four read sequences: read sequence a, read sequence b, read sequence c, and read sequence d. After comparing the four read sequences with the reference gene sequence, the computer device determines that read sequences a and c both contain read sequence fragment 1 (referred to as fragment 1) and read sequence fragment 2 (referred to as fragment 2), read sequence b contains fragment 1, and read sequence d contains fragment 2. The computer device can perform similarity detection on read sequences a, b, and c containing fragment 1. If the similarity is greater than the similarity threshold, read sequences a, b, and c are classified into the first read group. The computer device can perform similarity detection on read sequences a, c, and d containing fragment 2. If the similarity is greater than the similarity threshold, read sequences a, c, and d are classified into the second read group.

[0112] In other embodiments, steps 901 to 903 may also be performed simultaneously when the computer device compares the first read sequence with the reference gene sequence, or when the computer device aligns the first read sequence with the reference genome, but are not limited thereto.

[0113] Step 905: If the computer device determines that there is serial repetition in the first read segment sequence, it determines that all read segment sequences contained in the target read segment group have serial repetition.

[0114] It is understandable that if there is tandem repetition in any read sequence contained in the target read group, then all read sequence segments corresponding to all read sequences contained in the target read group will have tandem repetition.

[0115] In some embodiments, when the computer device determines that there is no tandem repeat in the first read sequence, it detects the third repeat number of each detection gene fragment in the second read sequence of the second read sequence based on multiple different detection gene fragments extracted from the reference gene sequence, in order to determine whether there is a tandem repeat in the second read sequence, until a read sequence included in the target read group is determined to have a tandem repeat, or the detection of each read sequence included in the target read group is completed. In some embodiments, the computer device may perform sequence analysis on each read sequence contained in the target read group, and if it is determined that the target detection gene sequence corresponding to the first read sequence is the longest tandem repeat sequence in the target read group, it may determine that the tandem repeat sequences corresponding to each read sequence contained in the target read group are all target detection gene sequences.

[0116] The target gene sequence is the gene fragment in the first read sequence that satisfies the condition that the number of the corresponding second repeat is not less than the number value, and the number of the corresponding second repeat is greater than the number of the corresponding first repeat.

[0117] Understandably, because the bases in the tandem repeat sequences of the UBTF gene are prone to mutation, the tandem repeat sequences may not be identified or may be split into multiple shorter repeat sequences. Sequence alignment of gene fragments can only detect the number of repetitions of each shorter repeat sequence, but cannot detect the number of repetitions of the longer mutated tandem repeat sequences. That is, it can only determine that there are tandem repeats in the first read sequence corresponding to the UBTF gene, and that the longest repeat sequence among the multiple shorter repeat sequences is a tandem repeat sequence, but cannot determine the actual longer tandem repeat sequence.

[0118] Therefore, if the target gene sequence corresponding to the first read sequence is determined to be the longest tandem repeat sequence in the target read group, it can be inferred that the other shorter repeat sequences in the target read group are all broken fragments caused by base mutations. Thus, the tandem repeat sequences corresponding to each read sequence in the target read group are uniformly identified as target gene sequences, thereby eliminating the identification bias of tandem repeat sequences caused by base mutations.

[0119] Furthermore, after determining the tandem repeat sequence of the target read group, the computer device can compare the tandem repeat sequences of each read sequence to identify the mutated bases in the tandem repeat sequence corresponding to each read sequence.

[0120] In this embodiment, the computer device classifies the first read sequence and the second read sequence into a target read group when the similarity between the first read sequence fragment of the first read sequence and the second read sequence fragment of the second read sequence is greater than a similarity threshold. This not only effectively reduces the number of times the same read sequence fragment is analyzed, further improving the efficiency of sequence analysis, but also avoids the problem that mutations in one or more bases in the read sequence fragment can lead to the inability to accurately obtain the second repeat number of the detection gene fragment in the read sequence fragment.

[0121] In some embodiments, the computer device extracts multiple different detection gene fragments from a reference gene sequence, including: the computer device traverses the reference gene sequence to extract multiple detection gene fragments with a base length of K, and the interval between the starting base positions of two adjacent detection gene fragments is N, where N is a positive integer and N is less than or equal to K, K is a positive integer greater than or equal to M, and M is a preset minimum base length.

[0122] It should be noted that the computer device traverses the reference gene sequence and extracts multiple detection gene fragments with a base length of K. In essence, it extracts multiple different detection gene fragments from the reference gene sequence through a sliding window with a window length of K and a sliding step size of N, thereby achieving high-density coverage of the reference gene sequence.

[0123] As one implementation method, the minimum base length M can be set according to the gene to which the reference gene sequence belongs. For example, if the reference gene sequence belongs to the UBTF gene, M can be 6; if the reference gene sequence belongs to the FLT3 gene, M can be 12.

[0124] Understandably, since the base length of the tandem repeat sequence of the FLT3 gene is usually between 12 and 500 bp, when the reference gene sequence belongs to the FLT3 gene, setting M to 12 can ensure that the extracted detection gene fragment can completely cover the shortest natural tandem repeat sequence in the FLT3 gene. Since the tandem repeat sequence of the UBTF gene is prone to mutation, causing a long tandem repeat sequence to be split into multiple shorter repeat sequences, when the reference gene sequence belongs to the UBTF gene, setting M to 6 can ensure that even if the tandem repeat sequence is broken due to base mutation, the resulting short fragments can still be included in the detection range, and these shorter repeat sequences derived from mutations will not be filtered out due to the minimum base length limitation.

[0125] Figure 11 This is a flowchart illustrating the extraction of multiple different detection gene fragments provided in an embodiment of this application. For example... Figure 11 As shown, the computer device traverses the reference gene sequence and extracts multiple detection gene fragments of length K, which may include the following steps: Step 1101: The computer device uses M as the current K to traverse the reference gene sequence and extract multiple detection gene fragments with a base length of K.

[0126] Step 1103: Determine whether the current K is greater than the target length threshold. If not, proceed to step 1105; if yes, proceed to step 1107.

[0127] The target length threshold can be set according to actual needs. For example, the target length threshold can be set to half the base length of the reference gene sequence.

[0128] Furthermore, when multiple detection gene fragments are extracted based on a reference sequence fragment, the target length threshold can be set to half the base length of the reference sequence fragment.

[0129] Step 1105: Update the current K to K+1, and re-execute the step of traversing the reference gene sequence and extracting multiple detection gene fragments with a length of K.

[0130] Step 1107: Extraction of multiple detection gene fragments.

[0131] For example, such as Figure 12 As shown, taking the reference gene sequence as the base sequence corresponding to the region of exon 13 on chromosome 17, and the gene to which it belongs is the UBTF gene, the reference gene sequence is as follows: CTTCTTCTTCTCAGACAGGTCGTTCCACATTCGGGCCAGCAGGCGGGTCAGCTCGCTCTCGGAGAGCTCAGGCCGCTCCTCCTGCAGCTGCCGCCGTTTCTCCTCCGAGAAGATGAACATGGCCGACACGGGCCGCTTGGGCTTCTCGGAGCCGCC The computer device can determine that M is 6 and N is 1. The target length threshold corresponding to the reference gene sequence is half the base length of the reference gene sequence, that is, 78. Therefore, the computer device can first set the base length K to 6, traverse the reference gene sequence, and extract multiple detection gene fragments with a base length of 6. The interval between the starting base positions of two adjacent detection gene fragments is 1. For example, the first detection gene fragment with a base length of 6 can be "CTTCTT", the second detection gene fragment can be "TTCTTC", the third detection gene fragment can be "TTCTTC", until the last detection gene fragment with a base length of 6 is "GCCGCC".

[0132] The computer equipment then updates K sequentially to 7, 8, and up to 78. After extracting the last detection gene fragment with a base length of 78, "CTCCTGCAGCTGCCGCCGTTTCTCCTCCGAGAAGATGAACATGGCCGACACGGGCCGCTTGGGCTTCTCGGAGCCGCC", it is determined that the updated K+1 is 79, which is greater than the target length threshold of 78, thus completing the extraction of the reference gene sequence.

[0133] In this way, the computer equipment traverses the reference gene sequence and extracts multiple detection gene fragments of different base lengths, which helps to identify the longest repetitive sequence in the target read sequence fragment and improves the accuracy of tandem repeat sequence analysis.

[0134] In some embodiments, after extracting multiple different detection gene fragments from a reference gene sequence, the computer device can store the reference gene sequence and the corresponding multiple different detection gene fragments in a database.

[0135] In some embodiments, the computer device may pre-store in a database multiple different detection gene fragments contained in reference gene sequences corresponding to multiple genes. When the reference gene sequence corresponding to the first read sequence is determined, and the corresponding gene is determined, the computer device may obtain multiple different detection gene fragments of the reference gene sequence corresponding to the first read sequence from the database, so as to avoid the problem of low sequence analysis efficiency caused by extracting multiple different detection gene fragments from the reference gene sequence in real time.

[0136] In this embodiment, the computer device traverses the reference gene sequence and extracts multiple detection gene fragments of length K. The interval between the starting base positions of two adjacent detection gene fragments is N, so that the extracted detection gene fragments can achieve high-density coverage of the reference gene sequence. This helps to accurately identify repetitive sequences of different lengths in each read sequence fragment, and can effectively avoid the problem that the tandem repeat sequence is broken into shorter repeat sequences due to base mutations and cannot be identified. This improves the sensitivity and accuracy of tandem repeat sequence detection.

[0137] Figure 13 This is a schematic diagram illustrating a sequence analysis of tandem repeats provided in an embodiment of this application. Figure 13 As shown, the computer device can obtain reference gene sequence 1300, first read sequence 1301, second read sequence 1302 and third read sequence 1303, wherein the gene to which reference gene sequence 1300 belongs is the UBTF gene, and reference gene sequence 1300 is the gene sequence of the 13th exon region on chromosome 17.

[0138] The computer equipment can compare each read sequence with the reference gene sequence 1300 to determine that the read sequences that do not match the reference gene sequence 1300 are all base sequences with the same start and end base positions. Figure 13 (The gray part in the middle) By comparing the read sequence segments corresponding to each read sequence, it is determined that the similarity of the read sequence segments corresponding to each read sequence is higher than the similarity threshold, and the first read sequence 1301, the second read sequence 1302 and the third read sequence 1303 are assigned to the target read group.

[0139] The computer device can determine the reference sequence fragment 1304 corresponding to the target read sequence fragment in the reference gene sequence 1300 based on the start and end base positions of the read sequence fragment, and traverse the reference sequence fragment 1304 to extract multiple detection gene fragments with base lengths of 6, 7, ... 23 and 24, and detect the first repeat number of each detection gene fragment in the reference gene sequence 1304, as well as the second repeat number of each detection gene fragment in the read sequence fragment of the first read sequence 1301.

[0140] The computer equipment can determine that the first repeat of the detection gene fragment "AGGCCG" is once, the second repeat is twice, and the second repeat is not less than the number threshold. It can also determine that there is a tandem repeat in the first read sequence 1301, and that the detection gene fragment "AGGCCGCTCCTCCTG" has the longest base length under the above conditions (the base length of the detection gene fragment is half the length of the read sequence fragment, and the detection gene fragment is repeated twice). Thus, it can determine that the detection gene fragment is a tandem repeat sequence of the first read sequence 1301, and determine that the tandem repeat sequences of the second read sequence 1302 and the third read sequence 1303 in the target read group are both "AGGCCGCTCCTCCTG".

[0141] Furthermore, if the computer device first detects the second repeat count of each gene fragment in the second read sequence 1302, since the second base G in "AGGCCGCTCCTCCTG" is mutated to C, the computer device can only determine that the gene fragment "CCGCTCCTCCTG" is a repeat sequence in the second read sequence 1302, thus confirming the existence of tandem repeats in the second read sequence 1302; similarly, for the third read sequence 1303, the computer device can only determine that the gene fragment "AGGCCGCTC" is a repeat sequence in the third read sequence 1303. By comparing the base lengths of the repeat sequences in each read sequence, the computer device can determine that "AGGCCGCTCCTCCTG" in the first read sequence 1301 has the longest base length, thereby identifying the base sequence "AGGCCGCTCCTCCTG" as the tandem repeat sequence of the target read group.

[0142] In some embodiments, after determining that a first read sequence has a tandem repeat, the computer device may record data such as the tandem repeat sequence of the first read sequence, the start and end base positions of the tandem repeat sequence, the flanking sequences on both sides of the tandem repeat sequence, and the number of read sequences in the target read group that support the tandem repeat sequence, in order to calculate the allele frequency.

[0143] In some embodiments, the computer device may output a sequence analysis report of tandem repeats after determining whether tandem repeats exist in the first read sequence. If tandem repeats exist in the first read sequence, a positive report may be output, and if tandem repeats do not exist in the first read sequence, a negative report may be output. Medical professionals may review the positive or negative reports for further analysis.

[0144] Based on the tandem repeat sequence analysis method provided in the above embodiments, Figure 14 This is a structural block diagram of a tandem repeat sequence analysis device provided in an embodiment of this application. Figure 14As shown, in one embodiment, a tandem repeat sequence analysis device 1400 is provided, which includes a sequence acquisition module 1401, a sequence comparison module 1402, a fragment alignment module 1403, and a fragment detection module 1404.

[0145] The sequence acquisition module 1401 is used to acquire the first read sequence to be analyzed and the reference gene sequence. The first read sequence is a gene sequence fragment read from the original gene data, and the reference gene sequence is the gene sequence of the first exon region on the target chromosome.

[0146] The sequence comparison module 1402 is used to compare the first read sequence with the reference gene sequence to identify one or more read sequence fragments in the first read sequence that do not match the reference gene sequence.

[0147] The fragment alignment module 1403 is used to extract multiple different detection gene fragments from the reference gene sequence, and to detect the first repeat number of each detection gene fragment in the reference gene sequence, and to detect the second repeat number of each detection gene fragment in the target read sequence fragment, wherein the target read sequence fragment is any one of one or more read sequence fragments.

[0148] The fragment detection module 1404 is used to determine that there is a tandem repeat in the first read sequence if there is a second repeat number corresponding to any detection gene fragment that is not less than the number threshold and the corresponding second repeat number is greater than the corresponding first repeat number.

[0149] In some embodiments, the sequence comparison module 1402 is further configured to compare the first read sequence with the reference gene sequence, identify the differential bases in the first read sequence that do not match the reference gene sequence, and connect multiple differential bases with consecutive base positions or multiple differential bases with intervals between base positions not greater than a first interval threshold to obtain one or more read sequence fragments.

[0150] In some embodiments, the tandem repeat sequence analysis 1400 further includes a sequence extraction module.

[0151] The sequence extraction module is used to determine the reference sequence fragment corresponding to the target read sequence fragment in the reference gene sequence based on the start and end base positions of the target read sequence fragment. The start base position of the reference sequence fragment is less than or equal to the start base position of the target read sequence fragment, and the end base position of the reference sequence fragment is greater than or equal to the end base position of the target read sequence fragment.

[0152] The fragment alignment module 1403 is also used to extract multiple different detection gene fragments from the reference sequence fragment and to detect the first repeat of each detection gene fragment in the reference sequence fragment.

[0153] In some embodiments, the sequence comparison module 1402 is further configured to compare each read sequence fragment corresponding to the first read sequence with each read sequence fragment corresponding to the second read sequence to obtain the similarity between each read sequence fragment of the first read sequence and each read sequence fragment of the second read sequence, wherein the second read sequence is another gene sequence fragment read from the original gene data; if the similarity between the first read sequence fragment of the first read sequence and the second read sequence fragment of the second read sequence is greater than the similarity threshold, the first read sequence and the second read sequence are classified into the target read group.

[0154] The segment detection module 1404 is also used to determine that each segment sequence contained in the target segment group has cascaded repetition.

[0155] In some embodiments, the fragment alignment module 1403 is further configured to traverse the reference gene sequence, extract multiple detection gene fragments of length K, and the interval between the starting base positions of two adjacent detection gene fragments is N, where N is a positive integer and N is less than or equal to K, K is a positive integer greater than or equal to M, and M is a preset minimum base length.

[0156] In some embodiments, the fragment alignment module 1403 is further configured to use M as the current K to traverse the reference gene sequence and extract multiple detection gene fragments with a base length of K; if the current K is less than or equal to the target length threshold, the current K is updated to K+1, and the step of traversing the reference gene sequence and extracting multiple detection gene fragments with a base length of K is re-executed until the current K is greater than the target length threshold.

[0157] In some embodiments, the tandem repeat sequence analysis 1400 further includes a preprocessing module.

[0158] The preprocessing module is used to obtain multiple initial read sequences corresponding to the original gene data; preprocess the multiple initial read sequences to obtain multiple target read sequences that meet the first quality condition; align the multiple target read sequences with the reference genome respectively; if the alignment results determine that the multiple target read sequences meet the second quality condition, then the step of comparing the first read sequence with the reference gene sequence is executed; the first read sequence is any one of the multiple target read sequences.

[0159] The first quality condition includes one or more of the following: the read sequence does not contain a linker sequence fragment; the read sequence does not contain a polyadenylated or polyguanylic acid signal; the proportion of bases in the read sequence whose base information cannot be determined is lower than a first proportion threshold; the proportion of bases in the read sequence whose base quality value is greater than or equal to the target base quality value is higher than a second proportion threshold; and the proportion of guanine G and cytosine C bases in the read sequence is within the target range.

[0160] The second quality condition includes one or more of the following: among multiple target read sequences, the proportion of read sequences aligned to the reference genome is higher than the third proportion threshold; among multiple target read sequences, the proportion of read sequences aligned to exon regions of the reference genome is higher than the fourth proportion threshold; among multiple target read sequences, the proportion of read sequences aligned to ribosomal regions of the reference genome is lower than the fifth proportion threshold.

[0161] Figure 15 This is a structural block diagram of an electronic device provided in an embodiment of this application. Figure 15 As shown, the electronic device 1500 may include a memory 1502 and a processor 1501. The memory 1502 stores a computer program. When the computer program is executed by the processor 1501, the electronic device 1500 implements the serial repetition sequence analysis method as described in the above embodiments.

[0162] Processor 1501 may include one or more processing cores. Processor 1501 connects to various parts of the computer device using various interfaces and lines, and performs various functions and processes data by running or executing instructions, programs, code sets, or instruction sets stored in memory, and by calling data stored in memory. Optionally, processor 1501 may be implemented using at least one hardware form of digital signal processing, field-programmable gate array, or programmable logic array. Processor 1501 may integrate one or more of the following: Central Processing Unit (CPU), Graphics Processing Unit (GPU), and modem. The CPU primarily handles the operating system, user interface, and applications; the GPU is responsible for rendering and drawing displayed content; and the modem handles wireless communication. It is understood that the modem may also not be integrated into processor 1501 and may be implemented separately through a communication chip.

[0163] The memory 1502 may include random access memory (RAM) or read-only memory (ROM). The memory can be used to store instructions, programs, code, code sets, or instruction sets. The memory may include a program storage area and a data storage area. The program storage area may store instructions for implementing an operating system, instructions for implementing at least one function, instructions for implementing the various method embodiments described above, etc. The data storage area may also store data created during the use of the computer device.

[0164] This application discloses a computer-readable storage medium storing a computer program, wherein when the computer program is executed by a processor, the processor implements the serial repeat sequence analysis method as described in the above embodiments.

[0165] This application discloses a computer program product, which includes a computer program, and when executed by a processor, causes the processor to implement the serial repetition sequence analysis method as described in the above embodiments.

[0166] Those skilled in the art will understand that all or part of the processes in the above embodiments can be implemented by a computer program instructing related hardware. The program can be stored in a non-volatile computer-readable storage medium, and when executed, it can include the processes of the embodiments of the above methods. The storage medium can be a magnetic disk, optical disk, ROM, etc.

[0167] The above description is merely a specific example of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.

Claims

1. A method for analyzing tandem repeat sequences, characterized in that, The method includes: The computer device acquires a first read sequence to be analyzed and a reference gene sequence. The first read sequence is a gene sequence fragment read from the original gene data, and the reference gene sequence is a gene sequence of the first exon region on the target chromosome. The computer device compares the first read sequence with the reference gene sequence to determine one or more read sequence fragments in the first read sequence that do not match the reference gene sequence; The computer device extracts multiple different detection gene fragments from the reference gene sequence, and detects a first repeat of each detection gene fragment in the reference gene sequence, and detects a second repeat of each detection gene fragment in a target read sequence fragment, wherein the target read sequence fragment is any one of the one or more read sequence fragments; If any detected gene fragment has a second repeat number not less than the number threshold, and the corresponding second repeat number is greater than the corresponding first repeat number, the computer device determines that there is a tandem repeat in the first read sequence.

2. The method according to claim 1, characterized in that, The computer device compares the first read sequence with the reference gene sequence to determine one or more read sequence fragments in the first read sequence that do not match the reference gene sequence, including: The computer device compares the first read sequence with the reference gene sequence to determine the differential bases in the first read sequence that do not match the reference gene sequence. The computer device connects multiple differential bases with consecutive base positions or multiple differential bases with intervals between base positions not greater than a first interval threshold to obtain one or more read sequence fragments.

3. The method according to claim 2, characterized in that, After obtaining one or more read segment sequences, the method further includes: The computer device determines a reference sequence fragment in the reference gene sequence corresponding to the target read sequence fragment based on the start and end base positions of the target read sequence fragment. The start base position of the reference sequence fragment is less than or equal to the start base position of the target read sequence fragment, and the end base position of the reference sequence fragment is greater than or equal to the end base position of the target read sequence fragment. The computer device extracts multiple different detection gene fragments from the reference gene sequence and detects the first repeat number of each detection gene fragment in the reference gene sequence, including: The computer device extracts multiple different detection gene fragments from the reference sequence fragment and detects the first repeat number of each detection gene fragment in the reference sequence fragment.

4. The method according to claim 1, characterized in that, After determining one or more read sequence fragments in the first read sequence that do not match the reference gene sequence, the method further includes: The sequence fragments corresponding to the first read sequence are compared with the sequence fragments corresponding to the second read sequence to obtain the similarity between the sequence fragments of the first read sequence and the sequence fragments of the second read sequence, wherein the second read sequence is another gene sequence fragment read from the original gene data. If the similarity between the first read segment of the first read segment sequence and the second read segment of the second read segment sequence is greater than a similarity threshold, the computer device will classify the first read segment sequence and the second read segment sequence into a target read segment group. After the computer device determines that there is a cascaded repetition in the first read segment sequence, the method further includes: The computer device determines that each read segment sequence contained in the target read segment group has cascaded repetition.

5. The method according to claim 1, characterized in that, The computer device extracts multiple different detection gene fragments from the reference gene sequence, including: The computer device traverses the reference gene sequence and extracts multiple detection gene fragments of length K. The interval between the starting base positions of two adjacent detection gene fragments is N, where N is a positive integer and is less than or equal to K. K is a positive integer greater than or equal to M, and M is a preset minimum base length.

6. The method according to claim 5, characterized in that, The computer device traverses the reference gene sequence and extracts multiple detection gene fragments of length K, including: The computer device uses M as the current K to traverse the reference gene sequence and extract multiple detection gene fragments with a base length of K. If the current K is less than or equal to the target length threshold, then the current K is updated to K+1, and the step of traversing the reference gene sequence and extracting multiple detection gene fragments of length K is re-executed until the current K is greater than the target length threshold.

7. The method according to claim 1, characterized in that, Before the computer device acquires the first read sequence to be analyzed and the reference gene sequence, the method further includes: The computer device acquires multiple initial read sequences corresponding to the original gene data; The computer device preprocesses the multiple initial read segments to obtain multiple target read segments that meet the first quality condition; The computer device aligns the plurality of target read sequences with the reference genome. If the alignment results determine that the plurality of target read sequences meet the second quality condition, then the step of comparing the first read sequence with the reference genome sequence is executed; the first read sequence is any one of the plurality of target read sequences. The first quality condition includes one or more of the following conditions: The read sequence does not contain a connector sequence fragment; The read sequence does not contain polyadenylation or polyguanylate signals; The proportion of bases in the read sequence whose base information cannot be determined is lower than the first proportion threshold. The proportion of bases in the read sequence whose base quality value is greater than or equal to the target base quality value is higher than the second proportion threshold. The proportions of guanine G and cytosine C bases in the total number of bases in the read sequence are within the target range; The second quality condition includes one or more of the following conditions: Among the multiple target read sequences, the proportion of read sequences that are aligned to the reference genome is higher than the third proportion threshold. Among the multiple target read sequences, the proportion of read sequences aligned to the exon regions of the reference genome is higher than the fourth proportion threshold. Among the multiple target read sequences, the proportion of read sequences that are aligned to the ribosomal region of the reference genome is lower than the fifth proportion threshold.

8. A sequence analysis device for tandem repeats, characterized in that, The device includes: The sequence acquisition module is used to acquire the first read sequence to be analyzed and the reference gene sequence. The first read sequence is a gene sequence fragment read from the original gene data, and the reference gene sequence is the gene sequence of the first exon region on the target chromosome. The sequence comparison module is used to compare the first read sequence with the reference gene sequence to determine one or more read sequence fragments in the first read sequence that do not match the reference gene sequence. The fragment alignment module is used to extract multiple different detection gene fragments from the reference gene sequence, and to detect the first repeat number of each detection gene fragment in the reference gene sequence, and to detect the second repeat number of each detection gene fragment in the target read sequence fragment, wherein the target read sequence fragment is any one of the one or more read sequence fragments; The fragment detection module is used to determine that there is a tandem repeat in the first read sequence if there is a second repeat number corresponding to any detection gene fragment that is not less than the number threshold and the corresponding second repeat number is greater than the corresponding first repeat number.

9. An electronic device, characterized in that, The method includes a memory and a processor, wherein the memory stores a computer program that, when executed by the processor, causes the processor to perform the method as described in any one of claims 1 to 7.

10. A computer program product, the computer program product comprising a computer program, characterized in that, When the computer program is executed by the processor, it causes the processor to implement the method as described in any one of claims 1 to 7.