Subsequence search method, comparison method, system and equipment of gene sequence
Patent Information
- Application Number
- CN202380094923.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-06-21
- Publication Date
- 2025-10-21
AI Technical Summary
Existing gene sequence alignment software has problems of high delay and low search efficiency when processing gene sequences, which cannot meet the needs of gene sequence processing.
By batch obtaining the sequencing data of the gene sequence to be searched, and processing other sequencing data while the CPU is waiting for the prefetching results, the prefetching mechanism is used to obtain the base data of the search interval corresponding to each sequencing data on the reference gene sequence in advance, thereby improving the CPU's Computational efficiency.
The efficiency of gene search is improved, and the CPU's calculation time fills the time waiting for prefetch through batch processing, which significantly improves the performance of gene sequence search.
Smart Images

Figure CN120826741A_ABST
Abstract
Description
Gene sequence subsequence search method, comparison method, system and device Technical Field
[0001] The present invention relates to the field of nucleic acid information, and in particular to a subsequence search method, comparison method, system and equipment for gene sequences. Background Art
[0002] A gene sequence is a long string of four bases: AC, G, and G. One method of gene sequence alignment is to compare a short gene sequence with a reference gene sequence to identify its location within the reference sequence and any differences between the two. Existing alignment software suffers from high latency, low search efficiency, and insufficient performance when processing gene sequences, making it inadequate for gene sequence processing.
[0003] Summary of the Invention
[0004] The technical problem to be solved by the present invention is to overcome the defect of low search efficiency of gene search algorithms in the prior art and to provide a subsequence search method, comparison method, system and device for gene sequences.
[0005] The present invention solves the above technical problems through the following technical solutions:
[0006] The present invention provides a subsequence search method for a gene sequence, the subsequence search method comprising:
[0007] Obtaining a preset amount of sequencing data of the gene sequence to be searched;
[0008] Sequentially obtaining base data of a search interval corresponding to each sequencing data on a reference gene sequence, and performing a search based on the obtained results; the search interval corresponding to each sequencing data is determined based on the base sequence to be searched for each sequencing data;
[0009] Repeatedly obtain base data of the search interval for the next search and search based on the obtained results to obtain the target subsequence; the search interval for the next search is determined according to the next base to be searched in each sequencing data and the search interval of the previous search.
[0010] Preferably, the amount of sequencing data of the gene sequence to be searched is determined based on the pre-fetching duration and calculation duration of the base data in the search interval.
[0011] Preferably, before the step of sequentially obtaining base data of a search interval corresponding to each sequencing data on the reference gene sequence, the subsequence search method further comprises:
[0012] configuring a plurality of search phases for each sequencing data according to a search direction and a starting position corresponding to each sequencing data;
[0013] The steps of repeatedly obtaining base data of the search interval for the next search include:
[0014] When there is no corresponding search interval in the next search, pre-fetching of base data of the search interval is restarted in the search phase with the end position of the previous search as the starting position.
[0015] Preferably, after the step of resuming pre-fetching base data of the search interval within the search phase with the end position of the previous search as the starting position, the step of repeatedly acquiring base data of the search interval for the next search further comprises:
[0016] When all base data of the search interval corresponding to the base sequence to be searched in the current search phase are pre-fetched, the pre-fetching of base data of the search interval is restarted in the next search phase; and / or,
[0017] When all base data in the search interval corresponding to the forward base sequence to be searched of the current sequencing data are pre-fetched in the current search phase, the next sequencing data of the gene sequence to be searched is obtained; and / or,
[0018] When all base data of the search intervals in all search phases are pre-fetched, the next sequencing data of the gene sequence to be searched is obtained.
[0019] Preferably, the step of obtaining a preset amount of sequencing data of the gene sequence to be searched includes:
[0020] The sequencing data is received from the IO (input and output) thread using an empty queue of the buffer pool; the empty queue is composed of buffer areas that do not store data, and the buffer areas enter the full queue of the buffer pool after storing data;
[0021] The full queue is used to transmit the sequencing data to the comparison thread for subsequence search; after the comparison thread completes the subsequence search, the data stored in the buffer area is released.
[0022] The present invention also provides a gene sequence comparison method, the comparison method comprising:
[0023] Obtain the gene sequence to be compared;
[0024] Obtaining a target subsequence of the gene sequence to be compared using the gene sequence subsequence search method described above;
[0025] The target subsequence is expanded to obtain an alignment result of the gene sequence to be aligned.
[0026] The present invention also provides a subsequence search system for a gene sequence, the subsequence search system comprising:
[0027] A sequencing data acquisition module is used to obtain a preset amount of sequencing data of the gene sequence to be searched;
[0028] A search data acquisition module is used to sequentially acquire base data of a search interval corresponding to each sequencing data on a reference gene sequence, and perform a search based on the acquired results; the search interval corresponding to each sequencing data is determined based on the base sequence to be searched of each sequencing data;
[0029] A subsequence determination module is configured to repeatedly call the search data acquisition module to acquire base data of a search interval for a next search and perform a search based on the acquired results to obtain a target subsequence; the search interval for the next search is determined based on the next base to be searched in each sequencing data and the search interval of the previous search.
[0030] Preferably, the amount of sequencing data of the gene sequence to be searched is determined based on the pre-fetching duration and calculation duration of the base data in the search interval.
[0031] Preferably, the subsequence search system further includes:
[0032] A search phase configuration module, configured to configure multiple search phases for each sequencing data according to the search direction and starting position corresponding to each sequencing data;
[0033] The subsequence determination module is specifically configured to call the search data acquisition module to restart pre-fetching base data of the search interval in the search phase with the end position of the previous search as the starting position when there is no corresponding search interval in the next search.
[0034] Preferably, the subsequence determination module is specifically configured to, when pre-fetching of base data in the search interval corresponding to the base sequence to be searched in the current search phase is completed, call the search data acquisition module to restart pre-fetching of base data in the search interval in the next search phase; and / or,
[0035] The subsequence determination module is specifically configured to call the sequencing data acquisition module to acquire the next sequencing data of the gene sequence to be searched when all base data in the search interval corresponding to the forward base sequence to be searched of the current sequencing data are pre-fetched in the current search phase; and / or,
[0036] The subsequence determination module is specifically configured to call the sequencing data acquisition module to acquire the next sequencing data of the gene sequence to be searched when the base data of the search intervals in all search phases are pre-fetched.
[0037] Preferably, the sequencing data acquisition module is specifically configured to receive the sequencing data from the IO thread using an empty queue of the buffer pool; the empty queue is formed by arranging buffer areas that do not store data, and the buffer areas enter the full queue of the buffer pool after storing data;
[0038] The sequencing data acquisition module is specifically configured to utilize the full queue to transmit the sequencing data to the alignment thread for subsequence search; and the alignment thread releases the data stored in the buffer after completing the subsequence search.
[0039] The present invention also provides a gene sequence comparison system, the comparison system comprising:
[0040] A gene sequence acquisition module is used to obtain the gene sequence to be compared;
[0041] a target sequence determination module, configured to obtain a target subsequence of the gene sequence to be compared using the gene sequence subsequence search system described above;
[0042] The target sequence expansion module is used to expand the target subsequence to obtain the alignment result of the gene sequence to be aligned.
[0043] The present invention also provides an electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the computer program, the subsequence search method for a gene sequence or the gene sequence comparison method as described above is implemented.
[0044] The present invention also provides a computer-readable storage medium having a computer program stored thereon, wherein when the computer program is executed by a processor, the computer program implements the above-mentioned gene sequence subsequence search method or the above-mentioned gene sequence comparison method.
[0045] The positive progress effect of the present invention is:
[0046] The gene sequence subsequence search method provided by the present invention obtains sequencing data of the gene sequence to be searched in batches, and pre-extracts the base data of the search interval corresponding to each sequencing data on the reference gene sequence. By batch processing the sequencing data, the CPU can process the calculation of other sequencing data while waiting for the pre-fetch results. That is, the batch processing makes up for the CPU's calculation time by filling the waiting time for pre-fetching, thereby improving the efficiency of gene search. BRIEF DESCRIPTION OF THE DRAWINGS
[0047] To more clearly illustrate the technical solutions of the embodiments of this specification, the following briefly introduces the drawings required for describing the embodiments. Obviously, the drawings described below are merely examples or embodiments of this specification. For those skilled in the art, it is possible to apply this specification to other similar scenarios based on these drawings without inventive effort.
[0048] FIG1 is a flow chart of a subsequence search method for a gene sequence in Example 1 of the present invention.
[0049] FIG2 is an example diagram of sequencing data in FASTQ format in Example 1 of the present invention.
[0050] FIG3 is an example diagram of determining the amount of sequencing data of a gene sequence to be searched in Example 1 of the present invention.
[0051] FIG4 is an example diagram of parsed sequencing data in Example 1 of the present invention.
[0052] FIG5 is a schematic diagram of the setting search phase in Embodiment 1 of the present invention.
[0053] FIG6 is a flow chart of the gene sequence comparison method in Example 2 of the present invention.
[0054] FIG7 is a schematic diagram of the structure of a subsequence search system for a gene sequence in Example 3 of the present invention.
[0055] FIG8 is a schematic diagram of the structure of the gene sequence comparison system in Example 4 of the present invention.
[0056] FIG9 is a schematic structural diagram of an electronic device in Embodiment 5 of the present invention. DETAILED DESCRIPTION
[0057] The present invention is further described below by way of examples, but the present invention is not limited to the scope of the examples.
[0058] References herein to "embodiments" mean that a particular feature, structure, or characteristic described in connection with the embodiments may be included in at least one embodiment of the present application. The appearance of the phrase in various places herein does not necessarily refer to the same embodiment, nor does it constitute an independent or alternative embodiment that is mutually exclusive of other embodiments. It is understood, both explicitly and implicitly, by those skilled in the art that the embodiments described herein may be combined with other embodiments.
[0059] It should be understood that the terms "system," "device," "unit," and / or "module" used herein are a method for distinguishing different components, elements, parts, portions, or assemblies at different levels. However, if other terms can achieve the same purpose, the terms may be replaced by other expressions.
[0060] As used herein, unless the context clearly indicates otherwise, the terms "a," "an," "an," and / or "the" are not intended to refer to the singular but may include the plural. Generally speaking, the terms "include" and "comprise" only indicate the inclusion of the steps and elements specifically identified, and these steps and elements do not constitute an exclusive list. A method or apparatus may also include additional steps or elements.
[0061] The definition of inclusion herein, such as the terms “having”, “may have”, “include” or “may include” as used herein, indicates the existence of the corresponding functions, operations, elements, etc. herein, and does not limit the existence of one or more other functions, operations, elements, etc. In addition, it should be understood that the terms “including” or “having” as used herein indicate the existence of the features, numbers, steps, operations, elements, components or their combination described in the specification, and do not exclude the existence or addition of one or more other features, numbers, steps, operations, elements, components or their combination.
[0062] As used herein, the terms "A or B," "at least one of A and / or B," or "one or more of A and / or B" include any and all combinations of the words recited therewith. For example, "A or B," "at least one of A and B," or "at least one of A or B" means (1) including at least one A, (2) including at least one B, or (3) including both at least one A and at least one B.
[0063] Flowcharts are used herein to illustrate the operations performed by the systems according to the embodiments of the present invention. It should be understood that the preceding or following operations do not necessarily need to be performed in exact order. Instead, the steps may be processed in reverse order or simultaneously. Furthermore, other operations may be added to these processes, or one or more operations may be removed from these processes.
[0064] Example 1
[0065] The core algorithm of alignment software consists of two main parts: seed search and seed-based expansion. A seed refers to a subsequence of a sequenced base sequence. The efficiency of seed search significantly impacts alignment efficiency, so typical alignment software uses index structures such as suffix arrays (SAs) and hash tables to improve seed search efficiency during the seed phase.
[0066] Taking the SA-based binary search algorithm as an example, each search requires accessing different locations in the reference sequence, resulting in poor spatial locality and high memory latency. Furthermore, after searching for the MMP (Maximal Mappable Prefix), the matching interval must be reconfirmed in the SA array, resulting in low search efficiency and suboptimal performance. To address the shortcomings of the core algorithm of alignment software, which suffers from poor spatial locality and high memory latency, this embodiment proposes a subsequence search method for gene sequences.
[0067] Please refer to Figure 1, which is a flow chart of the subsequence search method of the gene sequence in this embodiment. Specifically, as shown in Figure 1, the subsequence search method includes:
[0068] S101. Obtain a preset amount of sequencing data for the gene sequence to be searched. Specifically, as shown in FIG2 , taking FASTQ (a file format for sequencing data) as an example, the file contains a record (name, sequence, annotation, quality value) every 4 lines, and the record is called a read.
[0069] S102. Sequentially obtain base data for a search interval corresponding to each sequencing data point on the reference gene sequence, and perform a search based on the obtained results. The search interval corresponding to each sequencing data point is determined based on the base sequence to be searched for each sequencing data point. Specifically, in an optional embodiment, the base data for the search interval can be obtained using a prefetching technique. Prefetching is a technique used by a CPU (Central Processing Unit) to extract instructions or data from a slower memory into a faster local cache before they are actually needed, thereby improving execution performance. Without a prefetching mechanism, due to dependencies between data points, most of the time is wasted waiting for related data in the memory, and the computing unit is idle most of the time. The prefetching mechanism enables the computing unit and related access units to maintain efficient operation and maximize the utilization of related hardware resources. In this embodiment, by batch processing sequencing data, the CPU can process the calculation of other reads while waiting for prefetch results. That is, batch processing allows the CPU's computation time to fill the time waiting for prefetching, which is equivalent to hiding the CPU's waiting time for prefetching, thereby improving the efficiency of gene search. Preferably, software prefetching is used.
[0070] In an alternative embodiment, the amount of sequencing data to be searched is determined based on the prefetch duration and calculation duration of the base data in the search interval. As shown in Figure 3, when the calculation of sequencing data R4 is completed, the CPU has just received the prefetch results of sequencing data R1 and can continue the search calculation based on the prefetch results of sequencing data R1. According to actual testing, this method achieves optimal performance when the number of sequencing data read in batches is 8 to 16.
[0071] S103, repeatedly obtaining base data of the search interval for the next search and searching based on the obtained results to obtain the target subsequence; the search interval for the next search is determined according to the next base to be searched in each sequencing data and the search interval of the previous search.
[0072] The steps of the above method are described in detail below by taking an example. In an optional embodiment, step S101 includes:
[0073] S1011, using an empty queue in the buffer pool to receive sequencing data from the IO thread; the empty queue is composed of buffer areas that do not store data, and the buffer areas store data before entering the full queue of the buffer pool;
[0074] S1012: Using the full queue, the sequencing data is transferred to the alignment thread for subsequence search; after the alignment thread completes the subsequence search, the data stored in the buffer is released.
[0075] In one solution, sequencing data can be analyzed by the following steps:
[0076] 1. Unzip the FASTQ file and write the reads data into the pipeline;
[0077] 2. The comparison phase is multi-threaded, with multiple threads competing for pipeline read permissions. Threads that fail the competition need to wait.
[0078] 3. The thread that successfully competes reads a certain amount of reads information from the pipeline, copies it to the thread's local cache, and then releases the read permission to wake up the thread that failed in this round of competition;
[0079] 4. The thread that failed the competition in this round will compete for the pipeline again, and the thread that failed the competition will continue to wait;
[0080] 5. If the comparison thread completes the comparison task, it will be added to the list of threads competing for pipeline reading permissions in the next round;
[0081] 6. When the FASTQ file is processed, the alignment thread ends.
[0082] In this sequencing data parsing method, reads data is copied multiple times, from the decompression cache to the pipeline, and from the pipeline to the thread-local cache pool; only one thread successfully competes each time, and the rest of the threads all wait, resulting in low thread execution efficiency.
[0083] As shown in FIG4 , this embodiment analyzes the sequencing data through the following steps:
[0084] 1. The IO thread is responsible for parsing the file, obtaining a buffer area from the empty queue of the buffer pool, writing a certain number of read records to the buffer area, and then joining the full queue;
[0085] 2. The comparison thread obtains the buffer area from the full queue, completes the MMP search and MMP expansion, and adds the buffer area to the empty queue;
[0086] It can be seen that if the comparison thread consumes the buffer faster than the IO thread generates the buffer, it will cause the comparison thread to wait for a "full" buffer, and the performance bottleneck is determined by the IO parsing thread; if the comparison thread consumes the buffer slower than the IO thread generates the buffer, it will cause the IO thread to wait for an empty buffer. In this case, the performance bottleneck is determined by the efficiency of the alignment algorithm. In this embodiment, only one read information copy is required, which reduces the number of data copies. Multiple threads execute in parallel, improving thread execution efficiency. The size of the buffer can be adjusted. Actual tests have shown that one buffer can accommodate 120k (thousand) reads.
[0087] In an optional implementation, before step S102, the subsequence search method further includes:
[0088] According to the search direction and starting position corresponding to each sequencing data, multiple search phases are configured for each sequencing data;
[0089] Specifically, a buffer is obtained from the full queue, parsed according to the read format, and the starting position of each read parameter in the buffer is recorded. All reads in the buffer are preprocessed, for example, by converting the bases in the sequence to 0, 1, 2, or 3, and initializing the parameters of each read. A fixed number of reads are processed at a time, each read searches for a single base and updates the corresponding search task.
[0090] Each read has multiple MMP search subprocesses, each with a different starting position and complex logic involved, such as searching in two directions and selecting the starting position. To simplify the process, this embodiment employs a stage structure: a stage represents a search transaction for a read. Depending on the search direction and starting position, a read can be configured with multiple stages. Each stage (off, dir) identifies the starting position and direction of the MMP search.
[0091] Each stage involves multiple MMP searches. To identify each MMP search state, this implementation uses a task structure. Task-(sa1, sa2, off, len, dir, flg) identifies the starting position (off), search direction (dir), and match length (len) of the current MMP search. The MMP range is within the search interval SA[sa1, sa2]. Tasks use FMIndex (an indexing structure used for sequence-based searches, characterized by a reverse search of one character at a time) to search one base at a time. If the search is successful, (sa1, sa2, len) is updated. If the search fails, the task ends and an MMP is recorded.
[0092] In this embodiment, step S103 may include:
[0093] S1031: When the next search does not find a corresponding search interval, prefetching base data for the search interval is restarted within the search phase, starting from the end position of the previous search. Specifically, if a task completes and sufficient matching bases remain, a new task is started at the non-matching position. Each newly created task is initialized with the pre-index structure, thereby reducing the search interval and improving FMIdex search efficiency.
[0094] In addition, after step S1031, step S103 further includes:
[0095] S1032. When the base data of the search interval corresponding to the base sequence to be searched in the current search phase are all pre-fetched, the pre-fetching of the base data of the search interval is restarted in the next search phase. Specifically, when a task structure ends, a new task needs to be started at the unmatched position, or a new task needs to be started after the stage transfer. If too few matching bases remain after a task ends, switch directly to the next stage.
[0096] S1033. When all base data in the search intervals of all search stages have been pre-fetched, the next sequencing data of the gene sequence to be searched is obtained. Specifically, when all stages of a read are executed, a new read is loaded from the buffer queue until all reads in the buffer are processed. As shown in Figure 5, the MMP search of a read involves transitions between multiple stages. For example, after stage-0 is completed, it switches to stage-1, and so on, until all stages of the read are completed and the MMP search of the current read is completed, i.e., the next sequencing data of the gene sequence to be searched is obtained.
[0097] S1034. When all base data in the search interval corresponding to the forward base sequence to be searched in the current sequencing data are pre-fetched in the current search phase, the next sequencing data of the gene sequence to be searched is obtained. Specifically, not all phases are executed, and a phase may be skipped. For example, if the MMP match length in the search phase corresponding to the forward base sequence to be searched (e.g., stage-0) is equal to the read length, the search phase corresponding to the reverse base sequence to be searched (e.g., stage-4) can be skipped, and the next sequencing data of the gene sequence to be searched is obtained.
[0098] The gene sequence subsequence search method provided in this embodiment obtains sequencing data of the gene sequence to be searched in batches and uses a prefetch mechanism to obtain base data of the search interval corresponding to each sequencing data on the reference gene sequence. By batch processing the sequencing data, the CPU can process the calculation of other sequencing data while waiting for the prefetch results. That is, the batch processing allows the CPU's calculation time to make up for the waiting time for prefetching, thereby improving the efficiency of gene search.
[0099] Example 2
[0100] Please refer to Figure 6, which is a flow chart of the gene sequence comparison method in this embodiment. Specifically, as shown in Figure 6, the comparison method includes:
[0101] S201, obtaining the gene sequence to be compared;
[0102] S202, using the gene sequence subsequence search method of Example 1 to obtain a target subsequence of the gene sequence to be compared;
[0103] S203: Expand the target subsequence to obtain the alignment result of the gene sequence to be aligned.
[0104] The gene sequence comparison method provided in this embodiment utilizes the above-mentioned subsequence search method, allowing the CPU to process other sequencing data calculations while waiting for prefetch results. That is, through batch processing, the CPU's calculation time makes up for the waiting time for prefetching, thereby improving the efficiency of gene search.
[0105] Example 3
[0106] Please refer to Figure 7, which is a schematic diagram of the structure of the subsequence search system of the gene sequence in this embodiment. Specifically, as shown in Figure 7, the subsequence search system includes:
[0107] The sequencing data acquisition module 1 is used to obtain a preset amount of sequencing data of the gene sequence to be searched; specifically, as shown in Figure 2, taking FASTQ as an example, every 4 lines of the file are a record (name, sequence, annotation, quality value), and the record is called reads.
[0108] The search data acquisition module 2 is used to sequentially acquire the base data of the search interval corresponding to each sequencing data on the reference gene sequence and perform a search based on the acquisition results; the search interval corresponding to each sequencing data is determined based on the base sequence to be searched of each sequencing data; specifically, in an optional embodiment, the search data acquisition module 2 can acquire the base data of the search interval by prefetching technology; prefetching is a technology used by the CPU to extract the data from the slower memory to the faster local cache (cache memory) before the instruction or data is actually needed, thereby improving execution performance; if there is no software prefetching mechanism, due to the dependency between the data, most of the time is wasted waiting for the relevant data in the memory, and the computing unit is idle most of the time. The prefetching mechanism enables the computing unit and the related access unit to maintain efficient operation and maximize the utilization of related hardware resources; in this embodiment, by batch processing the sequencing data, the CPU can process the calculation of other reads while waiting for the prefetch results, that is, by batch processing, the CPU's calculation time fills the time waiting for prefetching, which is equivalent to hiding the time the CPU waits for prefetching, thereby improving the efficiency of gene search. Preferably, the search data acquisition module 2 can adopt a software prefetching method.
[0109] In an optional embodiment, the number of sequencing data points to be searched is determined based on the prefetch duration and calculation duration of the base data in the search interval. As shown in Figure 3, when the calculation of sequencing data R4 is completed, the CPU has just received the prefetch results of sequencing data R1 and can continue the search calculation based on the prefetch results of sequencing data R1. According to actual testing, this method achieves optimal performance when the number of sequencing data points read in batches is 8 to 16.
[0110] The subsequence determination module 3 is used to repeatedly call the search data acquisition module to obtain the base data of the search interval for the next search and search based on the obtained results to obtain the target subsequence; the search interval for the next search is determined based on the next base to be searched in each sequencing data and the search interval of the previous search.
[0111] The steps of the above method are described in detail below by way of example. In an optional embodiment, the sequencing data acquisition module 1 is specifically configured to receive sequencing data from an IO thread using an empty queue of a buffer pool; the empty queue is composed of buffer areas that do not store data, and the buffer areas that store data enter the full queue of the buffer pool;
[0112] The sequencing data acquisition module 1 is specifically configured to utilize a full queue to transmit sequencing data to an alignment thread for performing a subsequence search; and the alignment thread releases the data stored in the buffer after completing the subsequence search.
[0113] In one solution, sequencing data can be analyzed by the following methods:
[0114] 1. Unzip the FASTQ file and write the reads data into the pipeline;
[0115] 2. The comparison phase is multi-threaded, with multiple threads competing for pipeline read permissions. Threads that fail the competition need to wait.
[0116] 3. The thread that successfully competes reads a certain amount of reads information from the pipeline, copies it to the thread's local cache, and then releases the read permission to wake up the thread that failed in this round of competition;
[0117] 4. The thread that failed the competition in this round will compete for the pipeline again, and the thread that failed the competition will continue to wait;
[0118] 5. If the comparison thread completes the comparison task, it will be added to the list of threads competing for pipeline reading permissions in the next round;
[0119] 6. When the FASTQ file is processed, the alignment thread ends.
[0120] In this sequencing data parsing method, reads data is copied multiple times, from the decompression cache to the pipeline, and from the pipeline to the thread-local cache pool; only one thread successfully competes each time, and the rest of the threads all wait, resulting in low thread execution efficiency.
[0121] As shown in FIG4 , the sequencing data acquisition module 1 of this embodiment specifically analyzes the sequencing data through the following steps:
[0122] 1. The IO thread is responsible for parsing the file, obtaining a buffer area from the empty queue of the buffer pool, writing a certain number of read records to the buffer area, and then joining the full queue;
[0123] 2. The comparison thread obtains the buffer area from the full queue, completes the MMP search and MMP expansion, and adds the buffer area to the empty queue;
[0124] As can be seen, if the comparison thread consumes the buffer faster than the IO thread generates it, it will wait for a full buffer, and the performance bottleneck will be determined by the IO parsing thread. If the comparison thread consumes the buffer slower than the IO thread generates it, it will wait for an empty buffer, and the efficiency of the alignment algorithm will determine the performance bottleneck. In this case, only one read information copy is required, reducing the number of data copies. Multiple threads can execute in parallel, improving thread execution efficiency. The buffer size can be adjusted. Actual tests have shown that one buffer can accommodate 120k reads.
[0125] In an optional embodiment, the subsequence search system further includes:
[0126] A search phase configuration module 4, configured to configure multiple search phases for each sequencing data according to the search direction and starting position corresponding to each sequencing data;
[0127] Specifically, a buffer is obtained from the full queue, parsed according to the read format, and the starting position of each read parameter in the buffer is recorded. All reads in the buffer are preprocessed, for example, by converting the bases in the sequence to 0, 1, 2, or 3, and initializing the parameters of each read. A fixed number of reads are processed at a time, each read searches for a single base and updates the corresponding search task.
[0128] Each read has multiple MMP search subprocesses, each with a different starting position and complex logic involved, such as searching in two directions and selecting the starting position. To simplify the process, this embodiment employs a stage structure: a stage represents a search transaction for a read. Depending on the search direction and starting position, a read can be configured with multiple stages. Each stage (off, dir) identifies the starting position and direction of the MMP search.
[0129] For each stage, multiple MMP searches occur. To identify each MMP search state, this implementation uses a task structure. Task-(sa1, sa2, off, len, dir, flg) identifies the starting position (off), search direction (dir), and match length (len) of the current MMP search. The MMP range is within the search interval SA[sa1, sa2]. Task searches one base at a time using FMIndex (an indexing structure used for sequence-based searches, characterized by a reverse search of one character at a time). If the search is successful, (sa1, sa2, len) is updated. If the search fails, the task ends and an MMP is recorded.
[0130] In this embodiment, subsequence determination module 3 is specifically configured to, when the next search does not contain a corresponding search interval, invoke search data acquisition module 2 to restart prefetching base data for the search interval within the search phase, starting from the end position of the previous search. Specifically, if a task completes and sufficient matching bases remain, a new task is started at the non-matching position. Each newly created task is initialized with a pre-index structure, thereby reducing the search interval and improving FMIdex search efficiency.
[0131] In addition, the subsequence determination module 3 is specifically used to call the search data acquisition module 2 to restart the pre-fetching of base data in the search interval in the next search phase when the base data in the search interval corresponding to the base sequence to be searched in the current search phase are all pre-fetched; specifically, when a task structure ends, it is necessary to start a new task at the unmatched position, or start a new task after the stage transfer; if there are too few remaining matching bases after a task ends, switch directly to the next stage.
[0132] The subsequence determination module 3 is specifically configured to, when all base data in the search interval corresponding to the forward base sequence to be searched in the current sequencing data are pre-fetched in the current search phase, call the sequencing data acquisition module 1 to obtain the next sequencing data of the gene sequence to be searched. Specifically, not all phases are executed, and a phase may be skipped. For example, if the MMP match length in the search phase corresponding to the forward base sequence to be searched (e.g., stage-0) is equal to the read length, the search phase corresponding to the reverse base sequence to be searched (e.g., stage-4) can be skipped, and the next sequencing data of the gene sequence to be searched can be obtained.
[0133] Subsequence determination module 3 is specifically configured to call sequencing data acquisition module 1 to retrieve the next sequencing data of the gene sequence to be searched after all base data in the search interval of all search stages have been pre-fetched. Specifically, after all stages of a read have been executed, a new read is reloaded from the buffer queue until all reads in the buffer are processed. As shown in Figure 5, the MMP search for a read involves transitioning between multiple stages. For example, after stage-0 is completed, it switches to stage-1, and so on, until all stages of the read are completed and the MMP search for the current read is completed, i.e., the next sequencing data of the gene sequence to be searched is obtained.
[0134] The gene sequence subsequence search system provided in this embodiment obtains sequencing data of the gene sequence to be searched in batches and uses a prefetch mechanism to obtain base data of the search interval corresponding to each sequencing data on the reference gene sequence. By batch processing the sequencing data, the CPU can process the calculation of other sequencing data while waiting for the prefetch results. That is, the batch processing allows the CPU's calculation time to make up for the waiting time for prefetching, thereby improving the efficiency of gene search.
[0135] Example 4
[0136] Please refer to Figure 8, which is a schematic diagram of the structure of the gene sequence comparison system in this embodiment. Specifically, as shown in Figure 8, the comparison system includes:
[0137] Gene sequence acquisition module 5, used to obtain the gene sequence to be compared;
[0138] a target sequence determination module 6, configured to obtain a target subsequence of the gene sequence to be compared using the gene sequence subsequence search system of Example 3;
[0139] The target sequence expansion module 7 is used to expand the target subsequence to obtain the alignment result of the gene sequence to be aligned.
[0140] The gene sequence alignment system provided in this embodiment utilizes the above-mentioned subsequence search system, allowing the CPU to process other sequencing data calculations while waiting for pre-fetch results. That is, through batch processing, the CPU's calculation time makes up for the waiting time for pre-fetching, thereby improving the efficiency of gene search.
[0141] Example 5
[0142] Figure 9 is a schematic diagram of the structure of an electronic device provided in Example 5 of the present invention. The electronic device includes a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the program, it implements the gene sequence subsequence search method of Example 1 or the gene sequence comparison method of Example 2. The electronic device 30 shown in Figure 9 is merely an example and should not limit the functionality or scope of use of the embodiments of the present invention.
[0143] As shown in FIG9 , the electronic device 30 may be a general-purpose computing device, such as a server device. Components of the electronic device 30 may include, but are not limited to, the at least one processor 31, the at least one memory 32, and a bus 33 connecting various system components (including the memory 32 and the processor 31).
[0144] The bus 33 includes a data bus, an address bus, and a control bus.
[0145] The memory 32 may include a volatile memory, such as a random access memory (RAM) 321 and / or a cache memory 322 , and may further include a read-only memory (ROM) 323 .
[0146] The memory 32 may also include a program / utility 325 having a set (at least one) of program modules 324, such program modules 324 including but not limited to: an operating system, one or more application programs, other program modules, and program data, each of which or some combination may include an implementation of a network environment.
[0147] The processor 31 executes various functional applications and data processing by running computer programs stored in the memory 32, such as the gene sequence subsequence search method of Example 1 or the gene sequence comparison method of Example 2 of the present invention.
[0148] The electronic device 30 can also communicate with one or more external devices 34 (e.g., a keyboard, pointing device, etc.). This communication can occur via an input / output (I / O) interface 35. Furthermore, the model-generating device 30 can also communicate with one or more networks (e.g., a local area network (LAN), a wide area network (WAN), and / or a public network, such as the Internet) via a network adapter 36. As shown, the network adapter 36 communicates with other modules of the model-generating device 30 via a bus 33. It should be understood that, although not shown, other hardware and / or software modules can be used in conjunction with the model-generating device 30, including but not limited to microcode, device drivers, redundant processors, external disk drive arrays, RAID (RAID) systems, tape drives, and data backup storage systems.
[0149] It should be noted that although several units / modules or sub-units / modules of the electronic device are mentioned in the detailed description above, this division is merely exemplary and not mandatory. In fact, according to embodiments of the present invention, the features and functions of two or more units / modules described above may be embodied in a single unit / module. Conversely, the features and functions of a single unit / module described above may be further divided and embodied by multiple units / modules.
[0150] Example 4
[0151] This embodiment provides a computer-readable storage medium having a computer program stored thereon. When the program is executed by a processor, the method for searching for subsequences of gene sequences in embodiment 1 or the method for comparing gene sequences in embodiment 2 is implemented.
[0152] The readable storage medium may include, but is not limited to, a portable disk, a hard disk, a random access memory, a read-only memory, an erasable programmable read-only memory, an optical storage device, a magnetic storage device, or any suitable combination thereof.
[0153] In a possible embodiment, the present invention can also be implemented in the form of a program product, which includes a program code. When the program product is run on a terminal device, the program code is used to enable the terminal device to execute the subsequence search method for the gene sequence of Example 1 or the gene sequence comparison method of Example 2.
[0154] The program code for executing the present invention may be written in any combination of one or more programming languages, and may be executed entirely on the user device, partially on the user device, as an independent software package, partially on the user device and partially on a remote device, or entirely on the remote device.
[0155] Although specific embodiments of the present invention have been described above, those skilled in the art will appreciate that these are merely illustrative and that the scope of the present invention is defined by the appended claims. Those skilled in the art may make various changes or modifications to these embodiments without departing from the principles and essence of the present invention, and such changes and modifications are intended to fall within the scope of the present invention.
Claims
1. A method for searching subsequences of a gene sequence, characterized in that: The subsequence search method comprises: Obtaining a preset amount of sequencing data of the gene sequence to be searched; Sequentially obtaining base data of a search interval corresponding to each sequencing data on a reference gene sequence, and performing a search based on the obtained results; the search interval corresponding to each sequencing data is determined based on a base sequence to be searched for each sequencing data; Repeatedly obtain base data of the search interval for the next search and search based on the obtained results to obtain the target subsequence; the search interval for the next search is determined according to the next base to be searched for each sequencing data and the search interval of the previous search.
2. The subsequence search method according to claim 1, characterized in that: The amount of sequencing data of the gene sequence to be searched is determined based on the pre-fetching duration and calculation duration of the base data in the search interval.
3. The subsequence search method according to claim 1, characterized in that: Before the step of sequentially acquiring the base data of the search interval corresponding to each sequencing data on the reference gene sequence, the subsequence search method further includes: According to the search direction and the starting position corresponding to each sequencing data, configuring multiple search phases for each sequencing data; The step of repeatedly obtaining base data of the search interval for the next search includes: When there is no corresponding search interval in the next search, the pre-fetching of base data in the search interval is restarted in the search phase with the end position of the previous search as the starting position.
4. The subsequence search method according to claim 3, characterized in that: After the step of re-starting the pre-fetching of base data of the search interval in the search phase with the end position of the previous search as the starting position, the step of repeatedly acquiring base data of the search interval for the next search further includes: When the base data of the search interval corresponding to the base sequence to be searched in the current search phase are all pre-fetched, the pre-fetching of the base data of the search interval is restarted in the next search phase; and / or, When the base data of the search interval corresponding to the forward base sequence to be searched of the current sequencing data in the current search phase are all pre-fetched, the next sequencing data of the gene sequence to be searched is obtained; and / or, When the base data of the search intervals in all search phases are pre-fetched, the next sequencing data of the gene sequence to be searched is obtained.
5. The subsequence search method according to claim 1, characterized in that: The steps of obtaining a preset amount of sequencing data of the gene sequence to be searched include: The sequencing data is received from the IO thread using an empty queue of the buffer pool; the empty queue is composed of buffer areas that do not store data, and the buffer areas enter the full queue of the buffer pool after storing data; The full queue is used to transmit the sequencing data to the comparison thread for subsequence search; after the comparison thread completes the subsequence search, the data stored in the buffer area is released.
6. A method for comparing gene sequences, characterized in that: The comparison method comprises: Obtaining the gene sequence to be compared; Obtaining a target subsequence of the gene sequence to be compared using the gene sequence subsequence search method according to any one of claims 1 to 5; The target subsequence is expanded to obtain the alignment result of the gene sequence to be aligned.
7. A subsequence search system for a gene sequence, characterized in that: The subsequence search system comprises: A sequencing data acquisition module, used to acquire a preset amount of sequencing data of the gene sequence to be searched; A search data acquisition module, used to sequentially acquire base data of a search interval corresponding to each sequencing data on a reference gene sequence, and perform a search based on the acquisition result; the search interval corresponding to each sequencing data is determined based on a base sequence to be searched for each sequencing data; The subsequence determination module is used to repeatedly call the search data pre-fetching module to obtain the base data of the search interval for the next search and search based on the obtained results to obtain the target subsequence; the search interval for the next search is based on the next base to be searched of each sequencing data and the previous base. The search interval for the search is determined.
8. A gene sequence comparison system, characterized in that: The comparison system comprises: A gene sequence acquisition module is used to acquire the gene sequence to be compared; A target sequence determination module, used to obtain a target subsequence of the gene sequence to be compared using the gene sequence subsequence search system as claimed in claim 7; The target sequence expansion module is used to expand the target subsequence to obtain the comparison result of the gene sequence to be compared.
9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the computer program, the subsequence search method for the gene sequence according to any one of claims 1 to 5 or the gene sequence comparison method according to claim 6 is implemented.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the subsequence search method for a gene sequence according to any one of claims 1 to 5 or the gene sequence comparison method according to claim 6 is implemented.