Nanopore-sequencing-based sequence processing method and apparatus, and device and storage medium

By statistically analyzing and trimming specific sequences in nanopore sequencing sequences, the sequencing error problem caused by homopolymers and dimers is solved, improving sequencing accuracy and quality, and making it suitable for genome analysis of new species.

WO2026156563A1PCT designated stage Publication Date: 2026-07-30BGI HANGZHOU CYCLONESEQ TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
BGI HANGZHOU CYCLONESEQ TECHNOLOGY CO LTD
Filing Date
2025-01-22
Publication Date
2026-07-30

AI Technical Summary

Technical Problem

The high reproducibility of homopolymers and dimers in nanopore sequencing leads to confusion in electrical signal patterns, increasing the risk of sequencing errors and affecting the quality of the sequencing sequence.

Method used

By statistically analyzing the proportion of special sequences in the original sequencing sequence, the detection threshold is determined. The tail detection window is then used to adjust and trim the original sequencing sequence, removing tail homopolymers and dimers to obtain high-quality target sequencing sequences.

Benefits of technology

It improves the accuracy and quality of sequencing sequences, reduces the difficulty of identifying tail homopolymers and dimers, and is suitable for analyzing genome data of new species or lacking comprehensive annotation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2025074080_30072026_PF_FP_ABST
    Figure CN2025074080_30072026_PF_FP_ABST
Patent Text Reader

Abstract

Provided are a nanopore-sequencing-based sequence processing method and apparatus, and a device and a storage medium, which belong to the technical field of single-molecule sequencing. The method comprises: acquiring a raw sequencing sequence, wherein the raw sequencing sequence is acquired by means of nanopore sequencing technology, and the raw sequencing sequence meets preset sequencing sequence requirement parameters; calculating the proportion of special sequences in a middle region of the raw sequencing sequence, in order to obtain first proportion information; on the basis of the first proportion information, determining a detection threshold; calculating, on the basis of a preset tail detection window, the proportion of special sequences in the portion of the raw sequencing sequence within the tail detection window, in order to obtain second proportion information; and on the basis of the second proportion information, the detection threshold and the tail detection window, performing trimming processing on the raw sequencing sequence, in order to obtain a target sequencing sequence. The provided method can improve the quality of sequencing sequences.
Need to check novelty before this filing date? Find Prior Art

Description

Nanopore sequencing-based sequence processing methods, devices, equipment, and storage media Technical Field

[0001] This application relates to the field of single-molecule sequencing technology, and in particular to a sequence processing method, apparatus, device, and storage medium based on nanopore sequencing. Background Technology

[0002] Nanopore sequencing uses nanopores embedded in biological membranes as sensors to detect the passage of individual DNA molecules. DNA molecules are drawn into the nanopore under an electric field, generating a change in electrical current, which is then used to read the DNA sequence. However, when homopolymeric structures (multiple identical nucleotides) or dimeric structures (e.g., the ATAT structure) pass sequentially through the nanopore, the high reproducibility can lead to confusion in the electrical signal patterns, increasing the risk of sequencing errors and affecting the quality of the sequence. Therefore, improving the quality of sequencing sequences has become a pressing technical problem. Summary of the Invention

[0003] The main objective of this application is to propose a sequence processing method, apparatus, device, and storage medium based on nanopore sequencing, aiming to improve the quality of sequencing sequences.

[0004] To achieve the above objectives, a first aspect of this application proposes a sequence processing method based on nanopore sequencing, the method comprising:

[0005] Obtain the raw sequencing sequence; wherein the raw sequencing sequence is obtained through nanopore sequencing technology, and the raw sequencing sequence meets the preset sequencing sequence requirement parameters;

[0006] The proportion of special sequences in the middle region of the original sequencing sequence is statistically analyzed to obtain the first proportion information;

[0007] The detection threshold is determined based on the first proportion information;

[0008] Based on the preset tail detection window, the proportion of special sequences in the portion of the original sequencing sequence within the tail detection window is statistically analyzed to obtain the second proportion information;

[0009] The original sequencing sequence is trimmed based on the second proportion information, the detection threshold, and the tail detection window to obtain the target sequencing sequence.

[0010] In some embodiments, the step of statistically analyzing the proportion of special sequences in the middle region of the original sequencing sequence to obtain first proportion information includes:

[0011] A reference region sequence is extracted from the original sequencing sequence to obtain a reference sequencing sequence; wherein the reference region indicates a region located in the middle of the original sequencing sequence;

[0012] Based on the regular expression matching pattern, special sequences within the reference sequencing sequence are extracted to obtain the selected sequencing sequence.

[0013] The ratio of the sequence lengths of the selected sequencing sequence to the reference sequencing sequence is obtained as the first proportion information.

[0014] In some embodiments, the selected sequencing sequence includes at least one of the following: homopolymers and dimers; the step of extracting specific sequences from the reference sequencing sequence to obtain the selected sequencing sequence according to a preset regular expression matching pattern includes at least one of the following steps:

[0015] Homopolymers within the reference sequencing sequence are extracted based on the regular expression matching pattern.

[0016] Dimers within the reference sequencing sequence are extracted based on the regular expression matching pattern.

[0017] In some embodiments, determining the detection threshold based on the first proportion information includes:

[0018] The preset Bayesian prior distribution parameters are updated based on the first proportion information to obtain the updated posterior distribution parameters.

[0019] Based on the posterior distribution parameters, a posterior distribution is constructed to determine the detection threshold.

[0020] In some embodiments, the step of calculating the proportion of special sequences in a portion of the original sequencing sequence within a preset tail detection window to obtain second proportion information includes:

[0021] Based on the tail detection window, obtain the tail sequencing sequence of the tail region of the original sequencing sequence;

[0022] Based on the regular expression matching pattern, special sequences are extracted from the tail sequencing sequence to obtain the selected sequencing sequence.

[0023] Obtain the length of the selected sequencing sequence to get the selected sequence length;

[0024] The ratio between the selected sequence length and the window length of the tail detection window is obtained to obtain the second proportion information.

[0025] In some embodiments, the step of trimming the original sequencing sequence according to the second proportion information, the detection threshold, and the tail detection window to obtain the target sequencing sequence includes:

[0026] The tail detection window is updated based on the second proportion information of the tail detection window and the detection threshold.

[0027] The detection threshold is updated according to the preset threshold adjustment ratio;

[0028] The final tail detection window is determined based on the updated tail detection window and detection threshold.

[0029] The original sequencing sequence is pruned at the tail region based on the final tail detection window to obtain the target sequencing sequence.

[0030] In some embodiments, updating the tail detection window based on the second proportion information of the tail detection window and the detection threshold includes:

[0031] If the second proportion information is greater than the detection threshold, the tail detection window is enlarged according to the preset enlargement ratio;

[0032] If the second percentage information is less than or equal to the detection threshold, the tail detection window is reduced according to the preset reduction ratio.

[0033] In some embodiments, determining the final tail detection window based on the updated tail detection window and detection threshold includes:

[0034] Repeat the steps of updating the tail detection window based on the second proportion information of the tail detection window and the detection threshold, and adjusting the proportion of the detection threshold according to a preset threshold, until the second proportion information of the updated tail detection window is not greater than the updated detection threshold, to obtain the final tail detection window.

[0035] In some embodiments, obtaining the raw sequencing sequence includes:

[0036] Obtain candidate sequencing sequences and sequencing sequence requirement parameters; wherein, the sequencing sequence requirement parameters include: reference length requirement information and sequence quantity requirement information;

[0037] The candidate sequencing sequences are screened according to the reference length requirement information, and the candidate sequencing sequences whose sequence length meets the requirement are selected as sequencing sequences.

[0038] The selected sequencing sequences are filtered according to the sequence quantity requirement information to obtain the original sequencing sequences.

[0039] In some embodiments, after trimming the original sequencing sequence according to the second proportion information, the detection threshold, and the tail detection window to obtain the target sequencing sequence, the method further includes:

[0040] Obtain the processing instruction request for the target sequencing sequence;

[0041] The target sequencing sequence is requested to be quality evaluated according to the processing instruction, and a quality evaluation value is obtained;

[0042] According to the processing instruction, a target downstream operation is selected from a preset pool of candidate downstream operations; wherein, the target downstream operation includes genome assembly operation and variant detection operation;

[0043] If the quality assessment value is greater than the preset quality threshold, then the target downstream operation is performed on the target single gene molecule according to the target sequencing sequence.

[0044] To achieve the above objectives, a second aspect of this application provides a sequence processing apparatus based on nanopore sequencing, the apparatus comprising:

[0045] A raw sequence acquisition module is used to acquire raw sequencing sequences; wherein, the raw sequencing sequences are acquired through nanopore sequencing technology, and the raw sequencing sequences meet preset sequencing sequence requirement parameters;

[0046] The first statistical module is used to calculate the proportion of special sequences in the middle region of the original sequencing sequence to obtain first proportion information.

[0047] The threshold setting module is used to determine the detection threshold based on the first proportion information;

[0048] The second statistical module is used to calculate the proportion of special sequences in the portion of the original sequencing sequence within the tail detection window according to the preset tail detection window, and obtain the second proportion information.

[0049] The sequence trimming module is used to trim the original sequencing sequence according to the second proportion information, the detection threshold, and the tail detection window to obtain the target sequencing sequence.

[0050] To achieve the above objectives, a third aspect of this application provides an electronic device, which includes a memory and a processor. The memory stores a computer program, and the processor executes the computer program to implement the method described in the first aspect.

[0051] To achieve the above objectives, a fourth aspect of the present application provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the method described in the first aspect.

[0052] This application proposes a nanopore sequencing-based sequence processing method, apparatus, device, and storage medium. It determines the proportion of special sequences in the original sequencing sequence and establishes a detection threshold based on this proportion. Then, when trimming the target sequencing sequence, the proportion of special sequences in the original sequencing sequence is further determined based on the tail detection window. The original sequencing sequence is then trimmed according to the proportion of special sequences, the detection threshold, and the tail detection window to obtain a target sequencing sequence with higher accuracy and quality. Therefore, the sequence processing method proposed in this application does not rely on a reference genome or database during the detection and trimming of special sequences in the sequencing sequence. It is suitable for analyzing genomes of new species or genome data lacking comprehensive annotation. Furthermore, by adjusting the tail detection window, it reduces the difficulty of identifying tail homopolymers and dimers, trimming the special sequences corresponding to tail homopolymers and dimers to obtain sequencing sequences with higher accuracy and quality. Attached Figure Description

[0053] Figure 1 is a flowchart of the sequence processing method based on nanopore sequencing provided in an embodiment of this application;

[0054] Figure 2 is a line graph showing the percentage of four bases at the tail of the candidate sequencing sequence in the embodiments of this application;

[0055] Figure 3 is a schematic diagram of the distribution of homopolymer and dimer sequences in the tail and middle regions in the embodiments of this application;

[0056] Figure 4 is a schematic diagram of the distribution of head, tail, and middle region dimer sequences in the candidate sequencing sequences in the embodiments of this application;

[0057] Figure 5 is a schematic diagram of the distribution of homopolymer sequences in the head, tail, and middle regions of the candidate sequencing sequences in the embodiments of this application;

[0058] Figure 6 is a line graph showing the percentage of the four bases at the tail of the target sequencing sequence in an embodiment of this application;

[0059] Figure 7 is a schematic diagram of the distribution of the homopolymer and dimer sequences of the target sequencing sequence in the tail and middle regions in the embodiments of this application;

[0060] Figure 8 is a schematic diagram of the distribution of the tail and middle region dimer sequences of the target sequencing sequence in the embodiments of this application;

[0061] Figure 9 is a schematic diagram of the distribution of the tail and middle region homopolymer sequences of the target sequencing sequence in the embodiments of this application;

[0062] Figure 10(a) is a schematic diagram of the tail quality of the original sequencing sequence in the embodiment of this application;

[0063] Figure 10(b) is a schematic diagram of the tail quality of the target sequencing sequence in an embodiment of this application;

[0064] Figure 11 is a schematic diagram of the sequence processing device based on nanopore sequencing provided in an embodiment of this application;

[0065] Figure 12 is a schematic diagram of the hardware structure of the electronic device provided in an embodiment of this application;

[0066] Figure 13 is a flowchart of a sequence processing method based on nanopore sequencing provided in another embodiment of this application. Detailed Implementation

[0067] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.

[0068] It should be noted that although functional modules are divided in the device schematic diagram and a logical order is shown in the flowchart, in some cases, the steps shown or described may be performed in a different order than the module division in the device or the order in the flowchart. The terms "first," "second," etc., in the specification, claims, and the aforementioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence.

[0069] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application belongs. The terminology used herein is for the purpose of describing embodiments of this application only and is not intended to limit this application.

[0070] First, let's analyze some of the terms used in this application:

[0071] Single-molecule sequencing is a high-throughput gene sequencing technology that can directly sequence individual DNA molecules without the need for PCR amplification. Its advantages include the ability to detect DNA modifications (such as methylation) and to handle longer reads, typically exceeding several thousand base pairs.

[0072] Nanopore sequencing is a single-molecule sequencing technology that stretches DNA molecules through nanopores and detects their sequence. This technology enables efficient, low-cost, and high-precision DNA sequence sequencing. The principle involves injecting DNA molecules into nanopores, then using an electric field or other methods to stretch the DNA molecules, and finally detecting changes in the electrical signals at both ends of the nanopore to determine the DNA sequence.

[0073] Nucleotide homopolymers are polymers composed of monomers of the same type linked together by phosphodiester bonds. Nucleotides are the basic building blocks of DNA and RNA, including adenine (A), thymine (T), cytosine (C), guanine (G), and uracil (U).

[0074] Nucleotide dimers are molecules formed by two nucleotide monomers linked by a phosphate bond. These nucleotide monomers can be the same type of nucleotide.

[0075] Bayes' theorem is an important theorem in probability theory. It describes the probability of one event occurring given that some other events have already occurred. More specifically, Bayes describes how to update our beliefs about the probability of an event after receiving new evidence or information.

[0076] Regular expressions are powerful text processing tools used to search, replace, and retrieve text that matches a given pattern (rule). A regular expression consists of a series of characters, some of which have special meanings and are used to define the search pattern.

[0077] FASTQ is a text file format used to store sequencing data, which includes nucleotide sequences, sequencing quality scores for each sequence position, and information describing the origin of the reads.

[0078] Base separation: This usually refers to the process of identifying and distinguishing the bases (A, T, C, G) in DNA or RNA molecules during sequencing or molecular biology experiments.

[0079] Single-molecule sequencing technology, especially nanopore sequencing, has been widely used in the field of genomics. Nanopore sequencing reads DNA sequence information in real time by detecting changes in electrical current generated when a single DNA molecule passes through a nanopore. Compared to other sequencing technologies, nanopore sequencing has a significant advantage in generating ultra-long reads, which provides higher-quality sequencing data for analyzing complex genomic structural variations, repetitive sequence regions, and whole-genome assembly. Specifically, nanopore sequencing uses nanopores embedded in biological membranes as sensors. DNA molecules are pulled into the nanopores under the influence of an electric field. When different nucleotides in the DNA molecule pass through the nanopore, they interfere with the ion current passing through the pore, thereby generating a unique electrical signal. The DNA sequence is obtained by decoding this electrical signal. Each time a DNA sequence is read, the nanopore simultaneously senses a fragment composed of 3 to 6 nucleotides. This sensing process complicates the recognition of specific sequence patterns (such as homopolymers and dimers, hereinafter referred to as homopolymers and dimers).

[0080] In nanopore sequencing technology, single-sensor and dual-sensor configurations are commonly used. In a single-sensor configuration, the electrical signal is captured by a single nanopore sensor; in a dual-sensor configuration, the electrical signal is sensed simultaneously through two adjacent nanopores. However, regardless of the configuration, when identifying homopolymers, the current signal generated when multiple identical nucleotides pass through the nanopore consecutively is difficult to distinguish between different homopolymer lengths, leading to sequencing errors, especially at the tail end of the read. When dimers pass through the nanopore, the high reproducibility of dimers makes the electrical signals easily confused, further increasing the risk of sequencing errors. Therefore, nanopore sequencing technology lacks accuracy in identifying homopolymers versus dimers, affecting the quality of sequencing data.

[0081] Based on this, embodiments of this application provide a sequence processing method, apparatus, device, and storage medium based on nanopore sequencing. The aim is to calculate the proportion of special sequences in the original sequencing sequence using regular expression matching patterns, and then determine a detection threshold based on this proportion. Therefore, after calculating the proportion of special sequences in the original sequencing sequence, the sequence trimming window is adjusted based on the calculated proportion of special sequences and abnormal detection regions. The original sequencing sequence is then trimmed according to the adjusted window to obtain a target sequencing sequence with smaller special sequences and higher quality. This not only improves the accuracy of the sequencing sequence but also brings greater flexibility and reliability to genomics research.

[0082] The sequence processing method, apparatus, device, and storage medium based on nanopore sequencing provided in this application are specifically described through the following embodiments. First, the sequence processing method based on nanopore sequencing in this application is described.

[0083] The embodiments of this application can acquire and process relevant data based on artificial intelligence technology. Artificial intelligence (AI) refers to the theories, methods, technologies, and application systems that use digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use that knowledge to obtain optimal results.

[0084] Foundational technologies for artificial intelligence generally include sensors, dedicated AI chips, cloud computing, distributed storage, big data processing, operating / interactive systems, and mechatronics. AI software technologies mainly encompass computer vision, robotics, biometrics, speech processing, natural language processing, and machine learning / deep learning.

[0085] The nanopore sequencing-based sequence processing method provided in this application relates to the field of single-molecule sequencing technology. This nanopore sequencing-based sequence processing method can be applied to a terminal, a server, or software running on either a terminal or a server. In some embodiments, the terminal can be a smartphone, tablet, laptop, desktop computer, etc.; the server can be configured as an independent physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms; the software can be an application implementing the nanopore sequencing-based sequence processing method, but is not limited to the above forms.

[0086] This application can be used in a wide variety of general-purpose or special-purpose computer system environments or configurations. Examples include: personal computers, server computers, handheld or portable devices, tablet devices, multiprocessor systems, microprocessor-based systems, set-top boxes, programmable consumer electronics, network PCs, minicomputers, mainframe computers, and distributed computing environments including any of the above systems or devices. This application can be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc., that perform specific tasks or implement specific abstract data types. This application can also be practiced in distributed computing environments where tasks are performed by remote processing devices connected via a communication network. In distributed computing environments, program modules can reside in local and remote computer storage media, including storage devices.

[0087] Figure 1 shows a flowchart of a nanopore sequencing-based sequence processing method provided in this disclosure. This method can be applied to a nanopore sequencing-based sequence processing device, which can be integrated into a computer device, which can be a terminal or a server. The nanopore sequencing-based sequence processing method may include:

[0088] Step 110: Obtain the raw sequencing sequence; wherein the raw sequencing sequence is obtained through nanopore sequencing technology and the raw sequencing sequence meets the preset sequencing sequence requirements parameters;

[0089] It should be noted that multiple candidate sequencing sequences are obtained by sequencing multiple single-molecule genes using nanopore sequencing technology. These candidate sequencing sequences are stored in a sequencing database in a preset format. When sequencing sequence processing is required, the candidate sequencing sequences are extracted from the sequencing database and preprocessed to obtain the raw sequencing sequences, ensuring that the raw sequencing sequences are correctly analyzed in subsequent steps. In this embodiment, the preset format is FASTQ format, which is a text-based file format for storing biological sequences and their corresponding base qualities. In other embodiments, the preset format can be FASTA, SAM, or VCF; this embodiment does not impose specific limitations on the preset format.

[0090] To further explain, when extracting the original sequencing sequence, the extraction conditions are first set as preset information. Based on the set conditions, N candidate sequencing sequences (e.g., the first 100 long reads) that meet the length requirements are extracted as the original sequencing sequence. The proportion of special sequences in the genome can be accurately calculated through the original sequencing sequence, and then the detection threshold can be estimated.

[0091] In some embodiments, obtaining the raw sequencing sequence includes:

[0092] Obtain candidate sequencing sequences and sequencing sequence requirement parameters; among which, the sequencing sequence requirement parameters include: reference length requirement information and sequence quantity requirement information;

[0093] Candidate sequencing sequences are screened based on reference length requirements, and those sequences that meet the length requirements are selected as sequencing sequences.

[0094] The selected sequencing sequences are filtered based on the sequence quantity requirements to obtain the original sequencing sequences.

[0095] It should be noted that the sequencing sequence requirement parameters, used as selection parameters for candidate sequencing sequences, are employed to obtain raw sequencing sequences that meet these requirements. These parameters include: reference length requirement information and sequence quantity requirement information. The reference length requirement information represents the length requirement for selecting candidate sequencing sequences, and the sequence quantity requirement information represents the quantity requirement for selecting candidate sequencing sequences. The candidate reference length information represents the length of the candidate sequencing sequences. Therefore, in the candidate sequencing sequence selection process, sequences whose length information matches the reference length requirement information are first selected from the candidate sequencing sequences. Then, raw sequencing sequences with a matching quantity are extracted from the selected sequences according to the sequence quantity requirement information.

[0096] Specifically, this embodiment uses an *E. coli* sample as an example, sequencing the *E. coli* sample using a nanopore sequencer to obtain candidate sequencing sequences. As shown in Figure 2, quality control revealed that the candidate sequencing sequences exhibited base separation at the tail positions of several hundred to several thousand bp. Furthermore, referring to Figures 3, 4, and 5, further analysis revealed a large number of homopolymers and dimers at the tail of the candidate sequencing sequences, while the homopolymers and dimers were found to be relatively evenly distributed in the middle region of the candidate sequencing sequences. Therefore, the candidate sequencing sequences obtained directly by the nanopore sequencer need to be processed to obtain more accurate sequencing sequences. For example, if the reference length requirement is 6kb and the sequence quantity requirement is 100, then 100 raw sequencing sequences with a length of 6kb are extracted from the sequence database and stored in a list for subsequent analysis. It should be noted that raw sequencing sequences longer than 6kb are considered long sequences, and long sequences have fewer sequencing errors in the middle region, which can be used to estimate the distribution of homopolymers and dimers in the normal region.

[0097] In this embodiment of the disclosure, original sequencing sequences that meet the preset requirements in both length and quantity are selected from the candidate sequencing sequences in order to facilitate the correction of anomalies in the original sequencing sequences.

[0098] Step 120: Calculate the proportion of special sequences in the middle region of the original sequencing sequence to obtain the first proportion information.

[0099] It should be noted that regular expressions are a string matching algorithm that uses special characters and metacharacters to match specific regions in a sequence. As previously disclosed, homopolymers and dimers can exist in candidate sequencing sequences, leading to low accuracy. Therefore, the regular expression matching mode in this embodiment mainly identifies homopolymers and / or dimers in the original sequencing sequence and calculates their proportion within the analysis window to determine the proportion of special sequences. It should be noted that when analyzing the original sequencing sequence, because the original sequence is very long, calculating special sequences across the entire original sequence requires significant computational resources. Therefore, by determining the analysis sequences within the original sequencing sequence according to the analysis window, and then calculating the sequences corresponding to homopolymers and dimers in the analysis sequence, the proportion of special sequences in the analysis sequence can be calculated. Finally, the proportion of special sequences in the entire original sequencing sequence can be derived, saving computational resources previously used for calculating the proportion of special sequences.

[0100] Specifically, the first proportion information is the proportion of special sequences in the middle region of the original sequencing sequence, specifically the proportion of homopolymers and / or dimers.

[0101] In some embodiments, the proportion of special sequences in the middle region of the original sequencing sequence is statistically analyzed to obtain first proportion information, including:

[0102] The reference region is extracted from the original sequencing sequence to obtain the reference sequencing sequence; the reference region indicates the region located in the middle of the original sequencing sequence.

[0103] Based on the regular expression matching pattern, special sequences within the reference sequencing sequence are extracted to obtain the selected sequencing sequence.

[0104] The ratio of the sequence lengths of the selected sequencing sequence to the reference sequencing sequence is obtained as the first proportion information.

[0105] As disclosed above, this embodiment uses the sequence corresponding to the middle region of the original sequencing sequence as the reference sequencing sequence. It should be noted that homopolymers and / or dimers exist in the genome. For third-generation sequencing sequences, the middle regions of longer sequencing sequences have fewer errors, and the proportion of special sequences in these regions can be used to estimate the true proportion of special sequences in the entire genome. Specifically, sequencing errors are prone to occur in the tail region of the sequencing sequence. Therefore, this embodiment uses the sequence corresponding to the middle region of the original sequencing sequence as the reference sequencing sequence so that the calculated first proportion information can characterize the proportion of special sequences in the entire original sequencing sequence, improving the accuracy of the special sequence proportion calculation.

[0106] In some embodiments, after selecting the reference sequencing sequence, special sequences are extracted from the reference sequencing sequence according to a regular expression matching pattern. That is, sequences corresponding to homopolymers and / or dimers in the reference sequencing sequence are extracted to obtain the selected sequencing sequence. The length of the selected sequencing sequence is then determined as the length of the special sequence. As disclosed above, calculating the proportion of special sequences requires calculating the length ratio of special sequences within the analysis window. Since the analysis window is the region corresponding to the reference sequencing sequence, obtaining the length of the reference sequencing sequence yields the reference length. The ratio between the length of the special sequence and the reference length is used as the first proportion information, which can accurately characterize the proportion of special sequences in the original sequencing sequence. This is beneficial for setting the detection threshold based on the first proportion information, which serves as the detection threshold for adjusting the candidate sequencing sequence pruning window.

[0107] For example, if the reference length of the reference sequencing sequence is 2000 bp, and the length of the special sequence detected in the original sequencing sequence is 400 bp, then the proportion of the special sequence in this reference sequencing sequence is determined to be 400 / 2000 = 0.2.

[0108] In this embodiment of the disclosure, by using the sequence in the middle region of the original sequencing sequence as a reference sequencing sequence, and calculating the first proportion information of the original sequencing sequence using the reference sequencing sequence, more accurate special sequence proportion information is obtained.

[0109] The selected sequencing sequences must include at least one of the following: homopolymers or dimers.

[0110] In some embodiments, the selected sequencing sequence is obtained by extracting special sequences from the reference sequencing sequence according to a preset regular expression matching pattern, including at least one of the following steps:

[0111] Homopolymers within the reference sequencing sequence are extracted based on regular expression matching patterns;

[0112] Dimers within the reference sequencing sequence are extracted based on regular expression matching patterns.

[0113] As previously disclosed, extracting abnormal sequences from the reference sequencing sequence primarily involves extracting homopolymers and / or dimers. It should be noted that the regular expression matching patterns include homopolymer matching and dimer matching patterns. Homopolymers from the reference sequencing sequence are extracted according to the homopolymer matching pattern. Specifically, homopolymer matching rules are determined based on the homopolymer matching pattern, and sequences conforming to these rules are extracted from the reference sequencing sequence as homopolymers. Simultaneously, dimer matching rules are determined based on the dimer matching pattern, and sequences conforming to these rules are extracted from the reference sequencing sequence as dimers. This identifies high-error-rate regions in the reference sequencing sequence, especially at the end of long reads where sequencing end signals are difficult to identify. Using both homopolymer and dimer matching patterns to identify the reference sequencing sequence effectively captures special sequences within the reference sequencing sequence.

[0114] In this embodiment of the disclosure, a regular expression matching pattern is designed for homopolymer and dimer sequences to identify special sequences in the original sequencing sequence, thereby achieving more accurate identification of special sequences and effectively identifying and pruning error regions caused by inherent defects in nanopore sequencing.

[0115] Step 130: Determine the detection threshold based on the first proportion information.

[0116] As previously disclosed, the detection threshold serves as the threshold for adjusting the trimming window during sequencing sequence trimming, determining the specific adjustment operations for the trimming window to ensure higher accuracy of the trimmed sequencing sequence. It should be noted that the detection threshold, as a threshold for detecting special sequences, serves as the standard for subsequent target sequencing sequence detection and screening, and the detection threshold is continuously adjusted during sequence detection to extract accurate and complete special sequences.

[0117] In some embodiments, determining the detection threshold based on the first proportion information includes:

[0118] The preset Bayesian prior distribution parameters are updated based on the first proportion information to obtain the updated posterior distribution parameters.

[0119] Based on the posterior distribution parameters, a posterior distribution is constructed to determine the detection threshold.

[0120] It should be noted that after determining the initial proportion information, the posterior distribution of the special sequence proportion is continuously adjusted using a Bayesian update method to dynamically determine the detection threshold, making the anomaly detection of subsequent sequencing sequences more flexible and accurate. Therefore, the pre-set Bayesian prior distribution parameters are updated using the initial proportion information to obtain updated posterior distribution parameters. A posterior distribution is then constructed based on these parameters, and the 99th quantile of the posterior distribution is extracted as the detection threshold. Specifically, the prior beta distribution is updated based on the initial proportion information, i.e., the α and β parameters in the beta distribution are updated. It should be noted that α and β parameters represent the proportion of special sequences and the proportion of normal regions, respectively. Then, the detection threshold is determined based on the 99th quantile of the beta distribution.

[0121] For example, if we set the α and β parameters in the prior beta distribution, i.e., determine the initial Bayesian prior parameters, and the Bayesian prior parameters are estimated based on the empirical ratio of special sequences and normal regions, for example, α=1, β=2. The parameters of the beta distribution are updated according to the first proportion information to achieve the update of the Bayesian posterior parameters, thus obtaining the updated posterior distribution parameters. Finally, the upper limit of the proportion of special sequences is determined by calculating the 99th quantile of the beta distribution, thereby determining the detection threshold.

[0122] In this embodiment of the disclosure, the Bayesian posterior distribution parameters are updated by updating the first proportion information to obtain the updated posterior distribution parameters, and then the detection threshold is determined based on the updated posterior distribution parameters. The detection threshold is then used as the standard for subsequent detection and screening of special sequences.

[0123] It should be noted that after setting the detection threshold, nanopore sequencing technology is used again to sequence the target single gene molecule to obtain the updated original sequencing sequence. It should also be noted that after trimming the original sequencing sequence, it is necessary to ensure that the modified target sequencing sequence meets the minimum quality requirements. This is because the quality value set for the original sequencing sequence corresponds to the probability of error for each base in the original sequencing sequence.

[0124] It should be further noted that the updated original sequencing sequences and their quality values ​​will be written into a new FASTQ file to ensure sequence integrity and traceability. Furthermore, the original read length information of the updated original sequencing sequences is preserved in the FASTQ file output for subsequent analysis and validation.

[0125] Step 140: Based on the preset tail detection window, calculate the proportion of special sequences in the partial sequencing sequences of the original sequencing sequence within the tail detection window to obtain the second proportion information.

[0126] It should be noted that the tail detection window serves as both a detection window for calculating the proportion of special sequences in the original sequencing sequence and a trimming window for the original sequencing sequence. For example, in this embodiment, the tail detection window is set to the tail 1kb, meaning that the original sequencing sequence is read from the beginning, and the tail 1kb of the original sequencing sequence is extracted as a reference sequence for calculating the proportion of special sequences, thereby determining the proportion of special sequences in the reference sequence and making the calculation of the proportion of special sequences in the original sequencing sequence more accurate.

[0127] In some embodiments, based on a preset tail detection window, the proportion of special sequences in the portion of the original sequencing sequence within the tail detection window is statistically analyzed to obtain second proportion information, including:

[0128] Based on the tail detection window, obtain the tail sequencing sequence of the tail region of the original sequencing sequence;

[0129] Based on the regular expression matching pattern, special sequences are extracted from the tail sequencing sequence to obtain the selected sequencing sequence.

[0130] Obtain the length of the selected sequencing sequence;

[0131] Obtain the ratio between the length of the selected sequence and the window length of the tail detection window to get the second proportion information.

[0132] It should be noted that the original sequencing sequence is first divided into tail sequencing sequences according to the tail detection window, and the length of the tail sequencing sequences changes synchronously as the tail detection window is continuously adjusted. It should be further noted that the process of extracting special sequences from the tail sequencing sequences based on regular expression matching patterns is similar to the process of extracting special sequences from the reference sequencing sequence, and will not be elaborated here. The length of the selected sequence is obtained by determining the length of the special sequences in the original sequencing sequence based on the selected sequencing sequence. The window length is obtained by acquiring the length of the tail detection window, and the ratio between the selected sequence length and the window length is used to determine the second proportion information, that is, to determine the proportion of tail special sequences in the original sequencing sequence.

[0133] In this embodiment of the disclosure, as disclosed above, the presence of special sequences at the tail of the sequencing sequence is quite obvious, and the pruning operation mainly targets these special sequences. Therefore, by extracting the tail sequencing sequence from the original sequencing sequence using a tail detection window and calculating the proportion of special sequences based on the tail sequencing sequence, the special sequences that need to be pruned in the original sequencing sequence can be detected more accurately. Thus, pruning the original sequencing sequence based on the second proportion information is more accurate.

[0134] Step 150: The original sequencing sequence is trimmed according to the second proportion information, detection threshold and tail detection window to obtain the target sequencing sequence.

[0135] As previously disclosed, after determining the second proportion information and the detection threshold, the system uses this information and threshold to determine whether the original sequencing sequence needs to be pruned and to set the pruning window. It should be noted that adjusting the size of the tail detection window is equivalent to adjusting the pruning window of the original sequencing sequence. Therefore, the adjusted pruning window allows the original sequencing sequence to be pruned into an accurate and high-quality target sequencing sequence, reducing the problems of over-pruning or incomplete pruning.

[0136] In some embodiments, the original sequencing sequence is pruned according to the second proportion information, the detection threshold, and the tail detection window to obtain the target sequencing sequence, including:

[0137] Update the tail detection window based on the second proportion information and the detection threshold.

[0138] The detection threshold is updated by adjusting the ratio according to the preset threshold.

[0139] The final tail detection window is determined based on the updated tail detection window and detection threshold.

[0140] The tail region of the original sequencing sequence is pruned based on the final tail detection window to obtain the target sequencing sequence.

[0141] Specifically, comparison information is obtained by comparing the second proportion information of the tail detection window with the detection threshold. This comparison information represents the result of comparing the second proportion information of the tail detection window with the detection threshold. If the comparison information indicates that the second proportion information of the tail detection window is less than or equal to the detection threshold, it indicates that the proportion of special sequences in the tail sequencing sequence is relatively small, indirectly reflecting that the tail detection window is set too large. Conversely, if the comparison information indicates that the second proportion information of the tail detection window is greater than the detection threshold, it indicates that the proportion of special sequences in the tail sequencing sequence is relatively large, indirectly reflecting that the tail detection window is set too small.

[0142] As disclosed above, when pruning the original sequencing sequence, the size of the selected detection window affects the pruning effect of the original sequencing sequence. In order to improve the quality of the pruned sequencing sequence to the best, the tail detection window is adjusted to obtain a reasonable and accurate candidate tail detection window.

[0143] It should be noted that the detection threshold is also adjusted simultaneously during the process of adjusting the tail detection window.

[0144] In some embodiments, adjusting the tail detection window based on comparison information to obtain a candidate tail detection window includes:

[0145] If the second proportion information is greater than the detection threshold, the tail detection window is enlarged according to the preset enlargement ratio;

[0146] If the second proportion information is less than or equal to the detection threshold, the tail detection window is shrunk according to the preset shrinkage ratio.

[0147] As previously disclosed, referring to Figure 13, the tail detection window setting is determined in advance based on the comparison result of the second proportion information and the detection threshold. If the second proportion information is greater than the detection threshold, it is determined that the tail detection window setting is too small, and the tail detection window needs to be gradually expanded according to the expansion ratio. Then, the proportion information of special sequences in the sequencing sequence after being pruned with the tail detection window is repeatedly calculated, and the detection threshold is updated synchronously until the length of the sequencing sequence is covered, or the updated second proportion information is equal to the corresponding updated detection threshold. Specifically, in this embodiment, the tail detection window is expanded according to the expansion ratio. If the expansion ratio is 1, the tail detection window is expanded by 1 time. Then, the original sequencing sequence is pruned according to the updated tail detection window to obtain the updated sequencing sequence. The adjustment operation of the tail detection window is determined by comparing the proportion information of special sequences in the updated sequencing sequence with the corresponding updated detection threshold. It should be further explained that if further adjustments are needed, the tail detection window should be expanded to obtain the final tail detection window. Ensure that there are obvious special sequences in the tail sequencing sequences extracted based on the final tail detection window, and that these special sequences can be effectively identified and pruned to reduce the possibility of incomplete pruning of special sequences.

[0148] It should be noted that, referring to Figure 13, if the second proportion information is less than the detection threshold, it indicates that the tail detection window is set too large. If the original sequencing sequence is pruned according to the tail detection window, over-pruning may occur. Therefore, the tail detection window is reduced according to a reduction ratio, which can be set to a fixed ratio or dynamically. For example, if the reduction ratio is dynamically set, the size of the candidate tail detection window can be 200bp, 600bp, or 1kb, etc. The tail region of the original sequencing sequence is pruned according to the updated candidate tail detection window to obtain the updated sequencing sequence. The second proportion information is updated based on the updated sequencing sequence, and the detection threshold is updated according to the threshold adjustment ratio. The updated detection threshold and the updated second proportion information are compared, and the tail detection window is updated again until the second proportion information after multiple updates is equal to or slightly greater than the updated detection threshold. Therefore, by continuously updating the tail detection window with the continuously updated detection threshold and second proportion information, the final tail detection window is obtained, which can achieve precise location of specific sequences in the original sequencing sequence.

[0149] In some embodiments, determining the final tail detection window based on the updated tail detection window and the detection threshold includes:

[0150] Repeat the steps of updating the tail detection window based on the second proportion information and the detection threshold, and adjusting the proportion of the detection threshold according to the preset threshold, until the second proportion information of the updated tail detection window is not greater than the updated detection threshold, and obtain the final tail detection window.

[0151] It should be noted that the adjustment operation of the tail detection window is as disclosed above and will not be repeated here. Therefore, this embodiment determines the final tail detection window by continuously adjusting the tail detection window and the detection threshold until the second proportion information of the updated tail detection window is not greater than the updated detection threshold. The original sequencing sequence is then pruned based on the final tail detection window, which can reduce over-pruning and obtain sequencing sequences with better accuracy and quality.

[0152] In this embodiment, the tail detection window is dynamically adjusted based on the comparison information between the second proportion information and the detection threshold. Then, the updated sequencing sequence is trimmed according to the updated tail detection window. The tail detection window is further adjusted based on the second proportion information of the updated sequencing sequence and the updated detection threshold to obtain the final tail detection window. The adjusted final tail detection window can trim a more accurate target sequencing sequence, thereby improving the sequence trimming effect.

[0153] Specifically, while adjusting the tail detection window, the detection threshold is adjusted simultaneously, and the threshold adjustment is based on the threshold adjustment ratio. For example, if the threshold adjustment ratio is 0.1, when adjusting the size of the tail detection window, whether the tail detection window is enlarged or reduced, the detection threshold will increase by 0.1, which means that the proportion of special sequences in the tail detection window is required to increase by 10%.

[0154] In some embodiments, the final tail detection window can serve as a reference window for dividing the tail of the target sequencing sequence. Therefore, the tail region in the target sequencing sequence is divided into target tail sequences according to the final tail detection window, and the target tail sequences are special sequences caused by homopolymers or dimers. Thus, the original sequencing sequence is pruned based on the target tail sequences, that is, the target tail sequences are removed from the original sequencing sequence to obtain the target sequencing sequence.

[0155] For example, after processing the original sequencing sequence of the E. coli sample, the target sequencing sequence of the E. coli sample is shown in Figure 6. As can be seen from Figure 6, the base separation at the tail of the target sequencing sequence has basically disappeared. In addition, as shown in Figures 7, 8, and 9, the homopolymer and dimer sequences are relatively evenly distributed, without any sudden peaks at the tail. Furthermore, Figures 10(a) and 10(b) are base quality comparison diagrams of the original sequencing sequence and the target sequencing sequence. As can be seen from Figures 10(a) and 10(b), after the above processing, the overall quality of the tail of the target sequencing sequence is improved.

[0156] As previously disclosed, trimming the original sequencing sequence using the final tail detection window enables effective sequence trimming, resulting in a target sequencing sequence with higher quality. Specifically, the low-quality regions at the tail of the original sequencing sequence are trimmed according to the final tail detection window. The trimmed sequencing sequence removes low-quality fragments caused by homopolymer or dimer errors, yielding a target sequencing sequence of superior quality.

[0157] It should be noted that the updated quality value of the target sequencing sequence is calculated based on the error estimation probability of each base in the target sequencing sequence, and the target sequencing sequence and the updated quality value are saved to a new FASTQ file so that downstream applications can select the target sequencing sequence based on the updated quality value.

[0158] In some embodiments, the sequence processing method based on nanopore sequencing is implemented simultaneously by multiple CPUs, so that different sequencing sequences can be pruned by processes corresponding to multiple CPUs, thereby improving the speed of sequencing sequence processing.

[0159] In some embodiments, after step 160, the sequence processing method based on nanopore sequencing further includes:

[0160] Request a processing instruction for the target sequencing sequence;

[0161] The target sequencing sequence is evaluated for quality according to the processing instructions, and a quality evaluation value is obtained.

[0162] Based on the processing instructions, the target downstream operation is selected from the preset candidate downstream operations; the target downstream operations include genome assembly operations and variant detection operations.

[0163] If the quality assessment value is greater than the preset quality threshold, then downstream operations are performed on the target single gene molecule based on the target sequencing sequence.

[0164] It should be noted that, in order to achieve high-quality downstream operations, i.e., high-quality application of the target sequencing sequence, it is necessary to first perform a quality assessment on the target sequencing sequence to obtain a quality assessment value. As disclosed above, the target sequencing sequence and the updated quality value are stored together in the FASTQ file, so the updated quality value is directly extracted as the quality assessment value to determine the quality of the target sequencing sequence. First, target downstream operations matching the processing instruction request are screened from the candidate downstream operations. If the processing instruction request is characterized as genome assembly, then the target downstream operation is a genome assembly operation; if the processing instruction request is characterized as variant detection, then the target downstream operation is a variant detection operation. Therefore, the matching target downstream operation is determined according to different processing instruction requests, i.e., the application scenario of the target sequencing sequence is determined, and the target sequencing sequence is processed accordingly according to the target downstream operation. Specifically, the target downstream operation in this embodiment is not limited to genome assembly and variant detection operations; other operations can be set as needed, and no restrictions are placed here.

[0165] It should be further explained that the target sequencing sequence is obtained through the above-mentioned special sequence detection and pruning. The target sequencing sequence not only effectively improves the quality of the sequencing sequence, but also provides a reliable data foundation for the application of the sequencing sequence. It has broad application prospects and great significance in fields such as genome assembly and variant detection.

[0166] In summary, this application's embodiments determine the detection threshold based on the proportion of special sequences in the pre-collected raw sequencing sequences. This detection threshold is then used as the detection threshold for aligning the proportion of special sequences in the raw sequencing sequences, thereby determining the size of the trimming window in the target sequencing sequence and achieving precise trimming of the raw sequencing sequence to obtain a higher-quality target sequencing sequence. Therefore, this application's embodiments do not rely on any reference genome or known sequence databases. By directly analyzing the sequence features in the sequencing sequences, it independently completes anomaly detection and trimming, making it suitable for research on new species genomes or unknown genomes. It also ensures high-quality sequencing sequences and improves the accuracy of downstream applications.

[0167] Please refer to Figure 11. This application also provides a sequence processing device based on nanopore sequencing, which can implement the above-described sequence processing method based on nanopore sequencing. The device includes:

[0168] The raw sequence acquisition module is used to acquire raw sequencing sequences; the raw sequencing sequences are acquired through nanopore sequencing technology and meet the preset sequencing sequence requirements parameters.

[0169] The first statistical module is used to calculate the proportion of special sequences in the middle region of the original sequencing sequence to obtain the first proportion information.

[0170] The threshold setting module is used to determine the detection threshold based on the first proportion information;

[0171] The second statistical module is used to calculate the proportion of special sequences in the partial sequencing sequence within the tail detection window of the original sequencing sequence according to the preset tail detection window, and obtain the second proportion information.

[0172] The sequence trimming module is used to trim the original sequencing sequence according to the second proportion information, the detection threshold, and the tail detection window to obtain the target sequencing sequence.

[0173] The specific implementation of this nanopore sequencing-based sequence processing device is basically the same as the specific embodiment of the nanopore sequencing-based sequence processing method described above, and will not be repeated here.

[0174] This application also provides an electronic device, which includes a memory and a processor. The memory stores a computer program, and the processor executes the computer program to implement the above-described sequence processing method based on nanopore sequencing. This electronic device can be any smart terminal, including tablet computers, in-vehicle computers, etc.

[0175] Please refer to Figure 12, which illustrates the hardware structure of an electronic device according to another embodiment. The electronic device includes:

[0176] The processor 1201 can be implemented using a general-purpose CPU (Central Processing Unit), microprocessor, application-specific integrated circuit (ASIC), or one or more integrated circuits, and is used to execute relevant programs to implement the technical solutions provided in the embodiments of this application.

[0177] The memory 1202 can be implemented as a read-only memory (ROM), a static storage device, a dynamic storage device, or a random access memory (RAM). The memory 1202 can store the operating system and other applications. When the technical solutions provided in the embodiments of this specification are implemented through software or firmware, the relevant program code is stored in the memory 1202 and is called and executed by the processor 1201 to execute the sequence processing method based on nanopore sequencing of the embodiments of this application.

[0178] The input / output interface 1203 is used to implement information input and output;

[0179] The communication interface 1204 is used to enable communication and interaction between this device and other devices. Communication can be achieved through wired means (such as USB, network cable, etc.) or wireless means (such as mobile network, WIFI, Bluetooth, etc.).

[0180] Bus 1205 transmits information between various components of the device (e.g., processor 1201, memory 1202, input / output interface 1203, and communication interface 1204);

[0181] The processor 1201, memory 1202, input / output interface 1203 and communication interface 1204 are connected to each other within the device via bus 1205.

[0182] This application also provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the above-described sequence processing method based on nanopore sequencing.

[0183] Memory, as a non-transitory computer-readable storage medium, can be used to store non-transitory software programs and non-transitory computer-executable programs. Furthermore, memory may include high-speed random access memory, and may also include non-transitory memory, such as at least one disk storage device, flash memory device, or other non-transitory solid-state storage device. In some embodiments, memory may optionally include memory remotely located relative to the processor, and these remote memories can be connected to the processor via a network. Examples of such networks include, but are not limited to, the Internet, intranets, local area networks, mobile communication networks, and combinations thereof.

[0184] The nanopore sequencing-based sequence processing method, apparatus, device, and storage medium provided in this application determine the proportion of special sequences in the original sequencing sequence and then determine a detection threshold based on this proportion. When trimming the target sequencing sequence, the proportion of special sequences in the original sequencing sequence is further determined based on the tail detection window. The tail detection window is adjusted based on the proportion of special sequences and the detection threshold, with the adjustment process continuously updating the detection threshold and the tail detection window until the proportion of special sequences does not exceed the updated detection threshold, thus determining the final tail detection window. Therefore, by completing the trimming operation of the target sequencing sequence through the final tail detection window, a target sequencing sequence with higher accuracy and quality can be obtained. Thus, in the process of special sequence detection and trimming of sequencing sequences, there is no need to rely on a reference genome or database, making it suitable for the analysis of genes from new species or genome data lacking comprehensive annotation. Furthermore, the adjustment of the tail detection window reduces the difficulty of identifying tail homopolymers and dimers, reduces the number of special sequences corresponding to tail homopolymers and dimers, and improves the accuracy of the sequencing sequence.

[0185] The embodiments described in this application are for the purpose of more clearly illustrating the technical solutions of the embodiments of this application, and do not constitute a limitation on the technical solutions provided by the embodiments of this application. As those skilled in the art will know, with the evolution of technology and the emergence of new application scenarios, the technical solutions provided by the embodiments of this application are also applicable to similar technical problems.

[0186] Those skilled in the art will understand that the technical solutions shown in the figures do not constitute a limitation on the embodiments of this application, and may include more or fewer steps than shown, or combine certain steps, or different steps.

[0187] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs.

[0188] Those skilled in the art will understand that all or some of the steps in the methods disclosed above, as well as the functional modules / units in the systems and devices, can be implemented as software, firmware, hardware, or suitable combinations thereof.

[0189] The terms “first,” “second,” “third,” “fourth,” etc. (if present) in the specification and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this application described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms “comprising” and “having,” and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.

[0190] It should be understood that in this application, "at least one (item)" means one or more, and "more than" means two or more. "And / or" is used to describe the relationship between related objects, indicating that three relationships can exist. For example, "A and / or B" can represent three cases: only A exists, only B exists, and both A and B exist simultaneously, where A and B can be singular or plural. The character " / " generally indicates that the preceding and following related objects are in an "or" relationship. "At least one (item) of the following" or similar expressions refer to any combination of these items, including any combination of single or plural items. For example, at least one (item) of a, b, or c can represent: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, and c can be single or multiple.

[0191] In the several embodiments provided in this application, it should be understood that the disclosed apparatus and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of the units described above is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between apparatuses or units may be electrical, mechanical, or other forms.

[0192] The units described above as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0193] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.

[0194] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to related technologies, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes multiple instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods of the various embodiments of this application. The aforementioned storage medium includes various media capable of storing programs, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0195] The preferred embodiments of the present application have been described above with reference to the accompanying drawings, but this does not limit the scope of the claims of the present application. Any modifications, equivalent substitutions, and improvements made by those skilled in the art without departing from the scope and substance of the embodiments of the present application shall be within the scope of the claims of the present application.

Claims

1. A sequence processing method based on nanopore sequencing, characterized in that, The method includes: Obtain the raw sequencing sequence; wherein the raw sequencing sequence is obtained through nanopore sequencing technology, and the raw sequencing sequence meets the preset sequencing sequence requirement parameters; The proportion of special sequences in the middle region of the original sequencing sequence is statistically analyzed to obtain the first proportion information; The detection threshold is determined based on the first proportion information; Based on the preset tail detection window, the proportion of special sequences in the portion of the original sequencing sequence within the tail detection window is statistically analyzed to obtain the second proportion information; The original sequencing sequence is trimmed based on the second proportion information, the detection threshold, and the tail detection window to obtain the target sequencing sequence.

2. The method according to claim 1, characterized in that, The percentage of special sequences in the middle region of the original sequencing sequence is statistically analyzed to obtain the first percentage information, including: A reference region sequence is extracted from the original sequencing sequence to obtain a reference sequencing sequence; wherein the reference region indicates a region located in the middle of the original sequencing sequence; Based on the regular expression matching pattern, special sequences within the reference sequencing sequence are extracted to obtain the selected sequencing sequence. The ratio of the sequence lengths of the selected sequencing sequence to the reference sequencing sequence is obtained as the first proportion information.

3. The method according to claim 2, characterized in that, The selected sequencing sequence includes at least one of the following: homopolymer or dimer; the step of extracting special sequences from the reference sequencing sequence to obtain the selected sequencing sequence according to a preset regular expression matching pattern includes at least one of the following steps: Homopolymers within the reference sequencing sequence are extracted based on the regular expression matching pattern. Dimers within the reference sequencing sequence are extracted based on the regular expression matching pattern.

4. The method according to claim 1, characterized in that, The step of determining the detection threshold based on the first proportion information includes: The preset Bayesian prior distribution parameters are updated based on the first proportion information to obtain the updated posterior distribution parameters. Based on the posterior distribution parameters, a posterior distribution is constructed to determine the detection threshold.

5. The method according to any one of claims 1 to 4, characterized in that, The step involves calculating the proportion of special sequences in a portion of the original sequencing sequence within a preset tail detection window to obtain second proportion information, including: Based on the tail detection window, obtain the tail sequencing sequence of the tail region of the original sequencing sequence; Based on the regular expression matching pattern, special sequences are extracted from the tail sequencing sequence to obtain the selected sequencing sequence. Obtain the length of the selected sequencing sequence to get the selected sequence length; The ratio between the selected sequence length and the window length of the tail detection window is obtained to obtain the second proportion information.

6. The method according to any one of claims 1 to 4, characterized in that, The step of trimming the original sequencing sequence according to the second proportion information, the detection threshold, and the tail detection window to obtain the target sequencing sequence includes: The tail detection window is updated based on the second proportion information of the tail detection window and the detection threshold. The detection threshold is updated according to the preset threshold adjustment ratio; The final tail detection window is determined based on the updated tail detection window and detection threshold. The original sequencing sequence is pruned at the tail region based on the final tail detection window to obtain the target sequencing sequence.

7. The method according to claim 6, characterized in that, The step of updating the tail detection window based on the second proportion information of the tail detection window and the detection threshold includes: If the second proportion information is greater than the detection threshold, the tail detection window is enlarged according to the preset enlargement ratio; If the second percentage information is less than or equal to the detection threshold, the tail detection window is reduced according to the preset reduction ratio.

8. The method according to claim 7, characterized in that, The step of determining the final tail detection window based on the updated tail detection window and detection threshold includes: Repeat the steps of updating the tail detection window based on the second proportion information of the tail detection window and the detection threshold, and adjusting the proportion of the detection threshold according to a preset threshold, until the second proportion information of the updated tail detection window is not greater than the updated detection threshold, to obtain the final tail detection window.

9. The method according to any one of claims 1 to 4, characterized in that, The process of obtaining the raw sequencing sequence includes: Obtain candidate sequencing sequences and sequencing sequence requirement parameters; wherein, the sequencing sequence requirement parameters include: reference length requirement information and sequence quantity requirement information; The candidate sequencing sequences are screened according to the reference length requirement information, and the candidate sequencing sequences whose sequence length meets the requirement are selected as sequencing sequences. The selected sequencing sequences are filtered according to the sequence quantity requirement information to obtain the original sequencing sequences.

10. The method according to any one of claims 1 to 4, characterized in that, After the original sequencing sequence is pruned according to the second proportion information, the detection threshold, and the tail detection window to obtain the target sequencing sequence, the method further includes: Obtain the processing instruction request for the target sequencing sequence; The target sequencing sequence is requested to be quality evaluated according to the processing instruction, and a quality evaluation value is obtained; According to the processing instruction, a target downstream operation is selected from a preset pool of candidate downstream operations; wherein, the target downstream operation includes genome assembly operation and variant detection operation; If the quality assessment value is greater than the preset quality threshold, then the target downstream operation is performed on the target single gene molecule according to the target sequencing sequence.

11. A sequence processing device based on nanopore sequencing, characterized in that, The device includes: A raw sequence acquisition module is used to acquire raw sequencing sequences; wherein, the raw sequencing sequences are acquired through nanopore sequencing technology, and the raw sequencing sequences meet preset sequencing sequence requirement parameters; The first statistical module is used to calculate the proportion of special sequences in the middle region of the original sequencing sequence to obtain first proportion information. The threshold setting module is used to determine the detection threshold based on the first proportion information; The second statistical module is used to calculate the proportion of special sequences in the portion of the original sequencing sequence within the tail detection window according to the preset tail detection window, and obtain the second proportion information. The sequence trimming module is used to trim the original sequencing sequence according to the second proportion information, the detection threshold, and the tail detection window to obtain the target sequencing sequence.

12. An electronic device, characterized in that, The electronic device includes a memory and a processor, the memory storing a computer program, and the processor executing the computer program to implement the sequence processing method based on nanopore sequencing as described in any one of claims 1 to 10.

13. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, it implements the sequence processing method based on nanopore sequencing as described in any one of claims 1 to 10.