Index identification method, device and electronic equipment for gene sequencing sequence

By extracting index signal data from the electrical signal data of nanopore sequencing time series and calculating the distance score, the index of the gene sequencing sequence is identified, which solves the problem of low index identification accuracy in the existing technology and achieves higher identification accuracy.

CN122290722APending Publication Date: 2026-06-26HANGZHOU HUADA XUFENG TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
HANGZHOU HUADA XUFENG TECHNOLOGY CO LTD
Filing Date
2024-12-18
Publication Date
2026-06-26

AI Technical Summary

Technical Problem

In existing technologies, the accuracy of index identification in gene sequencing sequences is low, especially when the gene sequencing quality is low or the index sequence is incorrect, blurred, or contaminated, the identification results are prone to errors.

Method used

By acquiring multiple index signal templates, extracting index signal data based on nanopore sequencing time-series electrical signal data, and calculating the distance fraction between the index signal data and each index signal template, the target index signal template is determined, thereby identifying the index of the gene sequencing sequence.

Benefits of technology

It improves the accuracy of gene sequencing sequence indexing and reduces the impact of inaccurate gene sequencing results on the identification results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122290722A_ABST
    Figure CN122290722A_ABST
Patent Text Reader

Abstract

This disclosure provides a method, apparatus, and electronic device for indexing and identifying gene sequencing sequences. The method includes: acquiring multiple index signal templates; acquiring nanopore sequencing time-series electrical signal data of the gene sequencing sequence to be identified; extracting index signal data from the nanopore sequencing time-series electrical signal data; calculating the distance fraction between the index signal data and each index signal template; determining a target index signal template from the multiple index signal templates based on the distance fraction; and determining the target index sequence corresponding to the target index signal template as the indexing and identification result of the gene sequencing sequence to be identified. This method can improve the accuracy of indexing and identifying gene sequencing sequences.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the field of bioinformatics, and in particular to a method, apparatus, and electronic device for indexing and identifying gene sequencing sequences. Background Technology

[0002] High-throughput sequencing technology allows for the sequencing of large numbers of gene samples in a short time. When processing multiple gene samples simultaneously, an index—a short DNA sequence—can be introduced into the gene sample to accurately identify sequences from different gene samples from a large pool of sequencing data. Introducing different types of indexes for different gene samples allows for the differentiation of sequencing sequences from different gene samples based on the identification of these indexes within a large pool of sequencing data.

[0003] However, the accuracy of identifying indexes in gene sequencing sequences using related technologies is low. Summary of the Invention

[0004] This disclosure provides a method, apparatus, and electronic device for indexing and identifying gene sequencing sequences, which can improve the accuracy of indexing and identifying gene sequencing sequences.

[0005] According to one aspect of this disclosure, a method for indexing and identifying gene sequencing sequences is provided, comprising:

[0006] Multiple index signal templates are obtained, each index signal template corresponding to an index sequence. The index signal templates are obtained based on the excerpt of nanopore sequencing time-series electrical signal data containing the index.

[0007] Obtain time-series electrical signal data of nanopore sequencing of the gene sequence to be identified;

[0008] Extract index signal data from the time-series electrical signal data of the nanopore sequencing;

[0009] Calculate the distance score between the index signal data and each index signal template;

[0010] Based on the distance score, a target index signal template is determined from multiple index signal templates, and the target index sequence corresponding to the target index signal template is determined as the index recognition result of the sequencing sequence of the gene to be identified.

[0011] According to one aspect of this disclosure, an indexing and identification device for gene sequencing sequences is provided, comprising:

[0012] The first acquisition unit is used to acquire multiple index signal templates, each of which corresponds to an index sequence. The index signal template is obtained based on the excerpt of nanopore sequencing time-series electrical signal data containing the index.

[0013] The second acquisition unit is used to acquire the nanopore sequencing time-series electrical signal data of the gene sequencing sequence to be identified;

[0014] The extraction unit is used to extract index signal data from the nanopore sequencing time-series electrical signal data;

[0015] A calculation unit is used to calculate the distance score between the index signal data and each of the index signal templates;

[0016] The determining unit is used to determine the target index signal template from multiple index signal templates based on the distance score, and to determine the target index sequence corresponding to the target index signal template as the index recognition result of the sequencing sequence of the gene to be identified.

[0017] Optionally, in one embodiment, the computing unit is specifically used for:

[0018] For each of the index signal templates, multiple sub-signal data are obtained from the index signal data;

[0019] Calculate the sub-distance scores between the multiple sub-signal data and the corresponding index signal template;

[0020] The distance score between the index signal data and the corresponding index signal template is determined based on multiple sub-distance scores.

[0021] Optionally, in one embodiment, the computing unit is specifically used for:

[0022] For each index signal template, obtain the corresponding sequence truncation window and sequence truncation step size;

[0023] Starting from a predetermined position in the index signal data, the sequence truncation window is moved according to the sequence truncation step size to obtain multiple sub-signal data.

[0024] Optionally, in one embodiment, the computing unit is specifically used for:

[0025] For each of the index signal templates, obtain the signal sequence length of the index signal template;

[0026] The sequence truncation window and sequence truncation step size are determined based on the length of the signal sequence.

[0027] Optionally, in one embodiment, the computing unit is specifically used for:

[0028] Create a distance matrix, wherein the number of rows in the distance matrix is ​​equal to the length of the sub-signal data and the number of columns is equal to the length of the index signal template, and each element in the distance matrix represents the distance between the sub-signal data and the index signal template at the corresponding position;

[0029] Calculate the signal distance between each signal value in the sub-signal data and each signal value in the index signal template, and fill the distance matrix with the signal distance;

[0030] Calculate the shortest path from the top left corner to the bottom right corner of the distance matrix using dynamic programming algorithm;

[0031] The sum of signal distances along the shortest path is calculated as the sub-distance score between the sub-signal data and the index signal template.

[0032] Optionally, in one embodiment, the gene sequencing sequence indexing and identification device further includes:

[0033] The first compression unit is used to compress multiple index signal templates to obtain multiple compressed index signal templates;

[0034] The second compression unit is used to compress the index signal data to obtain compressed index signal data;

[0035] The computing unit is specifically used for:

[0036] Calculate the distance fraction between the compressed index signal data and each of the compressed index signal templates.

[0037] Optionally, the first compression unit is specifically used for:

[0038] Obtain a compression window of a predetermined width, the compression window being used to slide backward from the first signal value of the index signal template according to a predetermined step size until the last value;

[0039] The signal values ​​of a predetermined width in the compression window are compressed until the compression window slides to the last position, resulting in multiple compressed index signal templates.

[0040] Optionally, the first compression unit is specifically used for:

[0041] Calculate the standard deviation of the signal values ​​within the predetermined width in the compressed window;

[0042] When the standard deviation is less than the compression threshold, the average value of the predetermined width of signal values ​​is used as the compressed value of the predetermined width of signal values ​​in the compression window.

[0043] Optionally, the determining unit is specifically used for:

[0044] Determine the target distance score from the distance scores corresponding to the multiple index signal templates;

[0045] When the target distance score is less than the distance threshold, the index signal template corresponding to the target distance score is determined as the target index signal template, and the target index sequence corresponding to the target index signal template is determined as the index recognition result of the sequencing sequence of the gene to be identified.

[0046] Optionally, the distance threshold is determined in the following way:

[0047] Obtain an index matching set, which stores multiple sequencing electrical signal data and their corresponding index electrical signal data;

[0048] Calculate the matching score between the sequencing electrical signal data and the corresponding index electrical signal data in the index matching set;

[0049] The distance threshold is determined based on multiple matching scores.

[0050] Optionally, the interception unit is specifically used for:

[0051] Determine the truncation position and truncation range length of the index signal data in the nanopore sequencing time-series electrical signal data of the gene sequencing sequence to be identified;

[0052] In the nanopore sequencing time-series electrical signal data of the gene sequencing sequence to be identified, a signal sequence of the specified length at the specified cut position is used as index signal data.

[0053] Optionally, the interception unit is specifically used for:

[0054] Obtain the template truncation position and template truncation length of multiple index signal templates in the corresponding nanopore sequencing time-series electrical signal data;

[0055] Based on the template truncation position and the template truncation length, the truncation position and truncation range length of the index signal data are determined in the nanopore sequencing time series electrical signal data of the sequencing sequence of the gene to be identified.

[0056] According to one aspect of this disclosure, a computer device is provided, including a memory and a processor, the memory storing a computer program, the processor executing the computer program to implement the gene sequencing sequence indexing and identification method as described above.

[0057] According to one aspect of this disclosure, a computer-readable storage medium is provided, the storage medium storing a computer program that, when executed by a processor, implements the gene sequencing sequence indexing and identification method as described above.

[0058] According to one aspect of this disclosure, a computer program product is provided, the computer program product comprising a computer program that is read and executed by a processor of a computer device, causing the computer device to perform the gene sequencing sequence indexing and identification method as described above.

[0059] The gene sequencing sequence indexing identification method in this embodiment of the present disclosure involves obtaining multiple index signal templates, each index signal template corresponding to an index sequence. The index signal templates are obtained by extracting nanopore sequencing time-series electrical signal data containing the index. The method then acquires nanopore sequencing time-series electrical signal data of the gene sequencing sequence to be identified; extracts index signal data from the nanopore sequencing time-series electrical signal data; calculates the distance score between the index signal data and each index signal template; determines the target index signal template from the multiple index signal templates based on the distance score; and identifies the target index sequence corresponding to the target index signal template as the indexing identification result of the gene sequencing sequence to be identified.

[0060] Therefore, in the gene sequencing sequence index identification method of this disclosure, the nanopore sequencing time-series electrical signal data of the gene sequencing sequence is used for index identification. Specifically, an index signal template is extracted from the nanopore sequencing time-series electrical signal data containing the index. The index signal template represents the electrical signal data of the corresponding index sequence. Since the gene sequencing data to be identified contains the introduced index sequence, the nanopore sequencing time-series electrical signal data of the gene sequencing data to be identified also contains the electrical signal data corresponding to the index sequence. The index signal data extracted from the nanopore sequencing time-series electrical signal data of the gene sequencing data to be identified contains the electrical signal data corresponding to the index sequence. The index signal data is matched with multiple index signal templates to find the target index signal template, thereby determining the index sequence in the gene sequencing sequence to be identified that introduces the index. The nanopore sequencing time-series electrical signal data reflects the changes in electrical signals of the sequencing instrument during the gene sequencing process. Even if the sequencing results obtained by the sequencing instrument are inaccurate, the changes in electrical signals can still reflect the gene sequence information relatively accurately. Therefore, using electrical signal data to identify the index in the gene sequencing sequence to be identified can avoid the influence of gene sequencing results on the identification results, which is beneficial to improving the accuracy of index identification.

[0061] Other features and advantages of this disclosure will be set forth in the following description and will be apparent in part from the description or may be learned by practicing the disclosure. The objectives and other advantages of this disclosure may be realized and obtained by means of the structures particularly pointed out in the description, claims and drawings. Attached Figure Description

[0062] The accompanying drawings are provided to further understand the technical solutions of this disclosure and constitute a part of the specification. They are used together with the embodiments of this disclosure to explain the technical solutions of this disclosure and do not constitute a limitation on the technical solutions of this disclosure.

[0063] Figure 1 This is a system architecture diagram of a gene sequencing sequence indexing and identification method applied according to an embodiment of the present disclosure;

[0064] Figure 2 This is a flowchart illustrating a gene sequencing sequence indexing and identification method according to an embodiment of the present disclosure;

[0065] Figure 3 This is a time series plot of electrical signal data corresponding to an index sequence in nanopore sequencing time series electrical signal data, according to one embodiment of the present disclosure;

[0066] Figure 4 It is a time series plot of index signal data extracted from the time series electrical signal data of nanopore sequencing of a gene sequencing sequence to be identified, according to one embodiment of the present disclosure;

[0067] Figure 5 This is a schematic diagram illustrating the determination of the truncation position and truncation range length of index signal data based on multiple index signal templates according to one embodiment of this disclosure;

[0068] Figure 6 This is a schematic diagram of a compression window sliding on an index signal template according to one embodiment of the present disclosure;

[0069] Figure 7 This is a schematic diagram of a sequence truncating window sliding over index signal data according to one embodiment of the present disclosure;

[0070] Figure 8 This is a schematic diagram of the structure of a gene sequencing sequence indexing and identification device according to an embodiment of the present disclosure;

[0071] Figure 9 This is a terminal structure diagram for implementing various methods according to an embodiment of the present disclosure;

[0072] Figure 10 This is a server structure diagram illustrating the implementation of various methods according to an embodiment of the present disclosure. Detailed Implementation

[0073] To make the objectives, technical solutions, and advantages of this disclosure clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and are not intended to limit the scope of this disclosure.

[0074] Before providing a further detailed description of the embodiments of this disclosure, the terms and concepts used in these embodiments are explained, and they are subject to the following interpretations:

[0075] Index (barcode): In high-throughput sequencing, multiple gene samples can be sequenced simultaneously. To distinguish the sequences of multiple gene samples, a short, unique DNA sequence can be inserted into the gene sequence of each sample. This inserted DNA sequence is the index of the gene sequence, which is used to accurately track and identify gene samples during the sequencing process.

[0076] Electrical signal data: The electrical signal data generated during gene sequencing is obtained through nanopore sequencing technology. The core of nanopore sequencing technology is a polymer membrane integrating multiple transmembrane channel proteins (i.e., nanopore proteins). By applying a voltage across the membrane, a stable current is generated that flows through the nanopore. When other objects pass through the nanopore, they affect the magnitude of the current, thus producing a recognizable change in the electrical signal.

[0077] Nanopore sequencing is a DNA sequencing technology within the third-generation sequencing field. It utilizes nanopore proteins to sequence single-stranded DNA in real time. This technology offers advantages such as long read lengths, high speed, and portability, and is widely used in genomics, transcriptomics, and epigenetics. The basic principle of nanopore sequencing is to identify DNA sequences by utilizing changes in electrical signals. When single-stranded DNA passes through a nanopore, different bases (A, T, C, G) cause changes in electrical signals due to their unique physical and chemical properties. By detecting these changes in electrical signals, the passing bases can be identified in real time, thereby enabling the determination of the DNA sequence.

[0078] Dynamic programming is a method for solving optimization problems in multi-stage decision processes. In dynamic programming, the original problem is decomposed into relatively simple subproblems, which are then solved first, and the solution to the original problem is obtained from the solutions to the subproblems.

[0079] In related technologies, index identification is generally based on gene sequencing sequences. However, when the sequencing quality of the gene sequencing sequence is low, or when the index sequence contains errors, is blurred, or is contaminated, index identification may fail. Therefore, the accuracy of index identification in related technologies is low. To address this, this disclosure provides a method for index identification of gene sequencing sequences, aiming to improve the accuracy of index identification.

[0080] The system architecture used in the embodiments of this disclosure

[0081] Figure 1This is a system architecture diagram of a gene sequencing sequence indexing and identification method according to embodiments of the present disclosure. It includes a terminal 140, an Internet 130, a gateway 120, a server 110, etc.

[0082] Terminal 140 includes various forms such as desktop computers, laptops, PDAs (personal digital assistants), mobile phones, in-vehicle terminals, home theater terminals, and dedicated terminals. Furthermore, it can be a single device or a collection of multiple devices. Terminal 140 can communicate with the Internet 130 via wired or wireless means to exchange data.

[0083] Server 110 refers to a computer system that can provide certain services to terminal 140. Compared to ordinary terminal 140, server 110 has higher requirements in terms of stability, security, and performance. Server 110 can be a single high-performance computer in a network platform, a cluster of multiple high-performance computers, a portion of a single high-performance computer (e.g., a virtual machine), or a combination of portions of multiple high-performance computers (e.g., virtual machines).

[0084] Gateway 120, also known as an internetwork connector or protocol converter, is a computer system or device that acts as a translator, enabling network interconnection at the transport layer. It bridges the gap between two systems using different communication protocols, data formats, languages, or even completely different architectures. Gateways can also provide filtering and security functions. Messages sent from terminal 140 to server 110 are forwarded to the corresponding server 110 via gateway 120. Messages sent from server 110 to terminal 140 are also forwarded to the corresponding terminal 140 via gateway 120.

[0085] The gene sequencing sequence indexing and identification method of this disclosure can be implemented entirely on the terminal 140; it can be implemented entirely on the server 110; or it can be implemented partly on the terminal 140 and partly on the server 110.

[0086] General Description of Embodiments in this Disclosure

[0087] According to one embodiment of this disclosure, a method for indexing and identifying gene sequencing sequences is provided.

[0088] like Figure 2 The diagram shown is a flowchart illustrating a gene sequencing sequence indexing and identification method provided in this disclosure. This method can be applied to a gene sequencing sequence indexing and identification device. The gene sequencing sequence indexing and identification method may include:

[0089] Step 210: Obtain multiple index signal templates;

[0090] Step 220: Obtain the nanopore sequencing time-series electrical signal data of the gene sequencing sequence to be identified;

[0091] Step 230: Extract index signal data from the nanopore sequencing time-series electrical signal data;

[0092] Step 240: Calculate the distance fraction between the index signal data and each index signal template;

[0093] Step 250: Determine the target index signal template from multiple index signal templates based on the distance score, and determine the target index sequence corresponding to the target index signal template as the index recognition result of the gene sequencing sequence to be identified.

[0094] In step 210, each index signal template corresponds to an index sequence, and the index signal template is obtained based on the excerpt of nanopore sequencing time-series electrical signal data containing the index.

[0095] An index signal template indicates the electrical signal data obtained during sequencing of the corresponding index sequence using a single-channel nanopore sequencer. The indexed nanopore sequencing time-series electrical signal data can be the electrical signal data generated during the sequencing of the gene sequence with the inserted index sequence by the nanopore sequencer; therefore, a segment of the electrical signal is the electrical signal data corresponding to the index sequence. For example... Figure 3 The image shows a segment of nanopore sequencing time-series electrical signal data, with the boxed portion representing the electrical signal data corresponding to the index sequence. Therefore, obtaining the index signal template allows extraction of the index-corresponding electrical signal data from the indexed nanopore sequencing time-series electrical signal data. Multiple index signal templates can be obtained from the electrical signal data corresponding to multiple gene sequences with different index sequences inserted.

[0096] When inserting an index sequence into the gene sequence used to acquire the index signal template, the insertion position and the length of the index signal are known precisely. The position of the electrical signal data corresponding to the index sequence within the electrical signal data of the entire gene sequence corresponds to the position where the index is inserted in the gene sequence. For example, if the index sequence is inserted at the beginning of the gene sequence, the position of the electrical signal data corresponding to the index sequence is also at the beginning of the electrical signal data of the entire gene sequence. When the length of the index sequence is known to be 100, 100 electrical signals from the beginning of the electrical signal data of the entire gene sequence can be extracted as the index signal template. Therefore, when extracting the index signal template, the extraction can be performed from the indexed nanopore sequencing time-series electrical signal data based on the insertion position and length of the index sequence in the corresponding gene sequence.

[0097] Determining the index sequence corresponding to the index signal template allows us to know the specific information of the index sequence explicitly when inserting it. For example, if the inserted index sequence is "ATCACG", then the corresponding index sequence can be obtained after truncating the index signal template.

[0098] The index sequence corresponding to the index signal template can also be determined based on the electrical signal data of the index signal template after the index signal template is extracted from the nanopore sequencing time-series electrical signal data. Different bases exhibit different electrical signals due to their unique physical and chemical properties during nanopore sequencing. Therefore, the corresponding base information can be obtained based on the electrical signal data in the index signal template. Specifically, pre-trained machine learning algorithms or statistical models can be used to identify the type of base corresponding to each electrical signal in the index signal template, thereby determining the index sequence corresponding to the index signal template.

[0099] After obtaining multiple index signal templates, the index signal templates can be stored in correspondence with the corresponding index sequences. For example, they can be stored in CSV file format (using plain text to store table data, and using commas (or other delimiters) to distinguish the index signal template fields from the index sequence fields).

[0100] In step 220, nanopore sequencing time-series electrical signal data of the gene sequencing sequence to be identified can be obtained.

[0101] The gene sequencing sequence to be identified can be a gene sequencing sequence that requires indexing and identification. A modifying gene can be inserted at the front end of the gene sequencing sequence to reduce adverse interactions between different sequences. The nanopore sequencing time-series electrical signal data of the gene sequencing sequence can be the raw electrical signal data obtained by sequencing the gene sequence using a nanopore sequencing instrument. Due to the influence of the modified gene, the front end of the nanopore sequencing time-series electrical signal data of the gene sequencing sequence contains a characteristic signal. This characteristic signal can be used to locate the nanopore sequencing time-series electrical signal data corresponding to the gene sequencing sequence.

[0102] Obtaining the time-series electrical signal data of the gene sequencing sequence to be identified can be achieved by real-time acquisition of the generated electrical signal data during the sequencing process of the nanopore sequencing instrument.

[0103] In step 230, index signal data can be extracted from the nanopore sequencing time-series electrical signal data of the gene sequencing sequence to be identified.

[0104] Index signal data can be a segment of electrical signal data containing the electrical signal data corresponding to the index sequence within the nanopore sequencing time-series electrical signal data of the gene sequencing sequence to be identified. For example... Figure 4As shown, the labeled portion is the index signal data extracted from the nanopore sequencing time-series electrical signal data of the gene sequencing sequence to be identified. Since different index sequences may be inserted at different positions within the gene sequencing sequence, and their lengths may also vary, it is difficult to determine the insertion position of the index sequence within the gene sequencing sequence itself. Therefore, index signal data can be extracted by estimating the approximate position of the index sequence. The condition for extracting index signal data is that the extracted range must cover multiple index signal templates.

[0105] In one implementation, index signal data can be extracted from the nanopore sequencing time-series electrical signal data of the gene sequencing sequence to be identified by extracting the first predetermined number of electrical signals from the electrical signal data as index signal data. For example, the first 1000 points can be extracted as index signal data. Since the index sequence is generally inserted at a predetermined position at the beginning of the gene sequencing sequence to be identified, the electrical signal data corresponding to the index sequence is also generally located at a predetermined position at the beginning of the overall electrical signal data corresponding to the gene sequencing sequence to be identified. Therefore, the first predetermined number of electrical signals can be extracted from the electrical signal data as index signal data.

[0106] In another implementation, index signal data is extracted from the nanopore sequencing time-series electrical signal data of the gene sequencing sequence to be identified, including:

[0107] Determine the truncation position and truncation range length of the index signal data in the nanopore sequencing time-series electrical signal data of the gene sequencing sequence to be identified;

[0108] In the nanopore sequencing time-series electrical signal data of the gene sequencing sequence to be identified, the signal sequence of the truncation range length at the truncation position is used as index signal data.

[0109] In this implementation, the position of the electrical signal corresponding to the index sequence can be located more precisely. First, the truncation position and truncation range length of the index signal data can be determined from the nanopore sequencing time-series electrical signal data of the gene sequencing sequence to be identified. The truncation position can indicate the starting position from which the index signal data is truncated from the electrical signal data, for example, starting from the very beginning of the electrical signal data, or starting from the 100th bit of the electrical signal data. The truncation range length can indicate the length of the index signal data truncated from the electrical signal data, for example, 300 bits or 500 bits.

[0110] In one implementation, determining the truncation position and truncation range length of index signal data in the nanopore sequencing time-series electrical signal data of the gene sequencing sequence to be identified includes:

[0111] Obtain the template truncation position and template truncation length of multiple index signal templates in the corresponding nanopore sequencing time-series electrical signal data;

[0112] Based on the template truncation position and template truncation length, the truncation position and truncation range length of the index signal data are determined in the nanopore sequencing time series electrical signal data of the gene sequencing sequence to be identified.

[0113] In step 210, when extracting multiple index signal templates from the multiple nanopore sequencing time-series electrical signal data, the template extraction position of the index signal template can be accurately obtained. The template extraction position can be the starting position of the extracted index signal template; the template extraction length can be the length of the extracted index signal template.

[0114] Since the index signal template corresponds to multiple index sequences, the truncation position and truncation length of the index signal template can provide a reference for the truncation position and truncation length of the electrical signal data corresponding to the index sequence.

[0115] Since the truncation positions and lengths of different index signal templates may differ, to ensure that the electrical signal data corresponding to different index sequences can be included in the index signal data, when determining the truncation position and truncation range length of the index signal data based on the truncation position and length, the truncation position and truncation range length of the index signal data can cover the truncation position and length corresponding to each index signal template. Therefore, the truncation position of the index signal data can be the first position among the truncation positions corresponding to multiple index signal templates, and the truncation range length can be the range length between the first position and the last position of the corresponding truncation endpoint position. The truncation endpoint position of the index signal template can be obtained by adding the truncation position and the truncation length.

[0116] For example Figure 5 As shown, the truncation position of index signal template m1 is bit 100, and the truncation length is 105, meaning index signal template m1 truncates from bit 100 to bit 205; the truncation position of index signal template m2 is bit 70, and the truncation length is 120, meaning index signal template m2 truncates from bit 70 to bit 190; the truncation position of index signal template m3 is bit 110, and the truncation length is 150, meaning index signal template m3 truncates from bit 110 to bit 260. To ensure that the index signal data can cover each index signal template, the index signal data can be truncated from bit 70 to bit 260. Therefore, the truncation position of the index signal data is bit 70, and the truncation range length is 190.

[0117] To avoid calculation errors, an error value can be added to the truncation position and length calculated based on the template truncation positions and lengths corresponding to multiple index signal templates. For example, the truncation position can be shifted forward by 50 positions, and the truncation length can be increased by 100 positions. Assuming the truncation position calculated based on the template truncation positions and lengths corresponding to multiple index signal templates is the 70th position, and the truncation range length is 190, then after adding the error value, the truncation position will be the 20th position, and the truncation range length will be 290.

[0118] Determining the truncation position and truncation range length of index signal data based on the template truncation position and template truncation length in the corresponding nanopore sequencing time series electrical signal data can cover the electrical signal range of multiple index signal templates, which is beneficial to improving the accuracy of determining the truncation position and truncation range length of index signal data.

[0119] After determining the truncation position and truncation range length of the index signal data, the signal data of the truncation range length can be extracted from the nanopore sequencing time-series electrical signal data of the gene sequencing sequence at the truncation position as the index signal data. For example, if the truncation position is the 50th position and the truncation range length is 500, then the signal data from the 50th to the 550th position in the nanopore sequencing time-series electrical signal data of the gene sequencing sequence to be identified can be extracted as the index signal data.

[0120] After determining the truncation position and truncation range length of the index signal data, the location of the electrical signal data corresponding to the index sequence can be more accurately located, which is beneficial to reducing the length of the index signal data and improving the computational efficiency when matching the index signal template for the index signal data.

[0121] In step 240, after obtaining multiple index signal templates and index signal data, the distance score between the index signal data and each index signal template can be calculated, thereby matching the target index signal template to the sequencing sequence of the gene to be identified. However, the amount of index signal templates and index signal data is large, resulting in low computational efficiency. Therefore, in another embodiment, after extracting the index signal data from the nanopore sequencing time-series electrical signal data of the sequencing sequence of the gene to be identified, the method further includes:

[0122] Multiple index signal templates are compressed to obtain multiple compressed index signal templates;

[0123] The index signal data is compressed to obtain compressed index signal data.

[0124] Compressing multiple index signal templates and index signal data separately before calculating the distance score can reduce the amount of data to be calculated, which is beneficial to improving the efficiency of matching target index signal templates for index signal data.

[0125] In one implementation, multiple index signal templates are compressed to obtain multiple compressed index signal templates, including:

[0126] Get a compressed window of a predetermined width;

[0127] The signal values ​​within the compression window of a predetermined width are compressed until the compression window slides to the last position, resulting in multiple compressed index signal templates.

[0128] A compression window can be used to define the range of each compression within an index signal template. A predetermined width indicates the number of signal values ​​compressed each time. The compression window slides backward from the first signal value of the index signal template to the last value in a predetermined step. The predetermined step indicates the distance the compression window slides in one iteration. For example... Figure 6 As shown, the signal values ​​of the index signal template range from s1 to s6, with a predetermined width of 3 and a predetermined step size of 1 for the compression window. During the first compression, s1 to s3 are compressed, and the compression window slides sequentially forward, compressing s2 to s4, s3 to s5, and s4 to s6 in turn. This results in the compressed index signal template.

[0129] In one implementation, compressing signal values ​​of a predetermined width within a compression window includes:

[0130] Calculate the standard deviation of signal values ​​within a predetermined width in the compressed window;

[0131] When the standard deviation is less than the compression threshold, the average value of the predetermined width of signal values ​​is used as the compressed value of the predetermined width of signal values ​​in the compression window.

[0132] Standard deviation measures the dispersion of multiple signal values ​​within a compression window. A large standard deviation indicates significant differences between the values, potentially leading to inaccurate compression results. Therefore, a compression threshold can be used to determine if the signal values ​​within the window are compressible, calculated for a predetermined width of signal values. If the standard deviation is less than the threshold, the signal values ​​are similar and can be compressed collectively. The average standard deviation is then used as the compression value for the predetermined width of signal values ​​within the window.

[0133] For example, the predetermined width of the compression window is 3, and it contains signal values ​​of 1.8, 3.2, and 2.5, with a standard deviation of 0.7. The compression threshold is 2, and the standard deviation of the signal values ​​is less than the compression threshold. Therefore, their average value (2.5) is calculated as the compressed value of these three signal values ​​in the compression window.

[0134] If the standard deviation is greater than or equal to the compression threshold, the signal values ​​in the compression window do not need to be compressed. For example, the predetermined width of the compression window is 3, and it contains signal values ​​of 1.8, 7.2, and 0.9, with a standard deviation of 3.4. The compression threshold is 2, and the standard deviation of the signal values ​​is greater than the compression threshold; therefore, these three signal values ​​are not compressed.

[0135] Compressing the signal values ​​in the compression window after calculating the standard deviation of multiple signal values ​​allows for the compression of signal values ​​with similar values, while retaining signal values ​​with large differences without compression. This helps improve the accuracy of calculating distance scores using compressed data.

[0136] As the compression window slides across the index signal template, the signal values ​​within the compression window are compressed until the compression window reaches the last digit, resulting in the compressed index signal template corresponding to the index signal template.

[0137] By sliding a compression window of a predetermined width on the index signal template, the signal values ​​in the compression window are compressed to obtain the compressed index signal template, ensuring that each signal value can be compressed with adjacent signal values, thereby improving compression efficiency and accuracy.

[0138] The compressed index signal data is obtained by compressing the index signal data in the same way as compressing the index signal template, so it will not be described again here.

[0139] After compressing multiple index signal templates and index signal data, the distance score between the index signal data and each index signal template is calculated, including: calculating the distance score between the compressed index signal data and each compressed index signal template.

[0140] The method for calculating the distance fraction between the compressed index signal data and each compressed index signal template is the same as the process for calculating the distance fraction between the index signal data and each index signal template, which will be described in detail below.

[0141] In one implementation, calculating the distance fraction between the index signal data and each index signal template includes:

[0142] For each index signal template, multiple sub-signal data are obtained from the index signal data;

[0143] Calculate the sub-distance scores between multiple sub-signal data and their corresponding index signal templates;

[0144] The distance score between the index signal data and the corresponding index signal template is determined based on multiple sub-distance scores.

[0145] Because the index signal data is much larger than the index signal template, multiple sub-signal data can be obtained from the index signal data. Sub-signal data can be sub-sequences within the index signal data.

[0146] In one implementation, for each index signal template, multiple sub-signal data are obtained from the index signal data, including:

[0147] For each index signal template, obtain the corresponding sequence truncation window and sequence truncation step size;

[0148] Starting from a predetermined position in the index signal data, the sequence truncation window is moved according to the sequence truncation step size to obtain multiple sub-signal data.

[0149] A sequence truncation window can move over the indexed signal data to extract sub-signal data. The sequence truncation step size is the size of each movement of the sequence truncation window over the indexed signal data. Each time the sequence truncation window moves over the indexed signal data, the sequence within the sequence truncation window is captured as a sub-signal data. For example... Figure 7 As shown, the sequence truncation step size is 2, which means that every two units, the sequence truncation window is used to truncate the sub-signal data once. The width of the sequence truncation window is 5, which means that the length of the sub-signal data truncated each time is 5. Based on this sequence truncation window, 6 sub-signal data can be obtained from the index signal data.

[0150] In one implementation, for each index signal template, the sequence truncation window and sequence truncation step size are obtained, including:

[0151] For each index signal template, obtain the signal sequence length of the index signal template;

[0152] The sequence truncation window and sequence truncation step size are determined based on the length of the signal sequence.

[0153] The signal sequence length of the index signal template can be the electrical signal data length of the index signal template.

[0154] Different sequence truncating windows can be obtained for different index signal templates. The main difference between these windows is their width. To ensure a more accurate match between the sub-signal data and the index signal template, the length of the sub-signal data can be the same as the length of the signal sequence in the index signal template. Therefore, the width of the sequence truncating window can be the same as the length of the signal sequence in the index signal template. For example, if the length of one index signal template is 100, the width of the sequence truncating window can be 100, extracting multiple sub-signal data of length 100 from the index signal data; if the length of another index signal template is 120, the width of the sequence truncating window can be 120, extracting multiple sub-signal data of length 120 from the index signal data.

[0155] Different sequence truncation step sizes can be obtained for different index signal templates. When the signal sequence length of the index signal template is long, the length of the sub-signal data truncated based on the signal sequence length is also long. In this case, when calculating the sub-distance score between the index signal template and multiple sub-signal data, if the differences between different sub-signal data are small, the calculated sub-distance scores will also be relatively similar. For example, if the sequence truncation step size is 1, meaning that one sub-signal data is truncated for every unit the sequence truncation window moves, the difference between two adjacent sub-signal data is only an electrical signal. When calculating the sub-distance score using these two sub-signal data and the index signal template respectively, the two sub-distance score results will be quite similar. Therefore, for index signal templates with long signal sequence lengths, a sequence truncation step size that is too short will lead to insufficiently significant calculation results for different sub-signal data, resulting in a significant waste of computational resources.

[0156] Therefore, determining the sequence truncation step size based on the signal sequence length can be achieved by classifying the signal sequence length into different levels, and determining different sequence truncation step sizes for different levels of signal sequence length. Table 1 shows an example of a correspondence between different levels of signal sequence length and sequence truncation step sizes.

[0157] Signal sequence length Sequence truncation step size Less than 200 2 200-350 3 Greater than 350 4

[0158] Table 1

[0159] In Table 1, the signal sequence length is divided into three levels: less than 200, 200 to 350, and greater than 350. When the signal sequence length is less than 200, the corresponding sequence truncation step size is 2; when the signal sequence length is between 200 and 350, the corresponding sequence truncation step size is 3; and when the signal sequence length is greater than 350, the corresponding sequence truncation step size is 4.

[0160] For different index signal templates, the sequence truncation window is determined based on the corresponding signal sequence length. This allows the identification of the sub-signal data that best matches the index signal template within the index signal data, improving the accuracy of calculating the distance score between the index signal template and the index signal data. Determining the sequence truncation step size based on the signal sequence length ensures that the calculated sub-distance score between the index signal template and each sub-signal data has strong significance, reducing invalid calculations and improving the efficiency of distance score calculation.

[0161] After obtaining the sequence truncation window and the sequence truncation step size, the sequence truncation window can be moved from a predetermined position in the index signal data according to the sequence truncation step size. The starting position of the sequence truncation window on the index signal data can be preset, for example, starting from the first position of the index signal data, or starting from the k-th position of the index signal data. The starting position can be determined based on the index signal template for which distance fraction calculation is being performed. If the electrical signal data corresponding to the index signal template generally appears at the beginning of the index signal data, then the sub-signal data can be obtained starting from the first position of the index signal data; if the electrical signal data corresponding to the index signal template generally appears at the end of the index signal data, then the sub-signal data can be obtained starting from the middle position of the index signal data.

[0162] By using a sequence truncation window and a sequence truncation step size to obtain multiple sub-signal data from the index signal data, it is ensured that no electrical signal data in the index signal data is missed, thus improving the comprehensiveness of obtaining multiple sub-signal data.

[0163] After acquiring multiple sub-signal data, the sub-distance scores between the multiple sub-signal data and the index signal template can be calculated.

[0164] In one implementation, calculating the sub-distance scores between multiple sub-signal data and the index signal template includes:

[0165] Create a distance matrix where the number of rows equals the length of the sub-signal data and the number of columns equals the length of the index signal template. Each element in the distance matrix represents the distance between the sub-signal data and the index signal template at the corresponding position.

[0166] Calculate the signal distance between each signal value in the sub-signal data and each signal value in the index signal template, and fill the distance matrix with the signal distance;

[0167] Use dynamic programming to calculate the shortest path from the top left corner to the bottom right corner of the distance matrix;

[0168] The sum of signal distances along the shortest path is calculated as the sub-distance score between the sub-signal data and the index signal template.

[0169] Since both the sub-signal data and the index signal template are time series, and time series can sometimes experience delays, it can be difficult for the two time series to match precisely. A distance matrix can be used to ignore these time series delays and calculate the distance between each signal value in the sub-signal data and each signal value in the index signal template; the smaller the distance, the higher the similarity.

[0170] In the above implementation, the signal distance between each signal value in the sub-signal data and each signal value in the index signal template can be calculated using methods such as Euclidean distance, Manhattan distance, and cosine similarity. After calculating the signal distance, it can be filled into the corresponding position in the distance matrix. For example, the signal distance between the second signal value in the sub-signal data and the third signal value in the index signal template can be filled into the position of the second row and third column of the distance matrix.

[0171] After filling the distance matrix with the signal distances between each signal value in the sub-signal data and each signal value in the index signal template, a dynamic programming algorithm is used to calculate the shortest path from the top left corner to the bottom right corner of the distance matrix. The shortest path represents the optimal alignment of the sub-signal data and the index signal template on the time axis. The dynamic programming algorithm searches for the value with the smallest signal distance from the top left corner to the bottom right corner of the distance matrix one by one until the shortest path from the top left corner to the bottom right corner is determined.

[0172] The sum of distances along the shortest path represents the similarity distance between the sub-signal data and the index signal template, which is the sub-distance score. The smaller the sub-distance score, the more similar the sub-signal data and the index signal template are in shape.

[0173] The above method can ignore the time series delay, thereby calculating the similarity between the sub-signal data and the index signal template, which helps to improve the accuracy of calculating the sub-distance score between the sub-signal data and the index signal template.

[0174] The above implementation can calculate the sub-distance scores between sub-signal data and index signal templates of different lengths, but the computational complexity is high. Therefore, based on the above implementation, in another implementation, before creating the distance matrix, the method further includes: converting the sub-signal data and index signal template into sequences of the same length.

[0175] Transforming sub-signal data and index signal templates into sequences of the same length can be achieved by using interpolation to extend them to the same length, or by using downsampling to reduce them to the same length, etc.

[0176] Converting the sub-signal data and the index signal template into sequences of the same length and then calculating the sub-distance score between them using the distance matrix can improve computational efficiency. Experiments have verified that converting the sub-signal data and the index signal template into sequences of the same length improves the computational efficiency of calculating the sub-distance score between them by more than 50 times.

[0177] After calculating the sub-distance scores between multiple sub-signal data and the index signal template, the distance score between the index signal data and the index signal template can be determined based on these sub-distance scores. A smaller sub-distance score indicates a closer relationship between the corresponding sub-signal data and the index signal template; therefore, the smallest of the multiple sub-distance scores can be used as the distance score between the index signal data and the index signal template.

[0178] Determining the distance fraction between the index signal data and the index signal template based on the sub-distance fractions of multiple sub-signal data in the index signal data can reduce the amount of data required to calculate the distance fraction, thereby improving the efficiency of calculating the distance fraction.

[0179] In step 250, a target index signal template is determined from multiple index signal templates based on the distance score, and the target index sequence corresponding to the target index signal template is determined as the index recognition result of the gene sequencing sequence to be identified.

[0180] In one implementation, the index signal template with the smallest distance score to the index signal data among multiple index signal templates can be determined as the target index signal template. For example, if index signal template m1 has a distance score of 3 to the index signal data, index signal template m2 has a distance score of 5 to the index signal data, and index signal template m3 has a distance score of 1 to the index signal data, then index signal template m3 is determined as the target index signal template.

[0181] In another implementation, a target index signal template is determined from multiple index signal templates based on a distance score, and the index sequence corresponding to the target index signal template is determined as the index identification result of the sequencing sequence of the gene to be identified, including:

[0182] Determine the target distance score from the distance scores corresponding to multiple index signal templates;

[0183] When the target distance score is less than the distance threshold, the index signal template corresponding to the target distance score is determined as the target index signal template, and the target index sequence corresponding to the target index signal template is determined as the index recognition result of the gene sequencing sequence to be identified.

[0184] Determining the target distance score from the distance scores corresponding to multiple index signal templates involves identifying the smallest of these distance scores as the target distance score. However, even though the target distance score is the smallest among the distance scores corresponding to multiple index signal templates, this does not necessarily mean that the index signal template corresponding to the target distance score is the same as the target index signal template contained in the index signal data. This is because when the distance scores of multiple index signal templates and the index signal data are all relatively large, even if the target distance score is the smallest compared to other distance scores, the corresponding index signal template may still have a significant difference from the index signal data.

[0185] Therefore, after determining the target distance score, a distance threshold can be used to determine whether the target distance score is valid.

[0186] The distance threshold can be a preset value. In one implementation, the distance threshold can be determined in the following way:

[0187] Retrieve the index matching set;

[0188] Calculate the matching score between sequencing electrical signal data and corresponding index electrical signal data in the index matching set;

[0189] The distance threshold is determined based on multiple matching scores.

[0190] The index matching set stores multiple sequencing electrical signal data and their corresponding index electrical signal data. The index electrical signal data can be an index signal template that can match the sequencing electrical signal data. The index electrical signal data corresponding to the sequencing electrical signal data can be obtained manually by comparing the sequencing electrical signal data with multiple index signal templates; alternatively, it can be obtained by indexing the sequencing electrical signal data using the gene sequencing sequence indexing identification method of this disclosure, followed by manual review. After obtaining the sequencing electrical signal data and their corresponding index electrical signal data, the sequencing electrical signal data and index electrical signal data can be stored in the index matching set.

[0191] The matching score between sequencing electrical signal data and index electrical signal data in the index matching set can provide a reference for determining the distance threshold. The matching score indicates the similarity between sequencing electrical signal data and the corresponding index electrical signal data; the smaller the matching score, the higher the similarity between the sequencing electrical signal data and the corresponding index electrical signal data. The method for calculating the matching score between sequencing electrical signal data and index electrical signal data is the same as the process for calculating the distance score between index signal data and index signal template in the previous embodiment, and will not be repeated here.

[0192] Determining the distance threshold based on multiple matching scores allows us to select the highest matching score as the distance threshold. A higher matching score indicates a greater distance between the sequencing electrical signal data and the index electrical signal data. Since both the sequencing electrical signal data and their corresponding index electrical signal data in the index matching set are valid, a distance score less than the highest matching score in the index matching set can be considered a valid distance score.

[0193] Using matched sequencing electrical signal data from the index matching set and index electrical signal data to determine the distance threshold, and using the matching of historical sequencing electrical signal data as a reference, helps to improve the reliability of determining the distance threshold.

[0194] When using a distance threshold to determine the validity of a target distance score, if the target distance score is less than the threshold, it indicates that the target distance score is valid. Therefore, the index signal template corresponding to the target distance score can be identified as the target index signal template, and the target index sequence corresponding to the target index signal template can be identified as the index recognition result of the gene sequencing sequence to be identified. If the target distance score is greater than or equal to the distance threshold, it indicates that the distance between the index signal template corresponding to the target distance score and the index signal data is too large, making it unsuitable as a target index signal template, and re-identification is required.

[0195] After determining the validity of the target distance score using a distance threshold, determining the target index signal template can verify the distance score, which helps improve the accuracy of determining the target index signal template.

[0196] After determining the target index signal template, it can be assumed that the index signal data contains the target index signal template. In other words, the nanopore sequencing time-series electrical signal data of the gene sequencing sequence to be identified contains the target index signal template. The target index sequence corresponding to the target index signal template can be an index sequence inserted into the gene sequencing sequence to be identified. Therefore, the target index sequence is determined as the index identification result of the gene sequencing sequence to be identified.

[0197] Description of apparatus and devices according to embodiments of this disclosure

[0198] It is understood that although the steps in the above flowcharts are shown sequentially according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated in this embodiment, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some steps in the above flowcharts may include multiple steps or multiple stages. These steps or stages are not necessarily completed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential, but can be performed alternately or in turn with other steps or at least some of the steps or stages in other steps.

[0199] It should be noted that in the various specific embodiments of this disclosure, when processing is required based on data related to the characteristics of the target object, such as target object attribute information or a set of attribute information, the permission or consent of the target object will be obtained first. Furthermore, the collection, use, and processing of this data will comply with the relevant laws, regulations, and standards of the relevant regions. In addition, when this application embodiment needs to obtain target object attribute information, separate permission or consent from the target object will be obtained through pop-up windows or redirection to a confirmation page. Only after obtaining the target object's separate permission or consent will the necessary target object-related data for the normal operation of this application embodiment be obtained.

[0200] Figure 8 A schematic diagram of the structure of a gene sequencing sequence indexing and identification device 800 provided in this embodiment of the disclosure. The gene sequencing sequence indexing and identification device 800 includes:

[0201] The first acquisition unit 810 is used to acquire multiple index signal templates, each index signal template corresponding to an index sequence. The index signal template is obtained based on the excerpt of nanopore sequencing time series electrical signal data containing the index.

[0202] The second acquisition unit 820 is used to acquire the nanopore sequencing time-series electrical signal data of the sequencing sequence of the gene to be identified.

[0203] The extraction unit 830 is used to extract index signal data from the nanopore sequencing time-series electrical signal data of the gene sequencing sequence to be identified.

[0204] Calculation unit 840 is used to calculate the distance fraction between the index signal data and each index signal template;

[0205] The determining unit 850 is used to determine the target index signal template from multiple index signal templates based on the distance score, and to determine the target index sequence corresponding to the target index signal template as the index recognition result of the gene sequencing sequence to be identified.

[0206] Optionally, in one embodiment, the computing unit 840 is specifically used for:

[0207] For each index signal template, multiple sub-signal data are obtained from the index signal data;

[0208] Calculate the sub-distance scores between multiple sub-signal data and their corresponding index signal templates;

[0209] The distance score between the index signal data and the corresponding index signal template is determined based on multiple sub-distance scores.

[0210] Optionally, in one embodiment, the computing unit 840 is specifically used for:

[0211] For each index signal template, obtain the corresponding sequence truncation window and sequence truncation step size;

[0212] Starting from a predetermined position in the index signal data, the sequence truncation window is moved according to the sequence truncation step size to obtain multiple sub-signal data.

[0213] Optionally, in one embodiment, the computing unit 840 is specifically used for:

[0214] For each index signal template, obtain the signal sequence length of the index signal template;

[0215] The sequence truncation window and sequence truncation step size are determined based on the length of the signal sequence.

[0216] Optionally, in one embodiment, the computing unit 840 is specifically used for:

[0217] Create a distance matrix where the number of rows equals the length of the sub-signal data and the number of columns equals the length of the index signal template. Each element in the distance matrix represents the distance between the sub-signal data and the index signal template at the corresponding position.

[0218] Calculate the signal distance between each signal value in the sub-signal data and each signal value in the index signal template, and fill the distance matrix with the signal distance;

[0219] Use dynamic programming to calculate the shortest path from the top left corner to the bottom right corner of the distance matrix;

[0220] The sum of signal distances along the shortest path is calculated as the sub-distance score between the sub-signal data and the index signal template.

[0221] Optionally, in one embodiment, the gene sequencing sequence indexing and identification device 800 further includes:

[0222] The first compression unit (not shown) is used to compress multiple index signal templates to obtain multiple compressed index signal templates.

[0223] The second compression unit (not shown) is used to compress the index signal data to obtain compressed index signal data.

[0224] The computing unit 840 is specifically used for:

[0225] Calculate the distance fraction between the compressed index signal data and each compressed index signal template.

[0226] Optionally, the first compression unit (not shown) is specifically used for:

[0227] Obtain a compressed window of a predetermined width. The compressed window is used to slide backward from the first signal value of the index signal template to the last value according to a predetermined step size.

[0228] The signal values ​​within the compression window of a predetermined width are compressed until the compression window slides to the last position, resulting in multiple compressed index signal templates.

[0229] Optionally, the first compression unit (not shown) is specifically used for:

[0230] Calculate the standard deviation of signal values ​​within a predetermined width in the compressed window;

[0231] When the standard deviation is less than the compression threshold, the average of the signal values ​​of a predetermined width is used as the compressed value of the signal values ​​of the predetermined width in the compression window.

[0232] Optionally, the determination unit 850 is specifically used for:

[0233] Determine the target distance score from the distance scores corresponding to multiple index signal templates;

[0234] When the target distance score is less than the distance threshold, the index signal template corresponding to the target distance score is determined as the target index signal template, and the target index sequence corresponding to the target index signal template is determined as the index recognition result of the gene sequencing sequence to be identified.

[0235] Optionally, the distance threshold is determined in the following way:

[0236] Obtain the index matching set, which stores multiple sequencing electrical signal data and their corresponding index electrical signal data;

[0237] Calculate the matching score between sequencing electrical signal data and corresponding index electrical signal data in the index matching set;

[0238] The distance threshold is determined based on multiple matching scores.

[0239] Optionally, the interception unit 830 is specifically used for:

[0240] Determine the truncation position and truncation range length of the index signal data in the nanopore sequencing time-series electrical signal data of the gene sequencing sequence to be identified;

[0241] In the nanopore sequencing time-series electrical signal data of the gene sequencing sequence to be identified, the signal sequence of the truncation range length at the truncation position is used as index signal data.

[0242] Optionally, the interception unit 830 is specifically used for:

[0243] Obtain the template truncation position and template truncation length of multiple index signal templates in the corresponding nanopore sequencing time-series electrical signal data;

[0244] Based on the template truncation position and template truncation length, the truncation position and truncation range length of the index signal data are determined in the nanopore sequencing time series electrical signal data of the gene sequencing sequence to be identified.

[0245] Reference Figure 9 , Figure 9 The structural block diagram of a portion of the terminal 140 for implementing the gene sequencing sequence indexing and identification method of this embodiment is shown below. The terminal 140 includes: a radio frequency (RF) circuit 910, a memory 915, an input unit 930, a display unit 940, a sensor 950, an audio circuit 960, a wireless fidelity (WiFi) module 970, a processor 980, and a power supply 990, among other components. Those skilled in the art will understand that... Figure 9 The terminal 140 structure shown does not constitute a limitation on a mobile phone or computer, and may include more or fewer components than shown, or combine certain components, or have different component arrangements.

[0246] The RF circuit 910 can be used to receive and transmit signals during information transmission or calls. In particular, it receives downlink information from the base station and processes it with the processor 980; in addition, it transmits uplink data to the base station.

[0247] The memory 915 can be used to store software programs and modules. The processor 980 executes various functional applications of the terminal and the identification of lane line change points by running the software programs and modules stored in the memory 915.

[0248] The input unit 930 can be used to receive input numeric or character information, and to generate key signal inputs related to the terminal's settings and function control. Specifically, the input unit 930 may include a touch panel 931 and other input devices 932.

[0249] The display unit 940 can be used to display input or provided information, as well as various menus of the terminal. The display unit 940 may include a display panel 941.

[0250] Audio circuitry 960, speaker 961, and microphone 962 provide an audio interface.

[0251] In this embodiment, the processor 980 included in the terminal 140 can execute the gene sequencing sequence indexing and identification method of the previous embodiment.

[0252] The terminal 140 in this disclosure includes, but is not limited to, mobile phones, computers, intelligent voice interaction devices, smart home appliances, vehicle terminals, and aircraft. This invention can be applied to various scenarios, including but not limited to gene sequencing and biometrics.

[0253] Figure 10 This is a partial structural block diagram of a server 110 for implementing the gene sequencing sequence indexing and identification method according to embodiments of the present disclosure. The server 110 can vary significantly due to different configurations or performance, and may include one or more central processing units (CPUs) 1022 (e.g., one or more processors) and storage devices 1032, and one or more storage media 1030 (e.g., one or more mass storage devices) for storing application programs 1042 or data 1044. The storage devices 1032 and storage media 1030 can be temporary or persistent storage. The program stored in the storage media 1030 may include one or more modules (not shown in the diagram), each module including a series of instruction operations on the server 110. Furthermore, the central processing unit 1022 may be configured to communicate with the storage media 1030 and execute the series of instruction operations in the storage media 1030 on the server 110.

[0254] Server 110 may also include one or more power supplies 1026, one or more wired or wireless network interfaces 1050, one or more input / output interfaces 1058, and / or one or more operating systems 1041, such as Windows Server™, Mac OS X™, Unix™, Linux™, FreeBSD™, etc.

[0255] The central processing unit 1022 in server 110 can be used to execute the gene sequencing sequence indexing and identification method of the present disclosure embodiments.

[0256] This disclosure also provides a computer-readable storage medium for storing program code for executing the gene sequencing sequence indexing and identification method of the foregoing embodiments.

[0257] This disclosure also provides a computer program product comprising a computer program. A processor of a computer device reads and executes the computer program, causing the computer device to perform the above-described gene sequencing sequence indexing and identification method.

[0258] The terms “first,” “second,” “third,” “fourth,” etc. (if present) in this disclosure and the foregoing drawings are used to distinguish similar objects and are not necessarily used to describe a particular order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this disclosure described herein can be implemented, for example, in orders other than those illustrated or described herein. Furthermore, the terms “comprising” and “including,” and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or apparatus that includes a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatuses.

[0259] It should be understood that in this disclosure, "at least one item" means one or more, and "more than one" means two or more. "And / or" is used to describe the relationship between related objects, indicating that three relationships can exist. For example, "A and / or B" can represent three cases: only A exists, only B exists, and both A and B exist simultaneously, where A and B can be singular or plural. The character " / " generally indicates that the preceding and following related objects are in an "or" relationship. "At least one of the following" or similar expressions refer to any combination of these items, including any combination of single or plural items. For example, at least one of a, b, or c can represent: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, and c can be single or multiple.

[0260] It should be understood that in the description of the embodiments disclosed herein, "multiple" means two or more, "greater than", "less than", "exceeding" etc. are understood to exclude the number itself, and "above", "below", "within" etc. are understood to include the number itself.

[0261] In the several embodiments provided in this disclosure, it should be understood that the disclosed systems, apparatuses, and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces, indirect coupling or communication connection between apparatuses or units, and may be electrical, mechanical, or other forms.

[0262] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0263] Furthermore, the functional units in the various embodiments of this disclosure can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.

[0264] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this disclosure, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods of the various embodiments of this disclosure. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0265] It should also be understood that the various implementation methods provided in this disclosure can be combined arbitrarily to achieve different technical effects.

[0266] The above is a detailed description of the embodiments of this disclosure. However, this disclosure is not limited to the above embodiments. Those skilled in the art can make various equivalent modifications or substitutions without departing from the spirit of this disclosure. All such equivalent modifications or substitutions are included within the scope defined by the claims of this disclosure.

Claims

1. A method for indexing and identifying gene sequencing sequences, characterized in that, include: Multiple index signal templates are obtained, each index signal template corresponding to an index sequence. The index signal templates are obtained based on the excerpt of nanopore sequencing time-series electrical signal data containing the index. Obtain time-series electrical signal data of nanopore sequencing of the gene sequence to be identified; Extract index signal data from the time-series electrical signal data of the nanopore sequencing; Calculate the distance score between the index signal data and each index signal template; Based on the distance score, a target index signal template is determined from multiple index signal templates, and the target index sequence corresponding to the target index signal template is determined as the index recognition result of the sequencing sequence of the gene to be identified.

2. The method according to claim 1, characterized in that, The calculation of the distance score between the index signal data and each index signal template includes: For each of the index signal templates, multiple sub-signal data are obtained from the index signal data; Calculate the sub-distance scores between the multiple sub-signal data and the corresponding index signal template; The distance score between the index signal data and the corresponding index signal template is determined based on multiple sub-distance scores.

3. The method according to claim 2, characterized in that, For each of the index signal templates, obtaining multiple sub-signal data from the index signal data includes: For each index signal template, obtain the corresponding sequence truncation window and sequence truncation step size; Starting from a predetermined position in the index signal data, the sequence truncation window is moved according to the sequence truncation step size to obtain multiple sub-signal data.

4. The method according to claim 3, characterized in that, For each of the index signal templates, obtaining the sequence truncation window and the sequence truncation step size includes: For each of the index signal templates, obtain the signal sequence length of the index signal template; The sequence truncation window and sequence truncation step size are determined based on the length of the signal sequence.

5. The method according to claim 2, characterized in that, The calculation of sub-distance scores between the multiple sub-signal data and the corresponding index signal template includes: Create a distance matrix, wherein the number of rows in the distance matrix is ​​equal to the length of the sub-signal data and the number of columns is equal to the length of the index signal template, and each element in the distance matrix represents the distance between the sub-signal data and the index signal template at the corresponding position; Calculate the signal distance between each signal value in the sub-signal data and each signal value in the index signal template, and fill the distance matrix with the signal distance; Calculate the shortest path from the top left corner to the bottom right corner of the distance matrix using dynamic programming algorithm; The sum of signal distances along the shortest path is calculated as the sub-distance score between the sub-signal data and the index signal template.

6. The method according to claim 1, characterized in that, After extracting index signal data from the nanopore sequencing time-series electrical signal data of the gene sequencing sequence to be identified, the method includes: Multiple index signal templates are compressed to obtain multiple compressed index signal templates; The index signal data is compressed to obtain compressed index signal data; The calculation of the distance score between the index signal data and each index signal template includes: Calculate the distance fraction between the compressed index signal data and each of the compressed index signal templates.

7. The method according to claim 6, characterized in that, The step of compressing multiple index signal templates to obtain multiple compressed index signal templates includes: Obtain a compression window of a predetermined width, the compression window being used to slide backward from the first signal value of the index signal template according to a predetermined step size until the last value; The signal values ​​of a predetermined width in the compression window are compressed until the compression window slides to the last position, resulting in multiple compressed index signal templates.

8. The method according to claim 7, characterized in that, The compression of the predetermined width of signal values ​​in the compression window includes: Calculate the standard deviation of the signal values ​​within the predetermined width in the compressed window; When the standard deviation is less than the compression threshold, the average value of the predetermined width of signal values ​​is used as the compressed value of the predetermined width of signal values ​​in the compression window.

9. The method according to claim 1, characterized in that, The step of determining the target index signal template from multiple index signal templates based on the distance score, and determining the target index sequence corresponding to the target index signal template as the index identification result of the sequencing sequence of the gene to be identified, includes: Determine the target distance score from the distance scores corresponding to the multiple index signal templates; When the target distance score is less than the distance threshold, the index signal template corresponding to the target distance score is determined as the target index signal template, and the target index sequence corresponding to the target index signal template is determined as the index recognition result of the sequencing sequence of the gene to be identified.

10. The method according to claim 9, characterized in that, The distance threshold is determined in the following way: Obtain an index matching set, which stores multiple sequencing electrical signal data and their corresponding index electrical signal data; Calculate the matching score between the sequencing electrical signal data and the corresponding index electrical signal data in the index matching set; The distance threshold is determined based on multiple matching scores.

11. A gene sequencing sequence indexing and identification device, characterized in that, include: The first acquisition unit is used to acquire multiple index signal templates, each of which corresponds to an index sequence. The index signal template is obtained based on the excerpt of nanopore sequencing time-series electrical signal data containing the index. The second acquisition unit is used to acquire the nanopore sequencing time-series electrical signal data of the gene sequencing sequence to be identified; The extraction unit is used to extract index signal data from the nanopore sequencing time-series electrical signal data; A calculation unit is used to calculate the distance score between the index signal data and each of the index signal templates; The determining unit is used to determine the target index signal template from multiple index signal templates based on the distance score, and to determine the target index sequence corresponding to the target index signal template as the index recognition result of the sequencing sequence of the gene to be identified.

12. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, it implements the gene sequencing sequence indexing and identification method according to any one of claims 1 to 10.

13. A computer device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that, When the processor executes the computer program, it implements the gene sequencing sequence indexing and identification method according to any one of claims 1 to 10.

14. A computer program product comprising a computer program that is read and executed by a processor of a computer device, causing the computer device to perform the indexing and identification method for gene sequencing sequences according to any one of claims 1 to 10.