Data preprocessing method and device, electronic equipment and storage medium

By standardizing and clustering the sequences to be identified in solid nanopore analysis, the problem of difficulty in distinguishing biomolecules with similar sizes is solved, thus improving the accuracy of biomolecule classification.

CN116913403BActive Publication Date: 2026-04-28CHINA MOBILE CHENGDU INFORMATION & TELECOMM TECH CO LTD +1
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
CHINA MOBILE CHENGDU INFORMATION & TELECOMM TECH CO LTD
Filing Date
2023-06-29
Publication Date
2026-04-28

AI Technical Summary

Technical Problem

Existing technologies for analyzing biomolecular structures using solid-state nanopores have difficulty effectively distinguishing biomolecules of similar size, resulting in a lack of significant distinguishability in the characteristics of through-pore events and making it impossible to effectively distinguish a large number of overlapping through-pore signals in the intermediate region.

Method used

By determining the initial length of the sequence to be identified, standardizing it, and obtaining a target sequence of uniform length, the multiple data of the target sequence to be identified are divided into cluster sets to improve classification accuracy.

Benefits of technology

Clustering can clearly reveal the characteristics of data in the sequence to be identified, avoiding the inability to effectively distinguish biomolecules by threshold when their sizes are similar, thereby improving the accuracy of biomolecule classification.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116913403B_ABST
    Figure CN116913403B_ABST
Patent Text Reader

Abstract

Embodiments of the present application provide a data preprocessing method, comprising: determining an initial length of a to-be-identified sequence; the to-be-identified sequence is used to represent a biomolecule structure; performing standardization processing on the to-be-identified sequence based on the initial length to obtain a target to-be-identified sequence; performing clustering processing on a plurality of data in the target to-be-identified sequence to obtain a plurality of clustering sets; the clustering set is used to identify a category of the to-be-identified sequence. Embodiments of the present application also provide a data preprocessing method and device, an electronic device and a storage medium.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the fields of solid-state nanopore data analysis and artificial intelligence, and in particular to a data preprocessing method, apparatus, electronic device and storage medium. Background Technology

[0002] Currently, the analysis of biomolecular structures using perforation event data is still in its early stages. Published methods include: a method for detecting microRNA (miRNA) based on solid-state nanopore sensors, which involves statistically analyzing the Gaussian distribution of the amplitude (difference between current and baseline current values) of the perforation signals of two tumor markers, miRNA-21 and miRNA-486, and then manually observing and selecting an appropriate threshold for differentiation; and a method for estimating protein conformation and morphology characteristics based on nanopore perforation current, which statistically analyzes the relative blocking current and again selects an appropriate threshold for differentiation.

[0003] However, the magnitude of the amplitude depends on the degree of blockage of the nanopores, which is the ratio of the biomolecule size to the nanopore size. When it is necessary to identify biomolecules with similar sizes, the amplitude of the through-pore event does not have significant distinguishability. There are a large number of overlapping through-pore signals in the middle region, which makes it impossible to effectively distinguish biomolecules with similar sizes. Summary of the Invention

[0004] This application provides a data preprocessing method, apparatus, electronic device, and storage medium.

[0005] The technical solution of this application is implemented as follows:

[0006] This application provides a data preprocessing method, including:

[0007] Determine the initial length of the sequence to be identified; the sequence to be identified is used to characterize the structure of biomolecules;

[0008] Based on the initial length, the sequence to be identified is standardized to obtain the target sequence to be identified;

[0009] Multiple data points in the target sequence to be identified are clustered to obtain multiple cluster sets; the cluster sets are used to identify the category of the sequence to be identified.

[0010] This application provides a data preprocessing apparatus, including:

[0011] A determining unit is used to determine the initial length of the sequence to be identified; the sequence to be identified is used to characterize the structure of a biomolecule.

[0012] The first processing unit is used to perform standardization processing on the sequence to be identified based on the initial length to obtain the target sequence to be identified;

[0013] The second processing unit is used to perform clustering processing on multiple data in the target sequence to be identified, to obtain multiple cluster sets; the cluster sets are used to identify the category of the sequence to be identified.

[0014] This application provides an electronic device, including a memory and a processor. The memory stores a computer program that can run on the processor to execute the data preprocessing method provided above.

[0015] This application provides a computer-readable storage medium having a computer program stored thereon, the computer program causing a computer to perform the data preprocessing method provided above.

[0016] In some embodiments of this application, the technical solutions employ a sequence to be identified to characterize the structure of biomolecules, determine the initial length of the sequence to be identified, standardize the sequence to be identified based on the initial length to obtain a target sequence to be identified with uniform length, and divide multiple data of the target sequence to be identified using a clustering method to obtain multiple cluster sets. The data in each cluster set have similarity. In this way, by aggregating similar data into the same cluster set, the characteristics of the data in the sequence to be identified can be clearly understood. Thus, in subsequent processing, the cluster set can be used to improve the accuracy of the classification of the sequence to be identified, and avoid the inability to effectively distinguish biomolecules by using thresholds when the sizes of biomolecules are similar, thereby preventing the identification of the category of biomolecules. Attached Figure Description

[0017] Figure 1 A data preprocessing method flow provided in the embodiments of this application Figure 1 ;

[0018] Figure 2 A data preprocessing method flow provided in the embodiments of this application Figure 2 ;

[0019] Figure 3 A data preprocessing method flow provided in the embodiments of this application Figure 3 ;

[0020] Figure 4 A data preprocessing method flow provided in the embodiments of this application Figure 4 ;

[0021] Figure 5 A model architecture diagram provided for an embodiment of this application;

[0022] Figure 6A convolutional network structure diagram provided in an embodiment of this application;

[0023] Figure 7 Statistical charts of data on 8 different biomolecule perforation events provided for embodiments of this application;

[0024] Figure 8 An amplitude-time statistical graph provided for an embodiment of this application;

[0025] Figure 9 An experimental result diagram provided for an embodiment of this application;

[0026] Figure 10 This is a schematic diagram illustrating the structural composition of a data preprocessing apparatus provided in an embodiment of this application;

[0027] Figure 11 This is a schematic diagram of the hardware structure of an electronic device provided in an embodiment of this application. Detailed Implementation

[0028] The technical solutions of the embodiments of this application will now be described with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of this application, and not all of them. All other embodiments obtained by those skilled in the art based on the embodiments of this application without creative effort are within the scope of protection of this application.

[0029] To facilitate understanding of the technical solutions of the embodiments of this application, the relevant technologies of the embodiments of this application are described below. The following relevant technologies are optional solutions and can be combined with the technical solutions of the embodiments of this application in any way, and they all fall within the protection scope of the embodiments of this application.

[0030] In addition, in the embodiments of this application, "first," "second," etc. are used to distinguish similar objects, and are not necessarily used for a specific order or sequence.

[0031] Solid-state nanopores are powerful single-molecule sensing devices used in protein detection, biomolecule folding analysis, and SNP genotyping. A solid-state nanopore sequencing device consists of two liquid-filled containers connected by a nanopore. During the perforation experiment, an electric field is created by applying a voltage to the outside of the containers. This electric field drives biomolecules through the nanopore. As a molecule passes through the nanopore, a time-series current trajectory corresponding to its shape is generated, commonly referred to as a perforation event. Based on the characteristics of the current changes during perforation, certain analytical methods can be used to infer the type of biomolecule that passed through when the signal was generated.

[0032] Currently, the analysis of biomolecular structures using via-pore event data is still in its early stages. It primarily utilizes the characteristics of via-pore events, such as amplitude and relative blocking current, through manual observation and selection of appropriate thresholds for differentiation. However, on the one hand, threshold selection requires specialized knowledge, and the threshold needs to be adjusted accordingly with each change in experimental conditions, making the task require significant human intervention. On the other hand, the amplitude depends on the degree of blockage of the nanopore (the ratio of biomolecule size to nanopore size). When the sizes of the biomolecules to be identified are similar, the characteristics of via-pore events lack significant distinguishability, resulting in a large number of overlapping via signals in the intermediate region that cannot be effectively distinguished.

[0033] To address the existing problems in the technology, this application provides a data preprocessing method. This method uses a sequence to be identified to characterize the structure of biomolecules, determines the initial length of the sequence, standardizes the sequence based on the initial length to obtain a target sequence of uniform length, and then divides the data of the target sequence into multiple clusters using clustering. The data in each cluster are similar. By grouping similar data into the same cluster, the characteristics of the data in the sequence can be clearly understood. Therefore, in subsequent processing, the clusters can improve the accuracy of classifying the sequence and prevent the inability to effectively distinguish biomolecules by using thresholds when their sizes are similar, thus ensuring the classification of the biomolecules.

[0034] The data preprocessing method provided in this application can be applied to any field that requires molecular analysis of biomolecules, including but not limited to biological research, animal research, plant research, medical research, and biological genetics. This application does not impose any specific limitations on this.

[0035] The data preprocessing method provided in this application embodiment can be implemented by a device, which can be applied to electronic devices, including but not limited to laptops, tablets, desktop computers, mobile devices (for medical testing equipment), etc.

[0036] Figure 1 This is a flowchart of a data preprocessing method provided in an embodiment of this application, as shown below. Figure 1 As shown, the data preprocessing method includes:

[0037] Step S110: Determine the initial length of the sequence to be identified; the sequence to be identified is used to characterize the structure of biomolecules.

[0038] In this embodiment of the application, the sequence to be identified refers to a sequence that can characterize the structure of the biomolecule to be identified. The sequence to be identified refers to the time-series current trajectory corresponding to the molecular shape when the biomolecule to be identified passes through the nanopore, which can be called the through-pore event. The sequence to be identified contains multiple data.

[0039] In the embodiments of this application, the initial length of the sequence to be identified refers to the original length of the sequence to be identified. It should be understood that the original length of the sequence to be identified is related to the length of the biomolecule, which is the size of the biomolecule. Different biomolecules have different sizes. Therefore, when different biomolecules pass through the pore event, the initial length of the sequence to be identified will also be different. Different biomolecule structures will also produce different sequences to be identified when passing through the nanopore.

[0040] Biomolecules refer to macromolecules such as proteins, nucleic acids, and carbohydrates that exist within the cells of living organisms. Examples include microRNAs, deoxyribonucleic acid (DNA), glycoproteins, lipoproteins, and nucleoproteins. Different biomolecules have different molecular structures; therefore, the study of biomolecules essentially involves studying their structures to determine their classification.

[0041] The data preprocessing method provided in this application is applicable to the analysis and research of any biomolecule, and this application does not specifically limit the biomolecule.

[0042] Step S120: Based on the initial length, the sequence to be identified is standardized to obtain the target sequence to be identified.

[0043] In the embodiments of this application, standardization processing refers to unifying the initial length of the sequence to be identified. Different biomolecules have different lengths, resulting in different initial lengths of the sequences to be identified. In order to unify the subsequent processing, the initial length of the sequence to be identified is standardized.

[0044] In this embodiment of the application, after determining the initial length of the sequence to be identified, the initial length of the sequence to be identified is unified to obtain target sequences to be identified of the same length.

[0045] For example, the sequence to be identified is Where x i This represents the i-th via event in the original data. Let be the j-th current value in the i-th via event. First, determine the initial length of the sequence to be identified as N. After standardizing the sequence, obtain the target sequence to be identified. The length of the target sequence to be identified is L.

[0046] In this embodiment, the length of the target sequence to be identified can be the optimal value obtained by researchers through multiple experiments, or it can be the average of the initial lengths of multiple sequences to be identified, and the average value can be used as the length of the target sequence to be identified. Alternatively, the initial length covering more than 95% can be set as the length of the target sequence to be identified. This embodiment does not specifically limit the length of the target sequence to be identified.

[0047] In this embodiment of the application, by standardizing the sequence to be identified, the initial length is unified, which not only enables unified processing in the subsequent process, but also reduces training time while retaining information in the sequence to be identified.

[0048] Step S130: Perform clustering processing on multiple data in the target sequence to be identified to obtain multiple cluster sets; the cluster sets are used to identify the category of the sequence to be identified.

[0049] In this embodiment of the application, a cluster set refers to the classification result obtained after classifying multiple data in a target sequence to be identified. Each cluster set includes at least one data, and the data in each cluster set are similar. Grouping similar data into the same cluster set can effectively represent multiple similar data with a single feature.

[0050] In this embodiment of the application, clustering is used to classify multiple data into multiple different cluster sets. In subsequent processing, the cluster sets are used to identify the category of the sequence to be identified.

[0051] In this embodiment, a sequence to be identified for characterizing the structure of a biomolecule is determined, and its initial length is calculated. Based on the initial length, the sequence is standardized to obtain a target sequence of uniform length. Multiple data points in the target sequence are clustered to obtain multiple cluster sets. The data in each cluster set are similar. By standardizing the target sequences of varying lengths, not only can they be uniformly processed in subsequent steps, but the training time can also be reduced while retaining information in the target sequence. At the same time, clustering is used to group similar data in the target sequence into the same cluster set, and a single feature is used to represent multiple similar data points. In this way, by grouping similar data into the same cluster set, the characteristics of the data in the target sequence can be clearly understood. Therefore, in subsequent processing, the cluster set can be used to improve the accuracy of the classification of the target sequence and avoid the inability to effectively distinguish biomolecules by using thresholds when the sizes of the biomolecules are similar, thus preventing the inability to determine the category of the biomolecule.

[0052] In the embodiments of this application, reference is made to Figure 2As shown, step S120, based on the initial length, standardizes the sequence to be identified to obtain the target sequence to be identified, including:

[0053] Step S121: If the initial length is greater than the first threshold, remove the data in the part of the sequence to be identified whose initial length exceeds the first threshold to obtain the target sequence to be identified.

[0054] Step S122: If the initial length is less than the first threshold, then the initial length of the sequence to be identified is padded to be equal to the first threshold to obtain the target sequence to be identified.

[0055] In the embodiments of this application, the initial length refers to the original length of the sequence to be identified. The initial length is related to the size of the biomolecule to be identified. The longer the size of the biomolecule to be identified, the longer the initial length of the sequence to be identified.

[0056] In this embodiment of the application, the first threshold refers to the length of the target sequence to be identified. The first threshold is preset in advance. It can be the optimal value obtained by researchers through multiple experiments, or it can be the average of the initial lengths of multiple sequences to be identified after statistical analysis, and the average value can be used as the length of the target sequence to be identified. Alternatively, the initial length covering more than 95% can be set as the length of the target sequence to be identified. This embodiment of the application does not specifically limit the length of the target sequence to be identified.

[0057] In this embodiment, if the initial length is greater than the first threshold, the initial length of the sequence to be identified needs to be truncated. This involves cutting off the portion of the initial length exceeding the first threshold, so that the length of the truncated sequence is equal to the first threshold. If the initial length is less than the first threshold, the initial length of the sequence to be identified needs to be padded until it equals the first threshold. Here, if the initial length is equal to the first threshold, standardization of the sequence to be identified is not required.

[0058] In this embodiment of the application, the sequence to be identified is truncated or padded by comparing the initial length with the size of the first threshold, so that the length of the sequence to be identified after truncation or padding remains uniform. In this way, it is convenient to uniformly process the sequence to be identified in subsequent processing, thereby improving the classification speed of the sequence to be identified.

[0059] In this embodiment of the application, the step of filling the initial length of the sequence to be identified to be equal to the first threshold to obtain the target sequence to be identified includes: filling the head and / or tail of the sequence to be identified with noise data to obtain the target sequence to be identified.

[0060] In this embodiment of the application, noise refers to an irregular additional signal that does not exist in the original signal after passing through the device and does not change with the change of the original signal. Here, it refers to adding irregular signals that do not exist in the real data. Common noises include salt and pepper noise, Gaussian noise, Poisson noise, etc. This embodiment of the application does not specifically limit the type of noise data to be filled.

[0061] In this embodiment of the application, if the initial length of the sequence to be identified is less than the first threshold, the sequence to be identified needs to be filled. By filling the sequence to be identified with non-existent and irregular data, the target sequence to be identified is obtained.

[0062] In this embodiment, padding can be performed only at the beginning of the sequence to be identified, only at the end of the sequence to be identified, or partially at both the beginning and end of the sequence to be identified. The length of the padded sequence to be identified is equal to the first threshold.

[0063] By filling the head and / or tail of the sequence to be identified with noise points, a baseline is provided for subsequent processing.

[0064] In this embodiment of the application, the amount of noise data filled at the beginning and end of the sequence to be identified is equal or differs by 1.

[0065] In this embodiment, if padding is performed only at the beginning of the sequence to be identified or only at the end of the sequence to be identified, the amount of noise data to be filled is the difference between the first threshold and the sequence to be identified; if padding is performed at both the beginning and the end of the sequence to be identified, when the difference between the first threshold and the sequence to be identified is even, the amount of noise data to be filled at the beginning and the end is half of the difference, and when the difference between the first threshold and the sequence to be identified is odd, the amount of noise data to be filled at the beginning and the end differs by 1.

[0066] For example, the sequence to be identified is The initial length of the sequence to be identified is N, and the length of the target sequence to be identified is L. After padding the beginning and end of the sequence to be identified with noise data, the target sequence to be identified is obtained as follows:

[0067] In this embodiment of the application, the same noise data is filled at the beginning and end of the sequence to be identified to ensure the integrity of the original sequence to be identified and to avoid inaccurate identification of biomolecule categories due to the filling of noise data in subsequent processing.

[0068] In this embodiment of the application, the valid data in the sequence to be identified and the noise data are encoded using different types of data.

[0069] In this embodiment of the application, valid data refers to the original data in the sequence to be identified before padding, while noise data is the data filled into the sequence to be identified.

[0070] In this application embodiment, data encoding refers to digitizing the data in the sequence to be identified, making it a carrier of information. In order to transmit the data correctly, it is necessary to encode the data in the sequence to be identified.

[0071] In the embodiments of this application, different numbers can be used to distinguish between valid data and noise data in the target sequence to be identified. For example, "0" can be used to encode valid data and "1" to encode noise data, or "1" can be used to encode valid data and "0" to encode noise data. Other different numbers can also be used, or other methods can be used to distinguish between valid data and noise data in the sequence to be identified. This application does not specifically limit the type of data encoding.

[0072] For example, different data types are encoded as T i = {0,...0,1...,1,0,...,0}, where the number of 0s is the same as the number of noisy data in the target sequence to be identified, which is LN, and the number of 1s is the same as the number of valid data in the target sequence to be identified, which is N.

[0073] In this embodiment of the application, the sequence to be identified is standardized to obtain a target sequence to be identified with uniform length. By using different data type encodings, the effective data and the padded noise data in the sequence to be identified are distinguished, which is beneficial to the identification of the effective data in the sequence to be identified in subsequent processing, thereby accurately identifying the category of the sequence to be identified.

[0074] In the embodiments of this application, reference is made to Figure 3 As shown, in step S130, multiple data points in the target sequence to be identified are clustered to obtain multiple cluster sets, including:

[0075] Step S131: Calculate the first distance between the first data in the sequence to be identified and the other data in the sequence to be identified; if the first distance is less than the clustering threshold, then the first data and the data whose first distance is less than the clustering threshold are in the first cluster set; the clustering threshold is the average value of the noise data.

[0076] In this embodiment of the application, the first distance characterizes the similarity between the first data in the sequence to be identified and other data. The larger the first distance, the smaller the similarity between the two data; the smaller the first distance, the greater the similarity between the two data.

[0077] Alternatively, the distance can be calculated using Euclidean distance, Hamming distance, cosine distance, etc. This application does not specifically limit the method for calculating the first distance.

[0078] In this embodiment of the application, the clustering threshold is the average value of the noisy data. If the first distance is less than the clustering threshold, it means that the two data are very similar and can be classified into the same cluster set, thus obtaining the first cluster set. If the first distance is greater than the clustering threshold, it means that the two data are significantly different and do not belong to the same cluster set.

[0079] Step S132: Next, calculate the second distance between the first data in the remaining data of the sequence to be identified and the other data in the remaining data; if the second distance is less than the clustering threshold, then the first data in the remaining data and the data whose second distance is less than the clustering threshold are in the first cluster set.

[0080] In this embodiment of the application, the second distance represents the distance between the first data and other data in the remaining data. The larger the second distance, the smaller the similarity between the two data; the smaller the second distance, the larger the similarity between the two data.

[0081] Alternatively, the distance can be calculated using Euclidean distance, Hamming distance, cosine distance, etc. This application does not specifically limit the method for calculating the first distance.

[0082] In this embodiment of the application, if the second distance is less than the clustering threshold, it means that the first data in the remaining data is very similar to the data corresponding to the second distance and can be classified into the same cluster set to obtain the second cluster set; if the second distance is greater than the clustering threshold, it means that the first data in the remaining data is very different from the data corresponding to the second distance and does not belong to the same cluster set.

[0083] Step S133: Continue in this manner until the last data in the sequence to be identified is in the cluster set.

[0084] In this embodiment of the application, after obtaining the second cluster set, the third distance between the first data in the data not in the cluster set and other data is calculated. If the third distance is less than the clustering threshold, it means that the first data in the data not in the cluster set is similar to the data corresponding to the third distance and is divided into the same cluster set, thus obtaining the third cluster set. Then the distance between the remaining data is calculated, and the above operation is repeated until the last data is divided into a cluster set.

[0085] In this embodiment of the application, clustering is used to classify multiple data in the target sequence to be identified, resulting in multiple cluster sets. The clustering process involves taking the first data as the cluster center and calculating the distance between the other data and the cluster center. If the distance is less than the clustering threshold, it means that they belong to the same cluster set. Data with a distance less than the clustering threshold are all assigned to the same cluster set. Then, the distance between the first data in the data not in the cluster set and other data is calculated. The above operation is repeated until the last data is assigned to a cluster set.

[0086] By using the clustering method described above to cluster the data, the temporal sequence and continuity of the data in the sequence to be identified can be well guaranteed, and the integrity of the sequence to be identified can be guaranteed. In subsequent processing, using this cluster set can improve the accuracy of the sequence identification.

[0087] In this embodiment of the application, the minimum number of the plurality of cluster sets is 2, and the maximum number is related to the biomolecular structure.

[0088] In this embodiment of the application, the minimum number refers to the lower bound of the clustering. After the sequence to be identified is filled with noise data, the sequence to be identified contains at least two types of data: the filled noise data and the valid data. After the data is clustered, it is divided into at least two categories. That is to say, the minimum number of cluster sets is 2.

[0089] In this embodiment of the application, the biomolecular structure is characterized by the sequence to be identified. The size of the biomolecular can be seen from the biomolecular structure. The biomolecular needs to pass through a nanopore to obtain the sequence to be identified. The pore size of the nanopore needs to be determined. Therefore, the maximum number of cluster sets is related to the biomolecular structure.

[0090] In this embodiment of the application, by determining the minimum and maximum number of cluster sets, it is possible to avoid the situation where, during clustering, the data between cluster sets is similar due to improper partitioning, or the data between cluster sets is very different, which would result in uneven data partitioning between the target sequences to be identified and affect the accuracy of subsequent sequence identification.

[0091] In the embodiments of this application, reference is made to Figure 4 As shown, the method further includes:

[0092] S140: Input the multiple cluster sets into a pre-trained classification model to obtain the category of the sequence to be identified; wherein, the pre-trained classification model is used to determine the category of the sequence to be identified.

[0093] In this embodiment of the application, the cluster set is obtained after preprocessing the sequence to be identified to characterize the structure of biomolecules. The data in each cluster set are similar. Therefore, the cluster set can characterize the features of multiple similar data.

[0094] In this embodiment of the application, the pre-trained classification model is a pre-trained biomolecular structure processing model. Here, the pre-trained classification model can be obtained by the electronic device training on sample data, or it can be obtained by the electronic device from other servers that provide models.

[0095] Here, the pre-trained classification model is used to determine the category of the sequence to be identified. The cluster set obtained after data preprocessing is input into the pre-trained classification model. The pre-trained classification model is a neural network model. Furthermore, the category of the sequence to be identified can be obtained through the processing of the pre-trained classification model.

[0096] In this embodiment of the application, multiple data in the target sequence to be identified are classified by clustering to obtain multiple cluster sets. After obtaining multiple cluster sets, the multiple cluster sets are input into a pre-trained classification model to obtain the category of the sequence to be identified.

[0097] In this embodiment of the application, the category of a sequence to be identified can be determined by performing a single operation on a pre-trained classification model. In this way, the amount of computation can be reduced and the identification speed can be improved during the identification process of the sequence to be identified.

[0098] In this embodiment of the application, the method further includes:

[0099] Based on the multiple cluster sets, the signal-to-noise ratio of the multiple cluster sets is determined;

[0100] If at least two cluster sets have a signal-to-noise ratio less than the second threshold, then the cluster sets with a signal-to-noise ratio less than the second threshold are adjusted to obtain the adjusted cluster sets.

[0101] The adjusted cluster set is input into the pre-trained classification model to obtain the category of the sequence to be identified; wherein, the pre-trained classification model is used to determine the category of the sequence to be identified.

[0102] In the embodiments of this application, the signal-to-noise ratio refers to the ratio of signal to noise in an electronic device. Here, it refers to the ratio of the cluster set to the noise value in the sequence to be identified. When the signal-to-noise ratio is too small, it is difficult to distinguish between two cluster sets. Generally, when the signal-to-noise ratio is greater than or equal to 1.5, it can be fully guaranteed that there is a significant difference between the two cluster sets. When it is less than 1.2, it is difficult to distinguish them.

[0103] Here, the signal of each cluster can be characterized by the features of the cluster set. The features of the cluster set can be the amplitude of the cluster set (the difference between the current value and the baseline current value) or the relative blocking current, or other features. This application does not specifically limit the features.

[0104] In the embodiments of this application, the second threshold refers to a preset signal-to-noise ratio threshold. The signal-to-noise ratio threshold can be the optimal value obtained after multiple experiments, or it can be a value preset by researchers based on their professional knowledge. For example, if it is difficult to distinguish at 1.2, the second threshold can be set to 1.2. This application does not specifically limit the second threshold.

[0105] In this embodiment, after preprocessing the sequence to be identified, which characterizes the biomolecular structure, cluster sets are obtained. The features of each cluster are calculated, and the ratio of the features to the noise value of each cluster set is determined to obtain the signal-to-noise ratio (SNR). If two or more SNRs are less than a second threshold, it indicates that the cluster sets with SNRs less than the second threshold are not significantly different and difficult to distinguish. The cluster sets with SNRs less than the second threshold are adjusted to obtain an adjusted cluster set. The adjusted cluster set is then input into a pre-trained classification model, which outputs the category of the sequence to be identified. By determining the SNR of the cluster sets and adjusting the cluster sets with SNRs less than the second threshold, the adjusted cluster set is obtained.

[0106] In this embodiment of the application, the adjusted cluster set is input into the pre-trained classification module to obtain the category of the sequence to be identified. By adjusting the cluster set, the cluster sets are made to have obvious differences, which makes it easier to distinguish the cluster sets and further improves the accuracy of the sequence to be identified.

[0107] In this embodiment of the application, adjusting the cluster set whose signal-to-noise ratio is less than a second threshold to obtain the adjusted cluster set includes:

[0108] The cluster set whose signal-to-noise ratio is less than the second threshold is merged with the cluster set whose signal-to-noise ratio is greater in the two adjacent cluster sets to obtain the adjusted cluster set.

[0109] In this embodiment of the application, multiple data in the target sequence to be identified are classified by clustering to obtain a cluster set. The cluster set with a signal-to-noise ratio less than a second threshold is merged with the cluster set with a larger signal-to-noise ratio in the two adjacent cluster sets to obtain an adjusted cluster set.

[0110] By merging clusters with low signal-to-noise ratios with the clusters with high signal-to-noise ratios in their two neighboring clusters, the adjusted clusters are easier to distinguish, thereby improving the accuracy of the classification of the sequence to be identified.

[0111] In this embodiment of the application, the step of inputting the plurality of cluster sets into a pre-trained classification model to obtain the category of the sequence to be identified includes:

[0112] The first network of the pre-trained classification model is invoked to process the multiple cluster sets to obtain a first feature; the first feature includes local features of each cluster set in the multiple clustering results;

[0113] In the embodiments of this application, the first network refers to a convolutional neural network (CNN) that extracts local features of cluster sets from a pre-trained classification model. The CNN network includes convolutional layers, activation layers, and pooling layers.

[0114] In this embodiment, a feature represents the attribute or characteristic of a sample. The first feature is the output of the first network and is a set of local features. Each local feature is an attribute vector representing each cluster set.

[0115] In this application embodiment, the first network can be any convolutional neural network that extracts features from the cluster set. The backbone network in the structure of the convolutional neural network includes VGG, ResNet, LeNet, etc. This application does not specifically limit the backbone network and the number of convolutional layers of the first network.

[0116] In this embodiment of the application, a convolutional neural network is used to extract local features from multiple cluster sets to obtain the first feature.

[0117] The first feature is input into the second network of the pre-trained classification model to obtain the second feature; the second feature includes the feature after the local feature enhancement time sequence.

[0118] In the embodiments of this application, the second network refers to the network in the pre-trained classification model that can perform temporal enhancement of local features.

[0119] In the embodiments of this application, the second network can be any network that enhances temporal features, such as a bidirectional recurrent neural network, a temporal network, etc. This application does not specifically limit the structure of the second network.

[0120] In this embodiment of the application, the second feature is the output of the second network, which has temporal characteristics. Biomolecules have a temporal order when passing through nanopores. Therefore, the sequence to be identified has temporal characteristics, and the cluster set has certain temporal characteristics from front to back. The second feature is obtained by extracting the temporal characteristics in the cluster set through the second network.

[0121] The second feature is input into the third network of the pre-trained classification model to obtain the third feature; the third feature is used to characterize the differences between the features after the augmentation time series.

[0122] In the embodiments of this application, the third network refers to the network in the pre-trained classification model that enhances the differences between features.

[0123] In the embodiments of this application, the third network can be any network that enhances the differences between features, such as an attention mechanism network, a feature pyramid network (FPN), etc. This application does not specifically limit the structure of the third network.

[0124] In this embodiment of the application, the third feature is the output of the third network, which is also a feature vector. It is a feature vector after enhancing the difference of the second feature. The feature with enhanced difference can effectively distinguish the features of the cluster set.

[0125] Based on the third feature, the category of the sequence to be identified is determined.

[0126] In this application embodiment, the third feature is essentially a raw embedding vector. The inner product is calculated using the raw embedding vector obtained by the third network, and the final category of the biomolecule is obtained through a classification function.

[0127] In this embodiment, the classification function is used to calculate the possible probability of each class of molecules. Classification functions include the softmax function, support vector machine, etc. This embodiment does not specifically limit the classification function.

[0128] In this embodiment, after preprocessing the sequence to be identified, multiple cluster sets are obtained. These cluster sets are then input into a pre-trained classification model. The first network (convolutional neural network) of the pre-trained classification model extracts local features from the multiple cluster sets. Next, a second network processes the local features to obtain global features with enhanced temporal characteristics. Then, a third network processes the global features to enhance the diversity of the cluster set features, obtaining a third feature. Finally, a classification function is used to process the third feature to obtain the possible probabilities of the sequence to be identified, that is, the probabilities of the possible categories of the biomolecules, thereby determining the category of the sequence to be identified.

[0129] In this embodiment of the application, three different basic networks are used to extract different semantic features, and these features are combined to improve model performance and increase the accuracy of sequence recognition.

[0130] In this embodiment of the application, the verification process of the pre-trained classification model includes:

[0131] The sequences to be identified in the validation set are input into the pre-trained classification model, which outputs the category of the sequences to be identified; the validation set is used to verify the classification effect of the pre-trained classification model.

[0132] In this embodiment of the application, the validation set refers to a dataset used to assist in debugging. The validation set is used during the training process. Generally, after several batches of training are completed, the validation set is used to check the effect so that problems with the model or parameters can be found. Training can be terminated in time, and parameters or the model can be readjusted without waiting until the training is completed, thus avoiding wasting time.

[0133] The data preprocessing methods provided in the embodiments of this application will be described in detail below with reference to specific application scenarios.

[0134] The data preprocessing method provided in this application is a method for analyzing solid-state nanopore perforation data using a neural network model. First, by analyzing the data, signals of varying lengths are unified. Simultaneously, clustering is used to fragment the original data, obtaining features within each fragment. Based on the features of the preprocessed original data, a novel neural network is proposed, utilizing the characteristics of different base models to form the final network architecture. Finally, to verify the method's performance, a comprehensive analysis is performed on the raw sequencing data obtained through solid-state nanopores from biomolecules with publicly disclosed structures.

[0135] In this application scenario, specific implementation methods can include three aspects: data preprocessing, classification model, experiments and results.

[0136] First, data preprocessing: signals of varying lengths are standardized, and clustering is used to fragment the original data to obtain features within each fragment.

[0137] It should be understood that the signal is the sequence to be identified in the embodiments of this application, the original data is the data in the target sequence to be identified in the embodiments of this application, the fragmentation processing is the clustering processing in the embodiments of this application, and the fragment is the cluster set in the embodiments of this application.

[0138] Data preprocessing includes the following steps:

[0139] (1) Filling the original via events: In practical applications, there may be multiple biomolecules to be identified. Therefore, the real data may include multiple sequences to be identified. The input of the neural network is the current trajectory from a via event. Since different molecules have different lengths, the length of the original data obtained is different. The subsequent model input cannot be uniformly processed. Therefore, it is necessary to unify the length of the data obtained in the experiment.

[0140] Define the data to be obtained as x = {x1, ..., x2} i ,...,x M}, where X represents the number of sequences to be identified. Where x iThis represents the i-th via event (sequence to be identified) in the acquired data. Let be the j-th current value in the i-th sequence to be identified. First, calculate the original data length l = {len(x1),...,len(x2)} for each sequence to be identified. i ),...,len(x M )}={l1,...,l i ,...,l M In this application scenario, the length of the training samples was statistically analyzed, and the final via event length L was set to cover more than 95% of the training samples, in order to reduce training time while preserving as much data information as possible.

[0141] Therefore, regarding the data When N>=L, the data can be truncated, that is... For values ​​less than L, padding is performed using the following strategy: P = (LN) / 2 points are added before and after the original data. Then, the data without molecules passing through is fitted. In most cases, the noise data follows a Gaussian distribution with a mean of μ and a variance of δ. Therefore, P noise points are added before and after each iteration. Simultaneously, these pre- and post-padding noise points provide a baseline for subsequent multi-level feature extraction.

[0142] (2) Data fragmentation representation: For the data x processed in step (1) i Clustering is performed, and the clustering threshold d is set to the average value of the noise mentioned above. Given the biomolecule size as d1 and the nanopore diameter as d2, the maximum level is n = [d1 / d2] + 2, where 2 represents the noise points filled in before and after step (1). Therefore, the maximum number of clusters is set to n, and the boundary points of the clustering results are obtained. These boundary points are the initial segmentation positions of each level. The amplitude a = {a1,...,a1} within each segment is calculated. n}

[0143] It should be understood that data fragmentation is the process of clustering multiple data in the target sequence to be identified in this application embodiment to obtain multiple cluster sets.

[0144] To prevent over-segmentation of data points, the ratio of amplitude to search radius, f = a / d, is calculated. When there are two connected segment segments with a ratio less than 1.2, they are merged with the segment with the larger ratio from the two adjacent segments to form m segments, thus obtaining the final segmentation position and the amplitude within the segment.

[0145] It should be understood that the search radius is the average noise value, and the ratio of the amplitude to the search radius is the signal-to-noise ratio in the embodiments of this application.

[0146] (3) Input data construction: using sample x iFor example, in step (1), the data length is standardized. In order to distinguish the original current data from the filled noise data, it is also necessary to encode them with different data types. The position of the data point that is the original current is represented by 1, and the rest are represented by 0.

[0147] Then in step (2) x i After clustering, split point merging, and other processing, it can be represented as in Let S be the segmentation point location, and S be the segment embedding representation. An index count is introduced for each segment, and finally each data point can be represented as (segment index, segment magnitude).

[0148] The model's input format can ultimately be represented as follows: Original data: For x i After performing fill or stage operations: Embedded representation of different data types: T i ={0,...0,1...,1,0,...,0}; Data fragmentation representation:

[0149] Classification Model: Based on the features of the preprocessed raw data, a new neural network model is proposed: First, an embedding layer is introduced into the preprocessed data. Then, three different basic network structures are used to extract different semantic features and combine them to improve the model performance.

[0150] like Figure 5 As shown. Figure 5 It contains four components: (a) Input layer: preprocesses the raw data to obtain data point index T, hierarchical fragment information M, data type characteristics S, and current signal X; (b) Embedding layer: uses a linear expression layer to vectorize the data type and fragments; (c) Feature extraction layer: constructs a convolutional network (convolutional block) to extract local features, constructs a temporal recurrent neural network (RNN) to extract global features, and uses an attention mechanism network (Attention Netsork) to distinguish features; (d) Prediction layer: uses a softmax layer to predict the probability that each event belongs to a biomolecular category. Figure 5 In i1, i2...i c This represents the probability that each biomolecule belongs to a certain category, given that there are a total of C categories.

[0151] Input layer: using preprocessed data: current signal X, data point index T, data type S, and hierarchical fragment information M;

[0152] Embedding Layer: After processing the original current signal, three new types of data information are added: data point index T, which adds time-series information to each data point; data type S, which distinguishes the type information of the original data and the filled data; and hierarchical fragment information M, which is the fragment information of the data hierarchically divided. To represent this, three embedding layers are used. The role of the embedding layer is to change the linear expression into a non-linear expression, introduce implicit feature expressions, and perform embedding expressions on them separately. Simultaneously, a non-linear expression layer is added after the type information and fragment information to non-linearly express the features.

[0153] Feature layer: This layer mainly extracts features from the current signal. Currently, common feature extractors are divided into three categories: convolutional networks that handle strong local correlations; recurrent neural networks that handle strong temporal relationships; and attention networks that handle differential characteristics.

[0154] Based on the data characteristics, each layer represents the structural characteristics of a biomolecule within a certain range. Combining all extracted fragment features yields the final characteristics of the biomolecule. Considering this, a convolutional network is first used to extract local features within each layer. Since the layers exhibit temporal characteristics from beginning to end, and the front and back ends of the molecule cannot remain consistent during pore extraction, a bidirectional recurrent neural network is used to enhance its temporal features. Furthermore, considering the model's fault tolerance, an attention network is used in the last layer of the feature extraction layer to distinguish layer features.

[0155] Specifically, given a current signal, local feature extraction is first performed using multiple convolutional blocks. Each convolutional block consists of, for example,... Figure 6 As shown, it includes two one-dimensional convolutional networks. A convolutional network consists of: convolutional layers, batch normalization, and activation layers. Figure 6 The activation function in the middle activation layer is the ReLU activation function, which the convolutional network uses to extract features; the last layer of the convolutional block is max pooling, which is used for downsampling, increasing the receptive field, and enhancing the robustness of feature locations.

[0156] A convolutional network (Conv Layer) can be represented by the following formula (1):

[0157] Conv(x)=Relu(BatchNormal(CNN(X))) (1)

[0158] Where x represents the current signal, Conv(x) represents the convolution operation on x, ReLU is the activation function, which represents mapping a linear function to a nonlinear function, BatchNormal represents the standardization of the feature vector, and CNN(X) represents the feature extraction of the data.

[0159] A convolutional layer block (CLB) is defined as follows (2):

[0160] CLB(X)=MaxPool(Conv(Conv(x))) (2)

[0161] Where CLB(X) represents multiple convolutional networks, MaxPool represents downsampling the features extracted from multiple convolutional layers, and Conv(Conv(x)) represents performing multiple convolutions.

[0162] During data preprocessing, clustering is used to infer that the data contains approximately m segments. Considering the inherent errors in clustering methods, the dimension of the data extracted from the last convolutional block is defined as the final data dimension. Where l represents the segment error, d is the dimension of each segment feature, and l and d are hyperparameters that can be preset. Considering that the features extracted by the convolutional network only take into account locality, a temporal network is used to further enhance them, that is, to enhance the temporal information of each segment feature. The output data dimension is... To further differentiate the enhanced fragment features, an attention network is introduced for autonomous learning:

[0163]

[0164] Where X is the current signal, q is the eigenvector, and a i The value of i represents the importance of each feature segment, ranging from 1 to m+l. Therefore, the final output dimension of the feature extraction layer is...

[0165] Prediction layer: In the output layer, the original embedding vectors obtained from the previous feature layer are used to perform inner product calculation, and a softmax function is used to obtain the final probability of each class of biomolecules.

[0166] Experiments and Results: Experiments were conducted on real datasets and compared with other methods to verify the reliability of the method provided in the embodiments of this application.

[0167] The experiment was conducted in a conical nanopore with a diameter of 14 ± 3 nm, and the effective sensing length of the nanopore was approximately 200 nm. A voltage of -600 mV was applied across the nanopore to drive eight types of biomolecules through the solid nanopore, and a total of 6239 through-pore events were detected. Statistical results for each type are as follows: Figure 7As shown: the frequency of biomolecules in category 1 is approximately 1100 times, the frequency of biomolecules in category 2 is approximately 500 times, the frequency of biomolecules in category 3 is approximately 500 times, the frequency of biomolecules in category 4 is approximately 600 times, the frequency of biomolecules in category 5 is approximately 150 times, the frequency of biomolecules in category 6 is approximately 200 times, the frequency of biomolecules in category 7 is approximately 1550 times, and the frequency of biomolecules in category 8 is approximately 1500 times.

[0168] Figure 8 The scatter plot of "amplitude-via time" for eight biomolecules shows that a large number of repeated via events converge in the middle elliptical interval. By using only the two features of amplitude and duration to distinguish different types of biomolecules, it can be seen that when a large number of via events are repeated, amplitude and time alone cannot effectively distinguish them.

[0169] To evaluate the model's performance, the data was randomly divided into training and test sets in an 8:2 ratio, and 5-fold cross-validation was performed. The model was then compared with other existing models. Due to significant class differences among the data, four metrics—accuracy, precision, recall, and F1 score—were used for comprehensive evaluation. The experimental results are the average of the 5-fold cross-validation results.

[0170] Precision is the proportion of correctly classified samples out of the total number of samples. Accuracy is the proportion of samples that were predicted to be positive but were also actually positive out of the samples that were predicted to be positive. Recall is the proportion of samples that were actually positive but were also predicted to be positive out of the samples that were actually positive. The F1 score is the harmonic mean of precision and recall.

[0171] Experimental results are as follows Figure 9 As shown, compared with other existing models such as Naive Bayes (NB), Logistic Regression (LR), Temporal Convolutional Network (TCN), and SERSNET module, the method provided in this application embodiment has the highest four evaluation metrics.

[0172] This application first proposes a method for preprocessing raw signals: unifying signals of varying lengths and fragmenting the raw data using clustering to obtain features within each fragment. Based on the features obtained after preprocessing the raw data, a novel neural network model is proposed: first, an embedding layer is introduced into the preprocessed data; then, three different basic network structures are used to extract different semantic features, and these features are combined to improve model performance. Experiments are conducted on real-world datasets, and comparisons with other methods are made to verify the reliability of the method proposed in this application.

[0173] The preferred embodiments of this application have been described in detail above with reference to the accompanying drawings. However, this application is not limited to the specific details of the above embodiments. Within the scope of the technical concept of this application, various simple modifications can be made to the technical solutions of this application, and these simple modifications all fall within the protection scope of this application. For example, the various specific technical features described in the above specific embodiments can be combined in any suitable manner without contradiction. To avoid unnecessary repetition, this application will not describe the various possible combinations separately. Furthermore, various different embodiments of this application can also be arbitrarily combined, as long as they do not violate the spirit of this application, they should also be considered as the content disclosed in this application. Moreover, without conflict, the various embodiments and / or the technical features in the various embodiments described in this application can be arbitrarily combined with the prior art, and the resulting technical solutions should also fall within the protection scope of this application.

[0174] This application also provides a data preprocessing apparatus, referencing... Figure 10 As shown, the data preprocessing apparatus includes:

[0175] The determining unit 101 is used to determine the initial length of the sequence to be identified; the sequence to be identified is used to characterize the structure of a biomolecule.

[0176] The first processing unit 102 is used to perform standardization processing on the sequence to be identified based on the initial length to obtain the target sequence to be identified;

[0177] The second processing unit 103 is used to perform clustering processing on multiple data in the target sequence to be identified, to obtain multiple cluster sets; the cluster sets are used to identify the category of the sequence to be identified.

[0178] In this embodiment of the application, the first processing unit is further configured to, if the initial length is greater than a first threshold, remove the data portion of the initial length of the sequence to be identified that exceeds the first threshold to obtain the target sequence to be identified; if the initial length is less than the first threshold, fill the initial length of the sequence to be identified to be equal to the first threshold to obtain the target sequence to be identified.

[0179] In an embodiment of this application, the first processing unit is further configured to fill the head and / or tail of the sequence to be identified with noise data to obtain the target sequence to be identified.

[0180] In this embodiment of the application, the amount of noise data filled at the beginning and end of the sequence to be identified is equal or differs by 1.

[0181] In this embodiment of the application, the valid data in the sequence to be identified and the noise data are encoded using different types of data.

[0182] In this embodiment, the second processing unit is further configured to calculate a first distance between the first data in the sequence to be identified and other data in the sequence; if the first distance is less than a clustering threshold, the first data and the data whose first distance is less than the clustering threshold are in a first cluster set; the clustering threshold is the average value of the noise data; then, a second distance is calculated between the first data in the remaining data in the sequence to be identified and other data in the remaining data; if the second distance is less than the clustering threshold, the first data in the remaining data and the data whose second distance is less than the clustering threshold are in a second cluster set; and so on, until the last data in the sequence to be identified is in a cluster set.

[0183] In this embodiment of the application, the minimum number of the plurality of cluster sets is 2, and the maximum number is related to the biomolecular structure.

[0184] In this embodiment of the application, the second processing unit is further configured to determine the signal-to-noise ratio (SNR) of the plurality of cluster sets based on the plurality of cluster sets; if the SNR of at least two cluster sets is less than a second threshold, then adjust the cluster sets whose SNR is less than the second threshold to obtain an adjusted cluster set; input the adjusted cluster set into the pre-trained classification model to obtain the category of the sequence to be identified; wherein, the pre-trained classification model is used to determine the category of the sequence to be identified.

[0185] In this embodiment of the application, the second processing unit is further configured to merge the cluster set whose signal-to-noise ratio is less than the second threshold with the cluster set whose signal-to-noise ratio is greater than the two adjacent cluster sets to obtain the adjusted cluster set.

[0186] In this embodiment, the second processing unit is further configured to invoke the first network of the pre-trained classification model to process the plurality of cluster sets to obtain a first feature; the first feature includes local features of each cluster set in the plurality of cluster sets; input the first feature into the second network of the pre-trained classification model to obtain a second feature; the second feature includes the feature after time-series enhancement of the local features; input the second feature into the third network of the pre-trained classification model to obtain a third feature; the third feature is used to characterize the differences between the features after time-series enhancement; and determine the category of the sequence to be identified based on the third feature.

[0187] In this embodiment of the application, the second processing unit is further configured to input the sequence to be identified from the validation set into the pre-trained classification model and output the category of the sequence to be identified; the validation set is used to verify the classification effect of the pre-trained classification model.

[0188] Of course, in practical applications, such as Figure 11 As shown, the various components in this electronic device are coupled together via a bus system 112. It is understood that the bus system 112 is used to enable communication between these components. In addition to a data bus, the bus system 102 also includes a power bus, a control bus, and a status signal bus. However, for clarity, in... Figure 11 All buses are labeled as Bus System 112.

[0189] It is understood that the memory in this embodiment can be volatile memory or non-volatile memory, or both. Specifically, non-volatile memory can be read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), magnetic random access memory (FRAM), flash memory, magnetic surface memory, optical disc, or compact disc read-only memory (CD-ROM); magnetic surface memory can be disk storage or magnetic tape storage. Volatile memory can be random access memory (RAM), which is used as an external cache. By way of example, but not limitation, many forms of RAM are available, such as Static Random Access Memory (SRAM), Synchronous Static Random Access Memory (SSRAM), Dynamic Random Access Memory (DRAM), Synchronous Dynamic Random Access Memory (SDRAM), Double Data Rate Synchronous Dynamic Random Access Memory (DDRSDRAM), Enhanced Synchronous Dynamic Random Access Memory (ESDRAM), SyncLink Dynamic Random Access Memory (SLDRAM), and Direct Rambus Random Access Memory (DRRAM).The memories described in the embodiments of this application are intended to include, but are not limited to, these and any other suitable types of memories.

[0190] The methods disclosed in the embodiments of this application can be applied to a processor or implemented by a processor. A processor may be an integrated circuit chip with signal processing capabilities. During implementation, each step of the above methods can be completed by integrated logic circuits in the processor's hardware or by instructions in software form. The processor may be a general-purpose processor, a DSP, or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The processor can implement or execute the methods, steps, and logic block diagrams disclosed in the embodiments of this application. A general-purpose processor may be a microprocessor or any conventional processor, etc. The steps of the methods disclosed in the embodiments of this application can be directly manifested as execution by a hardware decoding processor, or execution by a combination of hardware and software modules in the decoding processor. The software modules may be located in a storage medium, which is located in memory. The processor reads information from the memory and, in conjunction with its hardware, completes the steps of the aforementioned methods.

[0191] This application also provides a computer storage medium, specifically a computer-readable storage medium. It stores computer instructions thereon. As a first implementation, when the computer storage medium is located in an electronic device, these computer instructions, when executed by a processor, implement any step in the data preprocessing method described above in this application.

[0192] In the several embodiments provided in this application, it should be understood that the disclosed devices and methods can be implemented in other ways. The device embodiments described above are merely illustrative. For example, the division of units is only a logical functional division, and in actual implementation, there may be other division methods, such as: multiple units or components can be combined, or integrated into another system, or some features can be ignored or not executed. In addition, the coupling, direct coupling, or communication connection between the various components shown or discussed can be through some interfaces, and the indirect coupling or communication connection between devices or units can be electrical, mechanical, or other forms.

[0193] The units described above as separate components may or may not be physically separate. The components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of the units may be selected to achieve the purpose of this embodiment according to actual needs.

[0194] In addition, each functional unit in the various embodiments of this application can be integrated into one processing unit, or each unit can be a separate unit, or at least two units can be integrated into one unit; the integrated unit can be implemented in hardware or in the form of hardware plus software functional units.

[0195] Those skilled in the art will understand that all or part of the steps of the above method embodiments can be implemented by hardware related to program instructions. The aforementioned program can be stored in a computer-readable storage medium. When the program is executed, it performs the steps of the above method embodiments. The aforementioned storage medium includes various media that can store program code, such as mobile storage devices, ROM, RAM, magnetic disks, or optical disks.

[0196] Alternatively, if the integrated units described above are implemented as software functional modules and sold or used as independent products, they can also be stored in a computer-readable storage medium. Based on this understanding, the technical solutions of the embodiments of this application, or the parts that contribute to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as mobile storage devices, ROM, RAM, magnetic disks, or optical disks.

[0197] It should be noted that the technical solutions described in the embodiments of this application can be combined arbitrarily without conflict.

[0198] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.

Claims

1. A data preprocessing method, characterized in that, include: Determine the initial length of the sequence to be identified; The sequence to be identified is used to characterize the structure of biomolecules; The sequence to be identified is a time-series current trajectory generated by the biomolecule to be identified through a nanopore. Based on the initial length, the sequence to be identified is standardized to obtain the target sequence to be identified; Multiple data points in the target sequence to be identified are clustered to obtain multiple cluster sets; the cluster sets are used to identify the category of the sequence to be identified. The process of standardizing the sequence to be identified based on the initial length to obtain the target sequence to be identified includes: If the initial length is greater than the first threshold, then remove the data in the part of the sequence to be identified whose initial length exceeds the first threshold to obtain the target sequence to be identified; If the initial length is less than the first threshold, noise data is filled into the head and / or tail of the sequence to be identified to obtain the target sequence to be identified with a length equal to the first threshold.

2. The method according to claim 1, characterized in that, The amount of noise data filled at the beginning and end of the sequence to be identified is equal or differs by 1.

3. The method according to claim 1, characterized in that, The valid data in the sequence to be identified and the noise data are encoded using different types of data.

4. The method according to claim 1, characterized in that, The clustering process is performed on multiple data points in the target sequence to be identified, resulting in multiple cluster sets, including: Calculate a first distance between the first data in the sequence to be identified and other data in the sequence to be identified; if the first distance is less than a clustering threshold, then the first data and the data whose first distance is less than the clustering threshold are in the first cluster set; the clustering threshold is the average value of the noise data; Next, calculate the second distance between the first data in the remaining data of the sequence to be identified and the other data in the remaining data; if the second distance is less than the clustering threshold, then the first data in the remaining data and the data whose second distance is less than the clustering threshold are in the second cluster set; This process continues until the last data point in the sequence to be identified is in the cluster set.

5. The method according to any one of claims 1 to 4, characterized in that, The minimum number of the multiple cluster sets is 2, and the maximum number is related to the biomolecular structure.

6. The method according to any one of claims 1 to 4, characterized in that, The method further includes: The multiple cluster sets are input into a pre-trained classification model to obtain the category of the sequence to be identified; wherein, the pre-trained classification model is used to determine the category of the sequence to be identified.

7. The method according to any one of claims 1 to 4, characterized in that, The method further includes: Based on the multiple cluster sets, the signal-to-noise ratio of the multiple cluster sets is determined; If at least two cluster sets have a signal-to-noise ratio less than the second threshold, then the cluster sets with a signal-to-noise ratio less than the second threshold are adjusted to obtain the adjusted cluster sets. The adjusted cluster set is input into a pre-trained classification model to obtain the category of the sequence to be identified; wherein, the pre-trained classification model is used to determine the category of the sequence to be identified.

8. The method according to claim 7, characterized in that, The adjustment of the cluster set whose signal-to-noise ratio is less than the second threshold to obtain the adjusted cluster set includes: The cluster set with a signal-to-noise ratio less than the second threshold is merged with the cluster set with a higher signal-to-noise ratio from the two adjacent cluster sets to obtain the adjusted cluster set.

9. The method according to claim 6, characterized in that, The step of inputting the multiple cluster sets into a pre-trained classification model to obtain the category of the sequence to be identified includes: The first network of the pre-trained classification model is invoked to process the multiple cluster sets to obtain a first feature; the first feature includes local features of each cluster set in the multiple cluster sets. The first feature is input into the second network of the pre-trained classification model to obtain the second feature; the second feature includes the feature after the local feature enhancement time sequence. The second feature is input into the third network of the pre-trained classification model to obtain the third feature; the third feature is used to characterize the differences between the features after the augmentation time series. Based on the third feature, the category of the sequence to be identified is determined.

10. The method according to claim 6, wherein the verification process of the pre-trained classification model includes: Input the sequence to be identified from the validation set into the pre-trained classification model, and output the category of the sequence to be identified; The validation set is used to verify the classification performance of the pre-trained classification model.

11. A data preprocessing apparatus, characterized in that, The device includes: A determining unit is used to determine the initial length of the sequence to be identified; the sequence to be identified is used to characterize the structure of a biomolecule; the sequence to be identified is a time-series current trajectory generated by the biomolecule to be identified through a nanopore; The first processing unit is used to perform standardization processing on the sequence to be identified based on the initial length to obtain the target sequence to be identified; The second processing unit is used to perform clustering processing on multiple data in the target sequence to be identified, to obtain multiple cluster sets; the cluster sets are used to identify the category of the sequence to be identified; The first processing unit is configured to, if the initial length is greater than a first threshold, remove the portion of the initial length of the sequence to be identified that exceeds the first threshold to obtain the target sequence to be identified; if the initial length is less than the first threshold, fill the head and / or tail of the sequence to be identified with noise data to obtain the target sequence to be identified with a length equal to the first threshold.

12. An electronic device, characterized in that, The device includes a memory and a processor, the memory storing a computer program that can run on the processor, the processor executing the program to implement the data preprocessing method according to any one of claims 1 to 10.

13. A computer-readable storage medium, characterized in that, It stores a computer program that, when executed by a processor, implements the data preprocessing method according to any one of claims 1 to 10.

Citation Information

Patent Citations

  • Novel coronavirus subgroup identification method based on machine learning

    CN113990390A