Method, system, computer program product and training method of a deep learning architecture for analyzing a mixture of substances

By converting translocation signals into image data through wavelet transformation and using CNNs with skip connections, the method achieves accurate classification of multiple proteins in nanopore experiments, addressing the limitations of conventional methods and enabling efficient, scalable protein identification.

DE102024121850A1Pending Publication Date: 2026-02-05UNIV STUTTGART KORPERSCHAFT DES OFFENTLICHEN RECHTS
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
DE102024121850
Authority / Receiving Office
DE · DE
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-07-31
Publication Date
2026-02-05

AI Technical Summary

Technical Problem

Conventional methods for classifying translocation signals from nanopore experiments are not sufficiently accurate for distinguishing a large number of proteins, particularly when they have similar physical properties, limiting the classification to a maximum of ten different proteins, which is insufficient for real-world applications requiring differentiation of 42 different proteins of interest.

Method used

A method involving wavelet transformation to convert time-dependent translocation signals into image data, followed by analysis using a deep learning architecture, specifically convolutional neural networks (CNNs) with skip connections, to classify components in substance mixtures, particularly proteins, peptides, and DNA in body fluids.

Benefits of technology

Enables the classification of up to 42 different proteins with high accuracy, overcoming the limitations of conventional methods by providing a scalable and flexible analysis capable of distinguishing proteins with similar signal traces, and can be performed efficiently on standard hardware without requiring highly specialized equipment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

In a method for analyzing a substance mixture based on a translocation signal, the substance mixture to be investigated is supplied and moved through at least one nanoscopic channel. An electrical channel current flows through this channel due to a voltage applied across the channel. The electrical channel current is measured as a time-dependent translocation signal while the substance mixture moves through the channel. The measured translocation signal is converted into image data using a wavelet transformation. The image data is then classified using a trained deep learning architecture to extract information about the substance mixture.
Need to check novelty before this filing date? Find Prior Art

Description

The invention relates to a method and a system for analyzing a substance mixture, a computer program product and a method for training a deep learning architecture for a substance mixture analysis.Understanding the chemical and biological composition of the human body can quickly provide medical practitioners with information on the health of the body to be examined. In this regard, translocation experiments through pores provide a method for assaying and classifying microscopic components of blood or other body fluid samples. In these studies, peptides, proteins, DNA or other components of the body fluids can be moved through a microscopic channel through which an electrical current can be measured. As the component passes through the pore, the measured current falls rapidly. By characterizing the way the signal changes when the current is thus impeded, it can be theoretically determined which component has just passed through the pore. In this case, in particular, some proteins can be classified as a component of the body fluid.Proteins are, like machines, in biological cells responsible for their function. The properties of proteins, including their correct function, depend on their structure, i.e. on the amino acids constituting the protein molecule. The development of methods that allow conclusions about the structure of these proteins can improve understanding of biological processes and lead, for example, to earlier detection of diseases, including cancer. A more recent approach to acquiring information about individual proteins is nanopore biological sequencing as a translocation experiment in which a protein is driven through a biological nanopore while a current is being measured through it. The flow of charged ions through the pore forms a circuit and the current through the pore decreases as the protein passes through the pore. Such a temporary drop of current when a molecule passes through is shown, for example, in Fig. 1. How much current drops depends on numerous factors, especially the geometry of the protein, its interaction with the pore, and the geometry of the pore itself.With such nanopore sequencings, data about individual proteins can be collected, but they do not conventionally provide an approach to accurately rank these data so that the protein passing through the pore can be accurately determined. Moreover, the translocation signals measured at the pore often look very similar and do not offer a clear possible interpretation.Several attempts have already been made to classify proteins on the basis of these translocation signals. In the publications of Ensslen et al. (Resolving Isomeric Posttranslational Modifications Using a Biological Nanopore as a Sensor of Molecular Shape. Journal of the American Chemical Society, 144(35), 16060-16068. https: / / doi.org / 10.1021 / jacs.2c06211from 2022) and Jain et al. (Discrimination of L- and Dichirality in oligoarginine peptides using the wild-type and R220s mutant aerolysin pores. Biophysical Journal, 122(3), 439a. https: / / doi.org / 10.1016 / j.bpj.2022.11.2374 of 2023) histograms of the current amplitudes across the nanopore data are used to identify large groups of signals that form clusters.Further evaluation tests aim at evaluating an amplitude of the translocation signals, a percentage blocking, or a dwell time in the pore. As a result, it has always been possible to create feature representations of at least four different proteins, which are then classified using classic statistical methods, for example.While several such conventional methods provide promising results, they still have several common problems. When evaluating information such as the signal amplitude, the dwell time or the percentage of pore blockage, it is assumed that individual proteins have significantly different values in these parameters. FIG. 2 shows amplitude histograms of translocation signals of several different proteins, which in part overlap strongly. As shown in FIG. 2, some of the characteristics of the translocation signals of the different proteins are very similar, which makes the evaluation more difficult. It has therefore been possible hitherto in some experiments to classify only a maximum of ten different proteins, which is significantly too low for real measurements in which 42 different proteins of interest would have to be distinguished at present. For substance mixtures with a large number of components, the evaluation of histograms obtained from the translocation signal has hitherto proved to be not practical.Conventional methods for classifying translocation signals are not sufficiently accurate for a larger number of classification classes or for systems where different analytes are physically similar and therefore generate similar signal traces.The object of the invention is therefore to make possible an improved analysis of substance mixtures on the basis of translocation signals. In this case, it may be an object in particular to be able to classify a larger number of components of the substance mixture, i.e. to scale the methods for a large number of components, such that it is possible to distinguish between analytes having similar physical properties. It may also be an object to be able to flexibly modify translocation experiments for the analysis of substance mixtures to applications in new problems. It may also be an object to enable analysis of substance mixtures in which hardware can be used efficiently, i.e. which can be carried out completely offline, e.g. in an end-to-end sequencing workflow.The object is achieved by the subject matters of the independent claims. The dependent claims relate to preferred embodiments.One aspect relates to a method for analyzing a substance mixture based on a translocation signal. The substance mixture to be examined is provided. The substance mixture is moved through at least one nanoscopic channel through which an electrical channel current flows due to a voltage applied via the channel. The channel electric current is measured as a time dependent translocation signal as the substance mixture is moved through the channel. The measured translocation signal is converted into image data by means of a wavelet transform. Finally, the image data are classified by means of a trained deep learning architecture in order to determine information about the substance mixture.The method thus relates to the analysis of substance mixtures on the basis of translocation signals which are obtained from a translocation experiment of the type mentioned at the beginning. In this case, a voltage is applied across the nanoscopic channel, wherein a first region with substance mixture upstream of the nanoscopic channel is at a different electrical potential than a second region downstream of the nanoscopic channel. As a result, ions as components of the substance mixture are excited to pass through the nanoscopic channel, wherein the channel current flows. The channel current is measured as a time-dependent translocation signal, i.e. for example as a time-dependent current flow. The translocation signal need not correspond directly to the channel current, but may be derived therefrom. Thus, for example, the normalized channel current can be used as the translocation signal. The translocation signal may also be referred to as a pore translocation signal.By means of the method, in particular liquid substance mixtures such as body fluids can be analyzed. Components of liquid substance mixtures can be moved particularly well through nanoscopic channels. The nanoscopic channel can be formed, for example, as a nanopore and / or comprise a nanopore. The nanoscopic channel is characterized in that it has dimensions in the nanometer range, i.e. for example an "inner diameter" in the nanometer range. The term inner diameter in the case of nanopores is not to be understood as being identical to the same term in the case of larger tubes, since nanopores usually do not have a circular geometry, but rather a more complex geometry. Nevertheless, the "inner diameter" of the nanoscopic channel can be understood, for example, as the smallest or mean inner extent orthogonal to the channel direction through the channel.By means of the deep learning architecture, microscopic and / or nanoscopic components of the substance mixture can be determined and / or sequenced, in particular peptides and / or proteins and / or amino acids and / or DNA contained in the substance mixture.The measured time-dependent translocation signal for a passage through the nanoscopic channel can be assigned to individual components of the substance mixture. The components can be, for example, peptides, proteins, DNA or other constituents of the substance mixture. The aim of the analysis can be to assign the time-dependent translocation signal, in particular time segments with a lowered channel current, to components of the substance mixture in such a way that the components can be recognized and / or classified. The method can thus be used to investigate which components are contained in the substance mixture.Since this assignment is not trivial, but rather is very complex in some cases, as shown in FIG. 2, for example, the time-dependent translocation signal is not directly evaluated, but is first converted into image data by means of the wavelet transformation. In this case, a time-limited section of the translocation signal can be converted as an individual signal by means of the wavelet transform. Thus, not the translocation signal per se is evaluated directly, but its wavelet transform.The translocation signal is transmitted into the frequency space by means of the wavelet transform. Both time-dependent and frequency-dependent information of the translocation signal can be converted into the image data. This enables the use of efficient algorithms of the deep learning architecture, which can evaluate the image data more efficiently than the directly measured translocation signal. The image data can be provided and / or evaluated in a common digital image format, in particular as raw data. The image data can be present as two-dimensional image data with an evaluable contrast, e.g. a light / dark contrast and / or a color contrast.The deep learning architecture can be trained to classify components of substance mixtures on the basis of such image data, which are obtained from the translocation signal by means of the wavelet transform. For this purpose, for example, sections of the translocation signal that are limited to be constant over time can be converted as individual signals into image data and evaluated.The deep learning architecture may be configured as a machine learning architecture including artificial neural networks having a plurality of intermediate layers (layers) between an input layer and an output layer. Thus, the deep learning architecture has a complex internal structure.It has been found that deep learning architectures can assign the translocation signals thus converted into image data to individual components of the substance mixture more efficiently and reliably than by evaluating the directly measured translocation signals. Thus, for example, all 42 of the proteins of current interest from body fluids can be classified by means of the method, which clearly goes beyond the methods best hitherto, in which, for example, a maximum of ten different proteins can be classified. The method is not limited to the classification of the currently relevant 42 proteins, but is flexibly scalable to the classification of further proteins and / or components.According to an embodiment, the deep learning architecture comprises a machine learning architecture having trainable and / or trained neural network layers, in particular convolutional layers with skip connections. Such convolutional layers are also called convolutional layers, of the English "convolutional layer". The skip connections pass data from one convolutional layer to another convolutional layer, and some intermediate layers may be skipped. This makes it possible, for example, to subsequently include additional information and / or training data in the deep learning architecture, for example for classifying further components of the substance mixture and / or in the case of another change in the method structure. It has been found that such convolutional deep learning architectures, which are also referred to as convolutional neural network, evaluate the image data particularly efficiently.As deep learning architecture suitable for the method, a convolutional neural network (CNN) and / or a residual neural network (RestNet) and / or a ResExt and / or a vision transformer (ViT) model can be used. It is evident that, in particular, the CNN efficiently evaluates the image data generated by means of wavelet transformation.In a development of this embodiment, the deep learning architecture has from about 15 to about 150 network layers, in particular a maximum of 120 network layers. This number of network layers may already be sufficient for a reliable classification of the components of the substance mixture. The use of convolutional deep learning architectures, in particular with this still manageable number of network layers, enables the analysis of substance mixtures by means of modern standard hardware such as FPGA, which can be executed offline. Consequently, no highly specialized and thus expensive hardware components are required for execution. In addition, the training effort of such algorithms can still be handled. Other architectures, such as transformer style architectures or else LSTM architectures, which comes from English long short-term memory, are significantly more complicated to train and therefore impractical.In one embodiment, the translocation signal is converted into the image data by means of the following wavelet transform:Here, x(t) denotes the time-dependent translocation signal, a selected wavelet function, a and b denote scaling factors corresponding to x and y coordinates of the image data, respectively, and X ψ(a, b) denotes the image data.By means of this wavelet transformation, the image data can be generated from the translocation signal, which contains both the time and the frequency information. This wavelet transform provides a reversible operation by means of which the raw signal can be converted into an image in which e.g. the colour of the pixels corresponds to the amplitude of the signal and the x and y coordinates, i.e. actually the a and b coordinates, provide information on the time and frequency properties. In this way, a representation of the translocation signal can be constructed that is well suited for deep learning without too much information being lost. Once the image data is created, it is forwarded to the deep learning architecture, in particular to a CNN, ResExt, ResNet or vision transformer model.In one embodiment, the nanoscopic channel comprises a nanopore having an inner diameter of about 1 nm to about 100 nm, in particular of about 1 nm to about 20 nm. The nanopore can have an inner diameter of 2 nm at the maximum, for example. This dimension is particularly well suited for classifying e.g. the relevant proteins currently known 42. The diameter is selected such that a distinct reaction of the channel stream is effected upon passage of at least one component of the substance mixture through the channel. Thus, for example, an aerolysin pore having a diameter of about 1.4 nm can be used to classify proteins and / or peptides.In one embodiment, a body fluid, in particular blood and / or urine and / or cerebrospinal fluid, is examined for its constituents as a substance mixture. These substance mixtures usually contain proteins and / or DNA, which can be investigated by means of the method as components of the substance mixture. Thus, they are well suited for analysis by the method. Furthermore, the detailed examination of these body fluids is relevant for a multiplicity of medical applications.In one embodiment, the deep learning architecture is trained to determine at least 42 different proteins from the translocation signal as a constituent of the substance mixture. The deep learning architecture is thus configured and designed to recognize all currently relevant 42 proteins in the substance mixture. The method thus overcomes the limitations of the conventional methods known up to now by significantly expanding the possible applications of the analysis. The deep learning architecture can in this case in particular have a CNN that is trained for classifying the at least 42 proteins.In one embodiment, the time-dependent translocation signal is divided into individual, time-limited event signals, which are each assigned to a passage of a component of the substance mixture through the nanoscopic channel. Here, the wavelet transformation can be carried out on the individual event signals. Thus, the time-dependent translocation signal is broken down into a plurality of individual event signals. These can either each comprise a predetermined, for example constant, time interval or can also be designed with different lengths. These individual signals are easier to handle than the entire, for example continuous, translocation signal and, in addition, are clearly characteristic of the passage of the different components of the substance mixture.In a development, the event signals are determined by analysis of statistical moments of the translocation signal, in particular by comparison of the current translocation signal with a mean translocation signal and / or a median of the translocation signal. In this case, in particular the start and / or the end of the individual signal is determined by comparison with the mean value and / or median value of the channel current. As shown on the translocation signal of a passage shown by way of example in FIG. 1, the channel current during the passage, i.e. approximately during the time period from approximately 300 μs to approximately 1200 μs, is significantly reduced in comparison to the otherwise customary mean value and / or median value of the channel current, i.e. in the example during the time before approximately 300 μs after approximately 1200 μs. As a result, this change in the channel flow during the passage can be reliably associated with an individual event such as a pore passage. The comparison with the mean value / median value of the translocation signal can be carried out in a computer-assisted manner, in particular automatically, which simplifies the evaluation.The deep learning architecture can be trained with image data which are created by means of wavelet transformation from individual signals of the time-dependent translocation signal. Thus, single signals can be used already during the training of the deep learning architecture, which single signals are also determined, for example, as described above by comparison with a mean value and / or median value of the translocation signal. Such training can provide an effective deep learning architecture for the method.The training of the deep learning architecture can be and / or can be carried out on the basis of a loss function and / or cost function which aims at the correct classification of the components. The error allowed in the loss function can be reduced further and further during the course of the training, in order to improve the training. For training, it is possible to use, for example, categoric cross entropy (also referred to as categoric cross entropy) and / or the mean square error (also referred to as MSE). An example of a cross entropy useful for training is shown, for example, in DataAmp- (2024)-Cross-Entropic Loss Function in Machine Learning: Enhancing Model Accuracy from https: / / www.datacamp.com / tutorial / the-cross-entropic-loss-function-in-machine-learning (e.g. archived at https: / / web.archive.org / web / 20240726123456 / https: / / www.datacamp.com / tutorial / the-cross-entropic-loss-function-in-machine-learning).One aspect relates to a system for analyzing a substance mixture based on a translocation signal, comprising:an input device for receiving the substance mixture to be examined;at least one nanoscopic channel through which a channel electrical current flows due to a voltage applied across the channel;a drive mechanism for moving the substance mixture through the nanoscopic channel;a measuring device that measures the electrical channel current as a time-dependent translocation signal while the substance mixture is moved through the channel; andan evaluation module which converts the measured translocation signal into image data by means of a wavelet transformation and classifies this image data by means of a trained deep learning architecture in order to determine information about the substance mixture.The system may be configured and / or used to perform the method described above. Therefore, all the statements regarding the method also relate to the system and vice versa. Individual modules of the system can be controlled and / or regulated by means of software and / or a processor, in particular the drive mechanism, the measuring device and / or the evaluation module. For example, the voltage applied to the channel by the drive mechanism may be controlled to apply a predetermined (e.g., constant) voltage to the channel. The measuring device can be switched on and / or off. In particular, the evaluation module, which has the deep learning architecture, can be controlled and / or regulated by means of software and / or processor.In one embodiment, the nanoscopic channel has a nanopore having an inner diameter of about 1 nm to about 100 nm, in particular of about 1 nm to about 20 nm, in particular of about 1 nm to about 2 nm. As already stated, this dimensioning is particularly suitable for classifying peptides and / or proteins, for example. The dimensioning is selected such that a distinct reaction of the channel stream takes place when at least one component of the substance mixture passes through the channel.In one embodiment, the deep learning architecture includes a convolutional neural network (CNN) with skip connections having from about 15 to about 150 network layers. As likewise already explained, this number of network layers may already be sufficient for reliable classification of the components of the substance mixture. The use of convolutional deep learning architectures with this manageable number of network layers enables the analysis of substance mixtures by means of standard hardware, which can be executed offline. In addition, the use of such deep learning architectures makes it possible to keep the training within the practicable scope compared to, for example, transformer style architectures.One aspect relates to a computer program product comprising computer readable instructions which, when loaded into a computer system and executed there, cause the computer system to carry out the operations of the evaluation module of the system according to the preceding aspect, including the conversion of the measured translocation signal into image data by means of a wavelet transformation and the classification of this image data by means of a trained deep learning architecture in order to determine information about the substance mixture.The computer program product can be designed as a software module, in particular as a pure software module, and / or it can additionally (simultaneously) comprise hardware. As hardware, the computer program product can have an FPGA, i.e. a field programmable gate array.In one embodiment of the computer program product, the deep learning architecture is configured and / or optimized for use on hardware with limited storage capacity and / or offline operation. The analysis can thus be carried out, for example, by means of a favorable analysis device, in particular by means of a handheld device, which can carry out the complete classification of the components in offline operation. This can be realized in particular by the selection of deep learning architectures suitable for this purpose, e.g. by the use of CCNs.One aspect relates to a method for training a deep learning architecture for substance mixture analysis, wherein a training data set is generated which comprises time-dependent translocation signals from known substance mixtures as training signals. The training signals are converted into image data by means of wavelet transformation. The deep learning architecture is trained using the image data for classifying components of substance mixtures and validated using a separate validation dataset.The method can be used for training the deep learning architecture, which is used in the method and / or in the system according to the preceding aspects for analyzing the substance mixture. Therefore, all the explanations concerning the analysis method and the analysis system also relate to the training method and vice versa.The training can initially take place on the basis of image data as training data, which are generated by means of wavelet transformation from translocation signals of previously known components. The training can be carried out with optimization of a target function, for example with optimization of a suitable cost function. Image data of previously known components can likewise be used for validation. The image data of the validation dataset differ from the image data of the training dataset.In the context of this invention, the terms "substantially" and / or "about" may be used to include a deviation of up to 5% from a numerical value following the term, a deviation of up to 5° from a direction following the term, and / or from an angle following the term.Preferred embodiments of the invention are described below by way of example with reference to figures. In this case, identical or similar reference numerals may identify identical or similar features of the embodiments. Individual elements of the described embodiments are not limited to the respective embodiment. Rather, individual elements of the embodiments can be combined with one another and new embodiments can be created thereby. It shows: FIG. 1 is a diagram showing a portion of a translocation signal of a channel current through a nanoscopic channel over time as a raw data signal; FIG. 2 shows an amplitude histogram of a plurality of mutually overlapping protein signals; FIG. 3 shows a flow chart of a method for analyzing components of a substance mixture; FIG. 4 is a diagram showing image data obtained by converting a translocation signal by wavelet transformation; FIG. 5 is a diagram illustrating a distribution of training data and test data used to train a deep learning architecture; FIG. 6 is a diagram of a confusion matrix generated for a deep learning architecture based on Resnet18; FIG. 7 is a diagram showing a distribution of correct positive predictions resulting for the deep learning architecture based on Resnet18; FIG. 8 is a diagram of a confusion matrix generated for a deep learning architecture based on ResExt 101; and FIG. 9 is a diagram showing a distribution of correct positive predictions resulting for the deep learning architecture based on ResExt101.FIG. 1 shows a diagram of a section of a translocation signal of a channel current through a nanoscopic channel, which results during a translocation setup. The nanoscopic channel is arranged in a boundary layer and / or boundary wall, by which a first spatial region with a substance mixture is separated from a second spatial region. A part of the substance mixture can also be arranged in the second spatial region, or alternatively originally only one buffer. The two spatial regions are separated from one another by means of the boundary layer and / or boundary wall and are connected to one another (in particular only) by the nanoscopic channel.The substance mixture can be formed as a liquid. The nanoscopic channel is dimensioned such that it allows the passage of individual components of the substance mixture through the channel.The first spatial region with the substance mixture is at a different electrical potential than the second spatial region behind the channel. An electrical voltage is thus applied across the channel, which ions as components of the substance mixture for passage through the nanoscopic channel move from the first into the second spatial region (and / or vice versa, depending on the applied voltage). These ion movements cause an electrical current flow through the channel, which is also referred to as channel current and is configured as time-dependent.The translocation signal is calculated from the channel current. The translocation signal may correspond directly to the channel current or depend directly on it, in particular proportionally. In the graph shown in FIG. 1, the normalized channel current is used as the translocation signal and plotted over time in μs. The translocation signal may be present and / or used as a raw data signal, fluctuate over time, and / or contain frequency information.FIG. 1 shows that the normalized channel current fluctuates about the value 1 in a time window from 0 μs to approximately 300 μs. This thus corresponds to the "normal" current flow of the translocation structure used by way of example. In a time window of about 300 μs to about 1300 μs, the normalized channel current is significantly reduced to about 60% of the normal current flow. Subsequently thereto, i.e. in a time window from 1300 μs, the normalized channel current again rises steeply by fluctuating by the value 1.The translocation signal shown in FIG. 1 can be interpreted in such a way that a translocation event has taken place in the time window of about 300 μs to about 1300 μs, in which a component of the substance mixture has been moved through the nanoscopic channel and in the process has reduced the current flow.Such a translocation event may be identified, for example, by comparing the current channel current to an average and / or median of the channel current over time. If the current channel current deviates significantly from this mean value and / or median value, if it is reduced in particular in comparison therewith, the time window with the reduced channel current can be defined and / or considered as an event signal. The event signal can also comprise at least one boundary time window in addition to the time window with the reduced channel current, i.e. in the example the time window of about 300 μs to about 1300 μs. Such a time-out window precedes and / or follows the event in time. It can be designed as a time window with a normal current flow and can comprise, for example, a fixed, predetermined time length.In the example shown in FIG. 1, the event signal also has two edge time slots each of about 300 μs in length; in the example, a first edge time slot extends from 0 μs to about 300 μs and a second edge time slot extends from about 1300 μs to about 1600 μs. Instead of a fixed time length of, for example, approximately 300 μs, the edge time window can also have a percentage of the length of the time window with the reduced channel current, for example, from approximately 20% to approximately 50% of its time length. A predetermined control for the length of the edge time slot or slots can improve both training and classification. In addition, the edge time window(s) can also contain information relating to the passage of pores, which can be concomitantly taken into account in the analysis.Thus, the continuous translocation event, for example, can be divided into a plurality of event signals, each event signal of which can be assigned to a passage of a component of the substance mixture through the nanoscopic channel.FIG. 1 shows such an event signal, i.e. a section of a longer translocation signal, which is determined from this translocation signal and can be assigned to the passage of a component of the substance mixture through the nanoscopic channel. The event signal is time dependent and intrinsically contains information about the component whose passage caused the change in the channel current. The event signal per se shown in FIG. 1 is not as readily accessible to deep learning architectures, except, if appropriate, extremely complicated architectures such as, for example, transformer style architectures. Training such an architecture, however, usually requires an extreme amount of training data and is therefore hardly practicable.Therefore, the event signal shown in FIG. 1, which is derived from the translocation signal, is first processed according to the invention in order to make it accessible to manageable deep learning architectures.FIG. 2 shows an amplitude histogram of a plurality of mutually overlapping protein signals. This means that the densities shown in FIG. 2 are each obtained from an event signal which can be assigned to the passage of a component, such as a protein or a peptide of a substance mixture, through the channel. As shown in FIG. 2, the individual signals overlap greatly, so that they are very difficult to distinguish. Conversion of the event signals into such amplitude signals has not been shown to be sufficient to achieve the desired reliability and / or scalability.FIG. 3 shows a flow chart of an exemplary embodiment of a method for analyzing components of a substance mixture. In this case, an event signal is initially assumed, which can be determined from the translocation signal, as shown in FIG. 1.The event signal is converted by means of a wavelet transform into image data, for example into an image whose coordinates correspond to the time and the frequency of the event signal. The image data can be two-dimensionally configured and have a contrast dependent on the event signal, e.g. a color contrast and / or a grayscale contrast.The image data determined in this way are fed to a trained deep learning architecture, which is designated as a vision model in FIG. 3. The deep learning architecture is trained on the basis of image data which are obtained by means of wavelet transformation from event signals of previously known components. The deep learning architecture can thus compare the image data of unknown components with the training data and in the process classify unknown components of the substance mixture.FIG. 4 shows the transformed event signal as image data which are supplied to the deep learning architecture, once again separately. The image data transformed in this way is more accessible to machine learning algorithms than the raw data of the event signal shown in FIG. 1. Thus, neural networks, also abbreviated as NN, are well suited for classifying complex image data. Therefore, the conversion of the graph shown in FIG. 1 to image data enables deep learning architectures to more reliably map the data.The wavelet transform mentioned above can be used as the wavelet transform for generating the image data, i.e.FIG. 4 shows the image data X ψ(a, b). In this case, for example, the color of each pixel of the image data can correspond to the signal amplitude of the translocation signal. The coordinates (a, b) may represent the time and frequency of the translocation signal, just as shown in Figure 4.The wavelet transform allows the translocation signal to be better evaluated on the one hand for the deep learning architectures used, and on the other hand information losses are kept small and / or minimized during the transformation, since the wavelet transform is reversible with low loss.Suitable deep learning architectures are in particular CNN architectures, i.e. architectures with convolutional or convolutional layers.Protein Translocation EmbodimentIn one embodiment, proteins are examined which are moved through a biological nanopore as a channel in a translocation experiment. Data were recorded in a continuous sequence during which hundreds of thousands of gating events occurred. To isolate the event signals, a statistical moment analysis was first performed during which the current signal of the channel current was compared with the mean channel current while the channel was open. Whenever a channel current occurs that has a value that differs from the background by about 3 to 4 standard deviations, this is defined as the beginning of a gating event and is separately established and / or used as an event signal. This produces event signals which look as shown in FIG. 1.Using wavelet-based denoising algorithms can improve the isolation of the event signals by reducing the number of erroneous event triggers.After the event signals have thus been extracted from the translocation signal, it is possible, for example, to train a ResNet18 network as a deep learning architecture. In one exemplary embodiment, a ResNet18 network was trained in this way on the basis of 250,000 images, which correspond to image data that were determined from 250,000 event signals by means of wavelet transformation. The event signals were generated from 42 different proteins. These are all proteins currently of interest in sequencing analyses.FIG. 5 shows a diagram of the resulting distribution of the training data and test data. As a result of the exemplary embodiment, an overall accuracy of 74.5% is obtained, with which the 42 proteins were correctly recognized in test data, i.e. in which "true positives" were determined.FIGS. 6 and 7 show an associated confusion matrix and the associated case numbers of the correctly recognized proteins. The overall accuracy is calculated by the number of correctly predicted and / or classified proteins compared to the total number of predictions and / or classifications made.In the confusion matrix, the rows correspond to the model predictions and the columns correspond to the target values. Therefore, the diagonal of the confusion matrix shows the frequency of the correctly predicted target value, and thus the correctly classified proteins, which represent components of the substance mixture. The higher the percentage of correct predictions, the stronger the color strength of the cell. Ideally, i.e. if only correct predictions have been made, confusion matrices correspond to unit matrices having only 1s in the diagonal.In a translocation experiment carried out as described above, the results shown in FIGS. 6 and 7, in particular the confusion matrix shown in FIG. 6, are broken down in column 2 of the following table 1: Table 1: Comparison of the performance of trained deep learning architectures based on ResNet 18 and on ResExt101. Table 1: Comparison of the performance of trained deep learning architectures based on ResNet 18 and on ResExt101.Components4242Accuracy74,4678,95Rectal74,4678,95F1 Score74,4678,95In addition to the results of the exemplary embodiment based on ResNet18, the results of a further exemplary embodiment based on ResNetx101 are also shown in Table 1. In addition, the rows also show the Recall and the F1 score under the precision.FIG. 7 shows the histogram of the correct predictions (true positives) associated with the test run.As a result, the deep learning architecture based on ResNet18 predicts results with an accuracy which is already of use in laboratory examinations, namely with the accuracy of 74.46%, for a relatively high number of different components, namely for the 42 different proteins.In a further test run with a somewhat more comprehensive and / or more complex model architecture, which is based on the ResExt101 architecture, an even higher accuracy of almost 79% was achieved.Analogously to FIGS. 6 and 7, FIGS. 8 and 9 show the confusion matrix belonging to this second test run and the case numbers of the correctly recognized proteins. The results of this second experiment are also shown in Table 1 in the third column.With additional training, the accuracy of the correct predictions can be further improved.However, since up to now no conventional method is able to predict a model for the high number (42) of proteins, in particular not in the accuracy achieved here, the method according to the invention makes possible a significant improvement of the previously known classification methods.The method according to the invention can also make it possible that the type of training can benefit greatly from pretraining and transfer learning. Once a deep learning architecture is trained, whether for only a small number of classes, additional data may then be augmented and the model may be trained supplementally based on the additional data.The method thus enables simplified addition of additional components to the deep learning architecture without having to retrain the architecture from reason to reason. Thus, the training time is reduced, and the performance can be even improved by the additional data based on the effect of the transfer learning. This scalability can be enabled in particular by using an architecture that has skip connections between the layers.In one embodiment, a basic model is first trained for specific components and / or groups of components. This basic model can then be supplemented by additional training data for the classification of additional components and / or groups. Thus, the method can be scaled to potentially thousands of components such as proteins, analogous to other image classification methods in other application fields.In one embodiment, the parameters of the wavelet transform used are automatically adjusted and / or optimized by the deep learning architecture used for an improved classification of the components. This may already have taken place and / or have taken place during the training. Thus, by using the neural network of the deep learning architecture for classification, an improved model can be generated, which classifies the components more reliably. This can be made possible in the deep learning architecture by including the parameters of the wavelet transformation in the optimization steps carried out during the training.The use of the deep learning architectures according to the invention, in particular CNNs, is well suited for use on current hardware. The invention thus makes it possible to carry out the classification offline, for example, and / or by means of a local computer system, for example, by means of a handheld device.More complex methods such as transform style architectures could theoretically be able to perform an immediate classification of signals from the time series itself, i.e. without wavelet transformation. However, they typically require much larger and more complex neural network architectures and hence much more training data. The approach according to the invention, in particular based on a ResNet-like architecture, enables the method to be carried out on significantly smaller networks and / or on modern FPGA hardware. Here, FPGA is an acronym for Field Programmable Gate Array, i.e. an integrated circuit (IC) of digital technology, into which a logic circuit can be loaded.Therefore, in one embodiment, the system includes FPGA hardware and is configured to perform the analysis and / or classification entirely offline. Thus, even in offline operation, a powerful implementation of the classification of translocation signals is made possible.References included in the specificationThis list of documents cited by the applicant has been produced in an automated manner and is only included for the better information of the reader. The list is not part of the German patent application or utility model application. The DPMA does not take any adhesion for any faults or omissions.Cited Non-Patent LiteratureEnsslen et al. (Resolving Isomeric Posttranslational Modifications Using a Biological Nanopore as a Sensor of Molecular Shape. Journal of the American Chemical Society, 144(35), 16060-16068. https: / / doi.org / 10.1021 / jacs.2c06211 from 2022

[0005] Jain et al. (Discrimination of L- and Dchirality in oligoarginine peptides using the wild-type and R220s mutant aerolysin pores. Biophysical Journal, 122(3), 439a. https: / / doi.org / 10.1016 / j.bpj.2022.11.2374 from 2023

[0005] DataAmp - (2024) - Cross-Entropy Loss Function in Machine Learning: Enhancing Model Accuracy from https: / / www.datacamp.com / tutorial / the-cross-entropy-loss-function in-machine learning (e.g. archived at https: / / web.archive.org / web / 20240726123456 / https: / / www.datacamp.com / tutorial / the -cross-entropy-loss-function in-machine learning

[0033]

Claims

Method for analyzing a substance mixture based on a translocation signal, comprising: - providing the substance mixture to be examined; - moving the substance mixture through at least one nanoscopic channel through which an electrical channel current flows on the basis of a voltage applied via the channel; - measuring the electrical channel current as a time-dependent translocation signal while the substance mixture is moved through the channel; - converting the measured translocation signal into image data by means of a wavelet transformation; and - classifying the image data by means of a trained deep learning architecture in order to determine information about the substance mixture.The method of claim 1, wherein the deep learning architecture comprises a machine learning architecture having trainable and / or trained neural network layers, in particular convolutional layers with skip connections.The method of claim 2, wherein the deep learning architecture has from about 15 to about 150 network layers.The method of any preceding claim, wherein the translocation signal is converted into the image data by the following wavelet transform: X ψ ( a, b ) = 1 | a | 1 2 nonyl - ∞ d t x ( t ) ψ ( t - b a ); wherein: x(t) denotes the time dependent translocation signal, ψ ( t - b a ) denotes a selected wavelet function, a and b denote scale factors respectively corresponding to x and y coordinates of the image data, and X ψ(a, b) denotes the image data.The method of any preceding claim, wherein the nanoscopic channel comprises a nanopore having an inner diameter of from about 1 nm to about 100 nm, in particular from about 1 nm to about 20 nm, in particular from about 1 nm to about 2 nm.Method according to one of the preceding claims, wherein a body fluid, in particular blood and / or urine and / or cerebrospinal fluid, is investigated for its constituents as a substance mixture.Method according to one of the preceding claims, wherein the deep learning architecture is trained to determine at least 42 different proteins as a constituent of the substance mixture from the translocation signal.Method according to one of the preceding claims, further comprising a subdivision of the time-dependent translocation signal into individual, time-limited event signals which are each assigned to a passage of a component of the substance mixture through the nanoscopic channel.Method according to claim 8, wherein the event signals are and / or are determined by analysis of statistical moments of the translocation signal, in particular by comparison of the current translocation signal with a mean translocation signal and / or a median of the translocation signal.A system for analyzing a substance mixture based on a translocation signal, comprising: - an input device for recording the substance mixture to be examined; - at least one nanoscopic channel through which an electrical channel current flows on the basis of a voltage applied via the channel; - a drive mechanism for moving the substance mixture through the nanoscopic channel; - a measuring device which measures the electrical channel current as a time-dependent translocation signal while the substance mixture is moved through the channel; and - an evaluation module which converts the measured translocation signal into image data by means of a wavelet transformation and classifies this image data by means of a trained deep learning architecture in order to determine information about the substance mixture.The system of claim 10, wherein the nanoscopic channel comprises a nanopore having an inner diameter of from about 1 nm to about 100 nm, in particular from about 1 nm to about 20 nm, in particular from about 1 nm to about 2 nm.The system of claim 10 or 11, wherein the deep learning architecture comprises a convolutional neural network (CNN) with skip connections having from about 15 to about 150 network layers.A computer program product comprising computer readable instructions which, when loaded into and executed by a computer system, cause the computer system to perform the operations of the evaluation module of the system according to any one of claims 10 to 12, including the conversion of the measured translocation signal into image data by means of a wavelet transform and the classification of this image data by means of a trained deep learning architecture to determine information about the substance mixture.The computer program product of claim 13, wherein the deep learning architecture is optimized for use on hardware with limited memory capacity and / or offline operation.Method for training a deep learning architecture for substance mixture analysis, comprising: - generating a training dataset which comprises time-dependent translocation signals from known substance mixtures as training signals; - converting the training signals into image data by means of wavelet transformation; - training the deep learning architecture using the image data for classifying components of substance mixtures; - validating the trained deep learning architecture using a separate validation dataset.

Citation Information

Patent Citations

  • Systems and methods of phenotype classification using shotgun analysis of nanopore signals

    WO2023215406A1