Deep learning classification method for screening high-confidence single-molecule fluorescence signals
By constructing a vision transformer network model based on spatial perception units, the problem of high confidence signal screening in single-molecule surface-induced fluorescence attenuation experiments is solved, and efficient and accurate single-molecule fluorescence signal recognition is achieved.
Patent Information
- Application Number
- CN202510131078.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-06
- Publication Date
- 2025-06-13
AI Technical Summary
The prior art is difficult to effectively screen and identify single-molecule fluorescence signals with high confidence in single-molecule surface-induced fluorescence attenuation (smSIFA) experiments, especially in the presence of multiple state changes and single-channel data.
Using a vision transformer network model based on space perception units, multi-layer spatial perception unit and vision transformer unit are constructed, combining multi-head self-attention mechanism and multi-layer perceptron, multi-scale features of single-molecular fluorescent signals are extracted and classified.
The efficient identification of the fluorescence attenuation signal induced by single molecule surface was achieved, which significantly improved the accurate identification of high confidence trajectories. The AUC on the test set reached 94.8%, and other performance indicators reached 89%.
Smart Images

Figure CN120148640A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of single-molecule fluorescence signal analysis, and particularly to a deep learning classification method for screening high-confidence single-molecule fluorescence signals. Background Art
[0002] The single-molecule surface-induced fluorescence attenuation (smSIFA) experiment characterizes the distance between a fluorescent dye and the surface of graphene oxide by measuring the fluorescence intensity. Compared with mainstream single-molecule techniques such as fluorescence resonance energy transfer and magnetic tweezers, the single-molecule surface-induced fluorescence attenuation technique can directly report the position of a membrane protein in a phospholipid membrane and more intuitively reflect the interaction between the membrane protein and the phospholipid membrane.
[0003] Single-molecule imaging usually generates a large amount of data. The manual screening and evaluation process may be time-consuming but is crucial. This is because at the single-molecule level, some undesirable phenomena may occur, such as sample aggregation, fluorescence interference, incomplete or incorrect labeling, complex photophysical behavior, and high noise. This data analysis method requires a large amount of time from experienced users, and this process is vulnerable to subjective human factors, thus affecting the accuracy and consistency of the analysis results.
[0004] The premise of calculating the dynamic process of membrane proteins and phospholipid membranes is to design a classification algorithm. This algorithm can select high-signal-to-noise fluorescence intensity change curves from a large number of low-signal-to-noise traces affected by the environment. The booming development of deep learning technology provides a new solution to this problem. The advantage of deep neural networks is that they can learn any complex function to best identify specific features in the input data and model complex non-linear relationships.
[0005] Recently, deep learning models have made remarkable progress in the field of single-molecule fluorescence. Deep learning models reduce human errors and biases while accelerating the analysis process. The leading models in this technical field include DeepFRET, AutoSIM, and Deep-LASI. Among them, DeepFRET can quickly classify single-molecule fluorescence resonance energy transfer traces, and this model includes the entire process from image preprocessing, trajectory extraction, trajectory selection to data analysis. AutoSIM innovatively converts one-dimensional fluorescence intensity data into two-dimensional images for quickly and automatically selecting single-molecule fluorescence resonance energy transfer trajectories, with a consistency of about 90% with manual selection, while significantly shortening the processing time. Deep-LASI classifies three-color single-molecule traces, can also determine the fluorescence resonance energy transfer correction factor, and classify state transitions in dynamic trajectories.
[0006] These deep learning models are trained using simulated fluorescence traces and do not guarantee that the models can exhibit good robustness in the real data generated by experiments. In addition, these models are trained using single-molecule fluorescence resonance energy transfer datasets and are not suitable for single-molecule surface-induced fluorescence decay experimental datasets with only one signaling pathway, such as the programmed necrosis executor protein MLKL and the human antimicrobial peptide LL-37.
[0007] For example, Chinese Patent CN118983007A discloses a method for identifying single-molecule fluorescence events based on local features. This method intercepts local fragments of a single-molecule fluorescence trace sequence through a sliding window, extracts local features using a long-term capture unit, and outputs the recognition result of single-molecule fluorescence event patterns based on a classification unit. This technology can identify dual-color and single-color fluorescence events, including fluorescence resonance energy transfer, the appearance and disappearance of donors and acceptors, etc., and is applicable to the classification of single-molecule fluorescence traces under balanced and unbalanced conditions.
[0008] This patented technology mainly focuses on the identification of dual-color fluorescence trajectories and stable single-color fluorescence trajectories. Since the dataset does not consider the two phenomena of fluorescence quenching and protein-induced fluorescence enhancement, it has limitations in processing single-color dynamic fluorescence trajectories with weak anti-interference ability and low contrast.
[0009] The above-mentioned models are mainly applied to the processing of single-molecule fluorescence resonance energy transfer (smFRET) data. However, such models are not suitable for single-molecule surface-induced fluorescence decay experiment (smSIFA) data because single-molecule surface-induced fluorescence decay data has only one signal channel and there are various state changes. So far, there has been no research on the deep learning automatic analysis of single-molecule surface-induced fluorescence decay data in the relevant field. Summary of the Invention
[0010] Aiming at the deficiencies of the existing technology, the purpose of the present invention is to provide a deep learning classification method for screening high-confidence single-molecule fluorescence signals.
[0011] To achieve the above purpose, the present invention provides the following technical solutions: A deep learning classification method for screening high-confidence single-molecule fluorescence signals, which includes the following steps: 1) Construct a vision transformer network model based on a spatial perception unit, use the fluorescence intensity trajectory of each single-molecule surface-induced fluorescence decay signal as input, and output the classification result or / and softmax prediction score of the trajectory after training, which includes: Several spatial perception units for embedding multi-scale context and local information into tokens; Several vision transformer units for further modeling local features and long-range dependencies in tokens; The output layer maps the learned features to the class probability space and outputs the final classification result, which is used to determine whether the input fluorescence intensity trajectory belongs to a high-confidence trajectory; 2) Capture the surface-induced fluorescence decay signal of each single molecule within a set time and obtain the fluorescence intensity trajectory of each single molecule surface-induced fluorescence decay signal; 3) Preprocess each fluorescence intensity trajectory in step 2) to obtain the intensity data of the fluorescence signal; 4) Input the intensity data of the fluorescence signal obtained in step 3) into the vision transformer network model based on the spatial perception unit constructed in step 1), and output the classification result or / and the softmax prediction score.
[0012] The spatial perception unit includes: The spatial pyramid reduction module is used to downsample the input time series and embed it into tokens with rich multi-scale context by using multiple convolutions with different dilation rates; The multi-head self-attention module is used to model the relationship between single-molecule fluorescence time series segments globally and capture long-term dependence features; The multi-layer perceptron is used to perform non-linear transformation on the features processed by each layer of attention.
[0013] A sequence-to-trajectory module is provided at the output end of each spatial perception unit.
[0014] A trajectory-to-sequence module is provided between the spatial perception unit and the vision transformer unit.
[0015] The vision transformer unit includes: The multi-head self-attention module is used to model the relationship between single-molecule fluorescence time series segments globally and capture long-term dependence features; The multi-layer perceptron is used to perform non-linear transformation on the features processed by each layer of attention.
[0016] The vision transformer network model based on the spatial perception unit includes three layers of spatial perception units and 7 layers of vision transformers, and the last layer of vision transformer is connected to the fully connected layer.
[0017] The three layers of spatial perception units are divided into: The first layer of spatial perception unit: at the initial stage, the input channel is 1, the convolution kernel is 7, and multi-scale features are extracted through a large dilation rate. The dilation rate dilations = [1, 2, 3, 4], and the number of output channels is 64; The second layer of spatial perception unit: at the intermediate stage, the number of output channels is 64, the convolution kernel is 5, and the dilation rate is reduced to dilations = [1, 2, 3]; The third - layer spatial perception unit: In the deep stage, the number of output channels is further increased to 320, the convolutional kernel is 3, and the dilation rate is the smallest, with dilations = [1, 2].
[0018] The vision transformer network model based on the spatial perception unit consists of four stages. The input and output feature dimensions of each stage are different, specifically as follows: Stage 1: The input feature dimension is 1. The convolutional kernel of the dilated convolution is 7, and the stride is 4. Through a relatively large dilation rate dilations = [1, 2, 3, 4], multi - scale features are extracted, and finally, features with a higher embedding dimension are output; Stage 2: The input feature dimension is 64. The convolutional kernel of the dilated convolution is 5, and the stride is 2. The dilation rate is reduced to dilations = [1, 2, 3], focusing on fine - grained feature extraction; Stage 3: The input feature dimension is 64. The convolutional kernel of the dilated convolution is 3, and the stride is 2. The dilation rate is the smallest, dilations = [1, 2], focusing on fusing local and global features; Stage 4: The input feature dimension is 320, the number of attention heads is 6, and the vision transformer module is stacked seven times.
[0019] The pre - processing in step 3) includes: 1. Normalize the data of the single - molecule surface - induced fluorescence decay signal. Use the min - max normalization method to normalize the time - series data into a set interval; 2. Through interpolation, expand or compress the original time - series to a unified length; 3. Perform Gaussian smoothing processing.
[0020] A computer - readable medium stores computer - executable instructions thereon. When the executable instructions are executed by a processor, the above - mentioned deep - learning classification method for screening high - confidence single - molecule fluorescence signals is implemented.
[0021] The beneficial effects of the present invention: Through the classification model established using a neural network, high - confidence single - molecule surface - induced fluorescence decay signals can be effectively identified. By using the spatial perception unit and the vision transformer network, the model extracts the features of the single - molecule surface - induced fluorescence decay signal and effectively distinguishes high - confidence and low - confidence trajectories. The evaluation of the benchmark dataset shows that the model reaches an AUC of 94.8% on the test set and 89% on other performance metrics. These results prove the potential of the model to replace manual screening and provide an efficient and objective method for identifying high - confidence single - molecule surface - induced fluorescence decay signals. Brief Description of the Drawings
[0022] Figure 1It is the overall design flow chart of the present invention.
[0023] Figure 2 It is a schematic diagram of the single-molecule surface-induced fluorescence decay trajectory of the programmed necrosis execution protein MLKL of the present invention, where (a-d) represent high-confidence trajectories, (e-h) represent fuzzy trajectories, and (i-l) represent low-confidence trajectories.
[0024] Figure 3 It is a schematic diagram of the vision transformer network model based on the spatial perception unit of the present invention.
[0025] Figure 4 It is a schematic diagram of the average accuracy rate of the present invention.
[0026] Figure 5 It is a schematic diagram of the AUC of the present invention.
[0027] Figure 6 It is a fluorescence curve graph of the scores of the vision transformer network model based on the spatial perception unit on the MLKL_2 test set and the corresponding fluorescence, where (a-c) represent three high-confidence fluorescence traces, (d-f) represent three fuzzy trajectories, and (g-i) represent three low-confidence fluorescence trajectories.
[0028] Figure 7 It is a schematic table of the results of various models on the programmed necrosis execution protein MLKL data set. Detailed implementation manners
[0029] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0030] It should be noted that all directional indications (such as up, down, left, right, front, back...) in the embodiments of the present invention are only used to explain the relative position relationship and movement conditions between components in a specific posture (as shown in the accompanying drawings). If the specific posture changes, the directional indications will also change accordingly.
[0031] The present invention provides a deep learning classification method for screening high-confidence single-molecule fluorescence signals, which includes the following steps: 1) Construct a vision transformer network model based on the spatial perception unit, use the fluorescence intensity trajectory of each single-molecule surface-induced fluorescence decay signal as the input, and output the classification result or / and softmax prediction score of the trajectory after training, which includes: A number of spatial perception units for embedding multi-scale context and local information into tokens; A number of vision transformer units for further modeling local features and long-range dependencies in the tokens; An output layer that maps the learned features to the class probability space and outputs the final classification result for determining whether the input fluorescence intensity trajectory belongs to a high-confidence trajectory; 2) Capture the surface-induced fluorescence decay signal of each single molecule within a set time and obtain the fluorescence intensity trajectory of each single molecule surface-induced fluorescence decay signal; 3) Preprocess each fluorescence intensity trajectory in step 2) to obtain the intensity data of the fluorescence signal; 4) Input the intensity data of the fluorescence signal obtained in step 3) into the vision transformer network model based on spatial perception units constructed in step 1) and output the classification result or / and softmax prediction scores.
[0032] The classification model established by using a neural network can effectively identify high-confidence single molecule surface-induced fluorescence decay signals. By using spatial perception units and vision transformer networks, the model extracts the features of single molecule surface-induced fluorescence decay signals and effectively distinguishes high-confidence and low-confidence trajectories. The evaluation on the benchmark dataset shows that the model reaches 94.8% AUC on the test set and 89% on other performance metrics. These results demonstrate the potential of the model to replace manual screening and provide an efficient and objective method for identifying high-confidence single molecule surface-induced fluorescence decay signals.
[0033] The present invention performs efficient feature extraction and classification on single molecule fluorescence experimental data. By utilizing spatial perception units and multi-head self-attention mechanisms, the method achieves the capture of multi-scale features, adaptively learns important features, thereby improving the performance and robustness of the model. Specifically: Model architecture design: A1. Spatial perception module: The main objective of this module is to embed multi-scale context and local information into tokens, rather than directly segmenting and flattening the trajectory through a linear 1D block embedding layer. This design introduces the intrinsic scale invariance and local inductive bias of convolution. This process achieves downsampling, reduces the computational amount, and obtains a feature map with spatial and channel information.
[0034] A2. Attention mechanism module: The core layer of the present invention consists of multiple basic units, and each unit integrates a feature interaction component and an information integration unit based on the multi-head attention mechanism internally. These components achieve the effective capture of local information and the feature communication between different regions by defining feature interactions within a specific region.
[0035] B. Model characteristics: B1. Spatial perception unit: It effectively combines time-series multi-scale context information and local information, making the model more precise when capturing features. By introducing the intrinsic inductive bias of convolution, it improves the performance of time-series classification. The multi-level dimensionality reduction design significantly reduces the computational overhead and enhances the efficiency.
[0036] B2. Multi-head attention mechanism: The multi-head attention mechanism contained in each basic unit allows the model to simultaneously focus on different parts of the time series, enabling it to process information of multiple scales in parallel. This mechanism not only enhances the understanding of local features but also captures different types of dependencies through multiple "heads", increasing the model's ability to learn complex patterns.
[0037] B3. Hierarchical structure: Through the hierarchical design, the model can capture multi-scale features, adaptively learn important features, and improve the performance and robustness of the model.
[0038] C. Model output: C1. Global feature fusion: Finally, the abstract features of each layer are fused into the global average pooling layer to extract global information.
[0039] C2. Classification result: The classification result is output through the fully connected layer, realizing the accurate classification of single-molecule surface-induced fluorescence decay signals.
[0040] It can effectively capture local and global features in single-molecule fluorescence signals, realize the effective utilization of time-series data, and optimize the training process. It provides an effective solution for manual screening, helping to eliminate the interference of low-confidence trajectories and fluorescence signals with photophysical artifacts.
[0041] Training of the Vision Transformer network model based on the spatial perception unit The present invention also provides a training process for the Vision Transformer neural network model based on the spatial perception unit, which is used to improve the accuracy of identifying high-confidence single-molecule fluorescence signals. Through the optimized dataset division strategy and model training process, this method makes full use of the overall spatial features of single-molecule surface-induced fluorescence decay signals, improving the generalization ability and prediction performance of the model.
[0042] A. Dataset division strategy: The dataset is divided into a training set, a validation set, and an external independent test set. Among them, the training set and the validation set adopt the method of five-fold cross-validation to further improve the generalization ability of the model and avoid overfitting.
[0043] B. Model training strategy: B1. Training data selection: During the model training process, the intensity information of single-molecule fluorescence signals after normalization, linear interpolation, and Gaussian smoothing is selected as the input of the training data. This can enable the model to focus on learning the key features of single-molecule fluorescence intensity signals, enhancing the recognition ability of the signal-to-noise ratio, the bleaching time, and the number of blinking events.
[0044] B2. Model testing phase: In the model testing phase, when the fluorescence intensity data of single-molecule fluorescence signals is input, the model can quickly and accurately identify high-confidence trajectories by comprehensively analyzing the overall trend characteristics of the trajectories.
[0045] C. Model optimization and evaluation: C1. Optimization algorithm: Appropriate optimization algorithms and loss functions are used for model training to improve the convergence speed and accuracy of the model.
[0046] C2. Performance evaluation: Through performance evaluation on the validation set and an external independent test set, indicators such as the accuracy of the model are calculated to ensure that the model has good generalization ability and stability.
[0047] Through the above method, the present invention effectively utilizes the characteristics of single-molecule fluorescence data, optimizes the model training process, improves the recognition accuracy of high-confidence single-molecule fluorescence signals, reduces the dependence on manual screening, and improves the overall effectiveness of single-molecule surface-induced fluorescence decay experiments in downstream applications.
[0048] Embodiment As Figure 1 shown, the specific design idea: Step 1: Construct a dataset and perform data preprocessing In this application, the single molecules that can be analyzed in this application can be exemplified as the programmed necrosis execution protein MLKL, the human antimicrobial peptide LL-37, etc. Using two different protein samples: MLKL and LL-37, 3 single-molecule monochromatic SIFA signal datasets were collected. Approximately 1500 traces were manually classified by two expert users in this field.
[0049] A. Dataset construction: The training set of this study comes from the fluorescence intensity trajectories generated by single-molecule surface-induced fluorescence decay experiments using proteins labeled with MLKL fluorophores, including 533 high-confidence trajectories and 557 low-confidence trajectories. We performed five-fold cross-validation on the entire dataset, which is a widely used and rigorous method for evaluating model performance, to ensure consistency and reliability. All time series data was randomly divided into a training set (80%) and a validation set (20%).
[0050] In addition, we used two external datasets to test the performance of the model: a single-molecule surface-induced fluorescence decay signal was obtained from the fluorophore of the MLKL label (n = 337, with 190 high-confidence traces and 147 low-confidence traces), and another single-molecule surface-induced fluorescence decay signal was obtained from the fluorophore of the LL-37 label (n = 122, with 32 high-confidence traces and 90 low-confidence traces). Two external datasets were obtained by experienced biologists using the same method as the internal dataset. Figure 2 Figure 2 shows the fluorescence intensity traces generated from the single-molecule surface-induced fluorescence decay experiment with the protein labeled with the MLKL fluorophore. (a-d) represent high-confidence traces, (e-h) represent ambiguous traces, and (i-l) represent low-confidence traces. Traces with a high signal-to-noise ratio, a long and stable bleaching duration, and fewer than three blinking events were labeled as high-confidence. Traces with a low signal-to-noise ratio, a short or unstable bleaching duration, and fluorescence interference nearby were labeled as low-confidence.
[0051] B. Data preprocessing: B1. Normalization: The single-molecule surface-induced fluorescence decay signal data was normalized using the min-max normalization method to normalize the time series data to the [0, 1] interval. This method can effectively eliminate the differences in the numerical range while retaining the relative relationships of the data, laying a foundation for subsequent processing.
[0052] B2. Standardization to a unified length: Through interpolation, the original time series was extended or compressed to a unified length of 1024. Interpolation does not change the trend or shape of the original data, but only adjusts the number of points to facilitate subsequent model training and validation. Linear interpolation is a computationally efficient method suitable for most scenarios of smooth time series.
[0053] B3. Gaussian smoothing: The smoothing intensity is determined by the standard deviation of the Gaussian kernel: a small standard deviation retains more details, and a large standard deviation improves the smoothing effect but may lose local information. Compared with simple moving average, this method is suitable for dealing with noise in time series while retaining the main trend.
[0054] B4. Random seed setting: To ensure the reproducibility of the experimental results, a random seed was set to ensure the consistency of data partitioning and model initialization.
[0055] Step 2: Construction of a vision transformer model based on a spatial perception unit The vision transformer network based on the spatial perception unit introduces the inductive bias in the convolutional neural network into the vision transformer. As Figure 3As shown in the figure, the vision transformer network based on the spatial awareness unit consists of two types of units: the spatial awareness unit and the vision transformer unit. Among them, the spatial awareness unit is responsible for embedding multi-scale context and local information into tokens, while the vision transformer unit is used to further model local features and long-range dependencies in the tokens. Finally, the prediction probability is calculated through the classification token output by the last vision transformer module.
[0056] Model architecture design: A1. Spatial Aware Cell: The first three layers of the model backbone network consist of three spatial pyramid reduction modules. By using multiple convolutions with different dilation rates, the input time series is downsampled and embedded into tokens with rich multi-scale context. In this way, it obtains an inherent inductive bias and enables the model to learn robust feature representations at various scales of single-molecule fluorescence time series.
[0057] The first pyramid reduction module (feature pyramid module) consists of 4 layers of dilated convolutions, the second pyramid reduction module (feature pyramid module) consists of 3 layers of dilated convolutions, and the third pyramid reduction module (feature pyramid module) consists of 2 layers of dilated convolutions.
[0058] The model has three spatial awareness units, and each spatial awareness unit has a feature pyramid module.
[0059] The feature pyramid module in the first spatial awareness unit consists of 4 dilated convolutions. The convolution kernel size of the dilated convolution is 7*7, the stride is 4, and the dilation rates of the 4 dilated convolutions are [1, 2, 3, 4] respectively; The feature pyramid module in the second spatial awareness unit consists of 3 dilated convolutions. The convolution kernel size of the dilated convolution is 5*5, the stride is 2, and the dilation rates here are [1, 2, 3]; The feature pyramid module in the third spatial awareness unit consists of 2 dilated convolutions. The convolution kernel size of the dilated convolution is 3*3, the stride is 2, and the dilation rates here are [1, 2]; Spatial Awareness Unit 1: In the initial stage, the input channel is 1, the convolution kernel is 7, through a relatively large dilation rate (dilations = [1, 2, 3, 4]), the output channel number is 64, extracting multi-scale features, and finally outputting features with a relatively high embedding dimension. Spatial Awareness Unit 2: In the middle stage, the output channel number is 64, the convolution kernel is 5, and the dilation rate is reduced (dilations = [1, 2, 3]), focusing on fine-grained feature extraction. Spatial Awareness Unit 3: In the deep stage, the output channel is further increased to 320, the convolution kernel is 3, and the dilation rate is the smallest (dilations = [1, 2]), focusing on fusing local and global features.
[0060] A2. Visual Transformer Module: Layers 4 to 10 of the model backbone network consist of 7 Visual Transformer modules, which are used to deeply capture global feature relationships and context information. The core structure of these Visual Transformer modules consists of a multi-head self-attention mechanism and a multi-layer perceptron, combined with residual connections and layer normalization to further enhance the stability of training and the expressive power of the model. Multi-Head Self-Attention (MSA): Models the relationships between single-molecule fluorescence time series segments globally and captures long-term dependence features. Multi-Layer Perceptron (MLP): Performs non-linear transformations on the features processed by each layer of attention to further enrich the feature representation ability. Through such a hierarchical design, the model effectively integrates the capabilities of local feature extraction and global context modeling, providing strong support for the classification of single-molecule fluorescence.
[0061] B. Model Structure Design: B1. Hierarchical Design: The model consists of four stages, with different input and output feature dimensions for each stage, specifically: Stage 1: The input feature dimension is 1, the convolutional kernel of the dilated convolution is 7, and the stride is 4. By using a relatively large dilation rate dilations=[1, 2, 3, 4], multi-scale features are extracted, and finally, features with a higher embedding dimension are output.
[0062] Stage 2: The input feature dimension is 64, the convolutional kernel of the dilated convolution is 5, and the stride is 2. The dilation rate is reduced dilations=[1, 2, 3], focusing on fine-grained feature extraction.
[0063] Stage 3: The input feature dimension is 64, the convolutional kernel of the dilated convolution is 3, and the stride is 2. The dilation rate is minimized dilations=[1, 2], focusing on fusing local and global features.
[0064] Stage 4: The input feature dimension is 320, and the number of attention heads is 6. The Visual Transformer module is stacked seven times.
[0065] C. Model Output Layer: C1. Fully Connected Layer: The output of the last Visual Transformer layer is input into the fully connected layer, and the model maps the learned features to the class probability space to output the final classification result, which is used to determine whether the single-molecule surface-induced fluorescence decay signal belongs to a high-confidence trajectory.
[0066] D. Model Parameter Settings: D1. Embedding Dimension: The initial embedding dimension is set to 64, and as the stage progresses, the feature channel dimension gradually increases to 320.
[0067] D2. Number of Stacked Spatial Perception Units: The number of stacked units is set to 3, which not only ensures the capture of multi-scale context information but also controls the computational complexity. The spatial perception unit embeds multi-scale context and local information into tokens, enabling the classification network to have an inherent inductive bias to capture more information. D3. Activation Function: The Gaussian Error Linear Unit (GELU) activation function is used in the multi-layer perceptron to improve the non-linear expression ability.
[0068] D4. Normalization Layer: Layer Normalization (LayerNorm) is used to normalize the features to stabilize the training process.
[0069] Step 3: Training of the Vision Transformer Model Based on the Spatial Perception Unit A. Dataset Partitioning: A1. Initial Partitioning: The collected dataset of 1090 MLKL fluorophore-labeled proteins is partitioned into a training set and a validation set at a ratio of 4:1. The training set and the validation set (20%) are used for model training and cross-validation.
[0070] A2. Five-Fold Cross-Validation: In 80% of the training data, the five-fold cross-validation method is adopted: the data is divided into five equal parts, and each time one part is selected as the validation set (accounting for 20% of the total data), and the remaining four parts are used as the training set (accounting for 80% of the total data), and this is repeated five times.
[0071] This method ensures that each data sample is used for validation once, making full use of the data and improving the generalization ability of the model.
[0072] B. Model Training: B1. Training Data Preparation: During training, for each single-molecule surface-induced fluorescence decay signal trajectory, preprocessing such as normalization, linear interpolation, and Gaussian smoothing is performed, and finally the intensity data of the fluorescence signal is input to the model.
[0073] B2. Model Training Process: Parameter Settings: Using the vision transformer model based on the spatial perception unit, set the learning rate to 0.0001, the batch size to 64, and the optimizer to the optimizer with weight decay (Adam with Weight Decay Regularization, AdamW), and the weight decay to 0.01.
[0074] Training process: In each fold of cross - validation, the training set is used to train the model, and the test set is used to evaluate the performance of the model on unseen data. During the training process, the cross - entropy loss function is used as the loss metric. A warm - up mechanism is adopted, and the warm - up stage is specified as the first 5 epochs. The learning rate will gradually increase during this period to avoid unstable training caused by too high an initial learning rate. In each fold of training, the model parameters with the best performance on the validation set are saved for subsequent evaluation.
[0075] C. Model testing: C1. Test data preparation: In the model testing stage, for each trajectory of the single - molecule surface - induced fluorescence decay signal, pre - processing such as normalization, linear interpolation, and Gaussian smoothing is performed, and finally the intensity data of the fluorescence signal is input into the model. This fluorescence intensity data is a one - dimensional array with 1024 columns.
[0076] C2. Model prediction: The model predicts the trajectory of the single - molecule surface - induced fluorescence decay signal, obtaining the classification result of the trajectory and the softmax prediction score. The prediction score ranges from 0 to 1. The closer it is to 0, the greater the possibility that the model believes the curve is a noise signal, and the closer it is to 1, the greater the possibility that the model believes the curve is a single - molecule surface - induced fluorescence decay signal with high confidence.
[0077] C3. Performance evaluation: Calculate metrics such as the loss and accuracy of the model on the validation set and the independent test set. Statistically analyze the average results of five - fold cross - validation to evaluate the stability and generalization ability of the model.
[0078] D. Model optimization and adjustment: D1. Learning rate adjustment: Use the Cosine Learning Rate Scheduler, which will slowly decrease the learning rate during training, gradually decreasing from the initial learning rate to the minimum learning rate according to the cosine curve.
[0079] D2. Model parameter adjustment: According to the training and validation results, adjust the hyperparameters of the model, such as the learning rate, batch size, number of training epochs, etc., to optimize the model performance.
[0080] Through the above specific implementation steps, the training and testing of the vision transformer model based on the spatial perception unit are completed, verifying the effectiveness and reliability of the model in screening high - confidence single - molecule surface - induced fluorescence decay trajectories, and providing a powerful auxiliary tool for single - molecule fluorescence analysis.
[0081] Step 4: Training results of the vision transformer model based on the spatial perception unit After completing the training and validation of the model, we evaluated the performance of the model, such as Figure 4and Figure 5 As shown below, the specific results are as follows: A. Five-fold cross-validation training set performance: A1. Training set accuracy: During the five-fold cross-validation process, the average accuracy of the model on the training set reached 0.968, and the AUC was 0.996. This indicates that the model has a good performance in learning the characteristics of the training data and can effectively fit the training data.
[0082] B. Five-fold cross-validation validation set performance: B1. Validation set accuracy: The accuracy of the model on the validation set was 0.913, and the AUC was 0.978. This means that the model has a high accuracy in screening high-confidence single-molecule surface-induced fluorescence decay trajectories.
[0083] C. Test set performance: C1. MLKL_2 test set accuracy: On the MLKL_2 test set, the accuracy of the model reached 0.89, and the AUC was 0.948.
[0084] C2. LL-37 test set accuracy: On the LL-37 external test set, the accuracy of the model reached 0.844, and the AUC was 0.942. This shows that the model still has good predictive ability on unseen data, reflecting the generalization performance of the model.
[0085] D. Visualization of the model scores on the MLKL_2 test set: D1. Figure 6 Shows the scores of the model on the MLKL_2 test set and the corresponding fluorescence curves. A total of 9 traces are shown, where subfigures (a-c) represent three high-confidence fluorescence traces, and subfigures (g-i) represent three low-confidence fluorescence traces. In addition, subfigures (d-f) represent three ambiguous traces, with scores ranging from 0.3 to 0.7. The closer the score is to 1, the greater the likelihood that the model considers the trace to belong to a high-confidence curve, while the closer the score is to 0, the greater the likelihood that the model considers the trace to belong to a low-confidence curve. By adjusting the score threshold generated by the softmax function, more refined classification of different fluorescence intensity traces can be achieved.
[0086] Result analysis: The vision transformer model based on the spatial perception unit achieved high accuracy on the training set, validation set, and test set. The accuracy of the test set was slightly lower than that of the training set and the validation set, but still remained at a high level, indicating that the model has good generalization ability and can adapt to unseen trace data. In addition, multiple different deep learning models were compared, and the experimental results are as Figure 7As shown. Although other models also show the possibility of screening high-confidence single-molecule surface-induced fluorescence decay trajectories, their performance is generally inferior to that of the vision transformer model based on the spatial perception unit. These results further confirm the unique advantages of the vision transformer model based on the spatial perception unit.
[0087] From the above experimental results, it can be seen that the vision transformer model based on the spatial perception unit has good performance in quickly screening high-confidence trajectories from a large amount of single-molecule surface-induced fluorescence decay experimental data. The stable performance of the model on multiple datasets verifies its effectiveness and reliability. Our method provides a valuable means for accelerating and improving the classification and analysis of single-molecule time trajectories in biophysics and analytical chemistry.
[0088] The present invention also provides a computer-readable medium, on which computer-executable instructions are stored. When the executable instructions are executed by a processor, the above-mentioned deep learning classification method for screening high-confidence single-molecule fluorescence signals is implemented.
[0089] By having a computer read the executable instructions, the screening of the credibility of fluorescence traces is achieved.
[0090] The embodiments should not be regarded as limitations of the present invention, but any improvements made based on the spirit of the present invention should be within the protection scope of the present invention.
Claims
1. A deep learning classification method for screening high-confidence single-molecule fluorescence signals, characterized by: It includes the following steps: 1) Construct a visual transformer network model based on spatial perception units, taking the fluorescence intensity trajectory of each single-molecule surface-induced fluorescence decay signal as input, and output the classification results or / and softmax prediction scores of the trajectory after training, which includes: Several spatially aware units to embed multi-scale context and local information into tokens; Several visual transformer units to further model local features and long-range dependencies in tokens; The output layer maps the learned features to the class probability space and outputs the final classification result, which is used to determine whether the input fluorescence intensity trajectory belongs to a high-confidence trajectory; 2) Capture each single-molecule surface-induced fluorescence decay signal within a set time and obtain the fluorescence intensity trajectory of each single-molecule surface-induced fluorescence decay signal; 3) Preprocessing each fluorescence intensity trajectory in step 2) to obtain intensity data of the fluorescence signal; 4) Inputting the intensity data of the fluorescence signal obtained in step 3) into the visual transformer network model based on the spatial perception unit constructed in step 1), and outputting the classification result and / or the softmax prediction score.
2. A deep learning classification method for screening high-confidence single-molecule fluorescence signals according to claim 1, characterized in that: The space perception unit comprises: A spatial pyramid reduction module for downsampling and embedding input time series into tokens with rich multi-scale contexts by using multiple convolutions with different dilation rates; A multi-head self-attention module is used to globally model the relationship between single-molecule fluorescence time series segments and capture long-term dependency features; Multilayer perceptron is used to perform nonlinear transformation on the features after each layer of attention processing.
3. A deep learning classification method for screening high-confidence single-molecule fluorescence signals according to claim 1 or 2, characterized in that: A sequence-to-trajectory module is provided at the output end of each spatial perception unit.
4. A deep learning classification method for screening high-confidence single-molecule fluorescence signals according to claim 3, characterized in that: A trajectory conversion sequence module is provided between the space perception unit and the visual converter unit.
5. The deep learning classification method for screening high-confidence single-molecule fluorescence signals according to claim 1, characterized in that: The visual converter unit comprises: A multi-head self-attention module is used to globally model the relationship between single-molecule fluorescence time series segments and capture long-term dependency features; Multilayer perceptron is used to perform nonlinear transformation on the features after each layer of attention processing.
6. The deep learning classification method for screening high-confidence single-molecule fluorescence signals according to claim 1, characterized in that: The visual transformer network model based on the spatial perception unit includes three layers of spatial perception units and seven layers of visual transformers, and the last layer of visual transformers is connected to the fully connected layer.
7. A deep learning classification method for screening high-confidence single-molecule fluorescence signals according to claim 6, characterized in that: The three-layer spatial perception unit is divided into: The first layer of spatial perception unit: In the initial stage, the input channel is 1, the convolution kernel is 7, and multi-scale features are extracted through a large dilation rate. The dilation rate is dilations=[1, 2, 3, 4], and the number of output channels is 64; The second layer of spatial perception unit: in the intermediate stage, the number of output channels is 64, the convolution kernel is 5, and the dilation rate is reduced by dilations=[1, 2, 3]; The third layer of spatial perception unit: In the deep stage, the number of output channels is further increased to 320, the convolution kernel is 3, the dilation rate is the smallest, and the dilation rate dilations=[1, 2].
8. The deep learning classification method for screening high-confidence single-molecule fluorescence signals according to claim 6, characterized in that: The visual transformer network model based on spatial perception units consists of four stages, each with different input and output feature dimensions, specifically: Stage 1: The input feature dimension is 1, the convolution kernel of the expanded convolution is 7, the step size is 4, and the multi-scale features are extracted through a large dilation rate dilations=[1, 2, 3, 4], and finally the features with higher embedding dimensions are output; Stage 2: The input feature dimension is 64, the convolution kernel of the extended convolution is 5, the step size is 2, the dilation rate is reduced by dilations=[1, 2, 3], and fine-grained feature extraction is emphasized; Stage 3: The input feature dimension is 64, the convolution kernel of the extended convolution is 3, the step size is 2, and the dilation rate is the minimum dilations=[1, 2], focusing on fusing local and global features; Stage 4: The input feature dimension is 320, the number of attention heads is 6, and the visual transformer modules are stacked seven times.
9. The deep learning classification method for screening high-confidence single-molecule fluorescence signals according to claim 1, characterized in that: The preprocessing in step 3) includes: First, normalize the single-molecule surface-induced fluorescence decay signal data, and use the minimum-maximum normalization method to normalize the time series data to the set interval; Second, expand or compress the original time series to a uniform length through interpolation; 3. Perform Gaussian smoothing.
10. A computer-readable medium having computer-executable instructions stored thereon, characterized in that: When the executable instructions are executed by the processor, the deep learning classification method for screening high-confidence single-molecule fluorescence signals as described in any one of claims 1 to 9 is implemented.
Citation Information
Patent Citations
Monomolecular fluorescence event identification method based on local features
CN118983007A