Method and system for detecting characteristics of particles or particle environment and method for training machine learning model, computer program or computer readable medium

By analyzing particle diffusion trajectories and using machine learning models, virus types and intracellular regions are identified, solving the problems of low sensitivity and slow speed in existing virus detection methods. This enables rapid and accurate pathogen detection, simplifies equipment, and reduces costs.

CN121752884APending Publication Date: 2026-03-27OXFORD UNIVERSITY INNOVATION LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-07-23
Publication Date
2026-03-27

AI Technical Summary

Technical Problem

Existing virus detection methods, such as RT-PCR and rapid antigen tests, suffer from low sensitivity, slow speed, or high cost, making it difficult to quickly and accurately identify pathogens. In particular, during pandemics, detection delays can accelerate transmission.

Method used

By analyzing particle diffusion trajectory data, using detectable markers bound to particles and combining them with machine learning models, particle types and environmental characteristics, especially virus types and intracellular regions, are identified. By leveraging the interaction between diffusion properties and the particle environment, a rapid and sensitive detection method is provided.

Benefits of technology

It enables rapid and accurate identification of viruses, especially the differentiation of different virus types, and can detect particle concentration and environmental changes within cells, providing high-performance pathogen detection while simplifying equipment and reducing costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121752884A_ABST
    Figure CN121752884A_ABST
Patent Text Reader

Abstract

Methods and systems for detecting characteristics of a particle or particle environment are disclosed. In one arrangement, trajectory data representing diffusion trajectories of particles detected in one or more respective particle environments is received. Each detected diffusion trajectory includes a sequence of trajectory segments. Characteristics of the particles and / or characteristics of one or more particle environments of the particles are identified by analyzing the trajectory segments.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] This disclosure relates to the detection of characteristics of particles or particle environments by analyzing diffusion behavior. This disclosure is applicable to the diagnosis or quantification of pathogens (such as viruses or bacteria), and to the identification of changes in particle composition and / or transitions of particles between different environments (such as between different parts of a biological cell).

[0002] Viral outbreaks have affected the world's population for millennia. Despite centuries of medical advancements, the world remains vulnerable to pandemics. Since the 20th century, there have been over 50 epidemics, including six global pandemics. In 2019–2021, COVID-19 swept the world.

[0003] While COVID-19 appears to have a lower fatality rate than SARS, it exhibits more severe symptoms in susceptible populations. Furthermore, specific testing is required to confirm infection because it shares similar symptoms with influenza. Additionally, COVID-19 has a long incubation period, during which patients can infect others even without showing any symptoms themselves. These characteristics cause COVID-19 to spread faster than other viruses such as influenza. Particularly in many underdeveloped areas, the availability and speed of testing have not kept pace with the expansion of the pandemic, leading to delays in test results and leaving many patients untested. There is an urgent need to provide improved, rapid, sensitive, and inexpensive testing methods not only for responding to COVID-19 but also for responding to potential future pandemics.

[0004] Many viruses have an external lipid envelope. Both coronaviruses and influenza viruses are encased in a lipid bilayer. Influenza is characterized by an eight-segmented single-stranded RNA genome, where the genome constitutes the total genetic information of an individual organism. Based on the genome, influenza is further classified into type A, type B, or type C, with decreasing probability of causing epidemics. Influenza A virus causes the majority of influenza, especially those that spread globally and cause pandemics. Influenza A is further subdivided by its surface proteins: hemagglutinin (HA) and neuraminidase (NA). Eighteen possible subtypes of HA and eleven possible subtypes of NA have been identified, giving a total of 198 combinations, but only three HA (H1, H2, and H3) and two NA (N1 and N2) subtypes have caused human epidemics. In the case of coronaviruses, it constitutes a non-segmented single-stranded RNA genome. For surface proteins, coronaviruses are covered by spike proteins and have no further subtypes.

[0005] There are various commercially available virus diagnostic tests with varying degrees of sensitivity, specificity, and latency.

[0006] Reverse transcriptase PCR (RT-PCR) is the current gold standard for viral diagnostics, which directly detects the presence of viral RNA genomes in clinical samples. The assay involves a process of detection, amplification, and output measurement to identify and quantify the presence of infectious agents in a sample.

[0007] RT-PCR requires multiple temperature changes to drive the reaction cycles, involving complex and expensive thermal equipment. Another way of proceeding is to use isothermal methods, such as Reverse Transcription Loop-Mediated Isothermal Amplification-Based Assay (RT-LAMP).

[0008] Rapid antigen tests (RAT) detect viral surface proteins, often in combination with lateral flow assays to show a colorimetric result that can be read with the naked eye. The process is much faster than RT-PCR, but tends to be less sensitive.

[0009] Enzyme-Linked Immunosorbent Assay (ELISA) is a microplate-based assay technique designed to detect and quantify substances such as proteins, hormones, and antibodies (IgG / IgM). However, there can be a time lag in antibody production in the early stages of infection, which can lead to false negative results.

[0010] It is an object of the present disclosure to provide improved methods and systems for identifying pathogens, such as viruses. More generally, it is also an object of the present disclosure to provide methods that extend the possibilities of monitoring particles and their environment.

[0011] According to an aspect of the invention, there is provided a computer-implemented method of detecting a feature of a particle or of a particle environment, comprising: receiving trajectory data representing diffusion trajectories of particles detected in one or more respective particle environments, each detected diffusion trajectory comprising a sequence of trajectory segments; and identifying a feature of the particle and / or of one or more particle environments of the particle by analysing the trajectory segments.

[0012] Thus, there is provided a method that uses multiple detected trajectory segments and from which extensive information about the particle and / or its environment can be extracted. Using individual trajectory segments rather than overall diffusion distances or diffusion coefficients allows more detailed information about the particle and / or its environment to be obtained compared to alternative methods, harnessing highly subtle aspects of the interaction between diffusion characteristics and the properties of the particle and its environment.

[0013] In one embodiment, the identifying features comprises identifying a particle type of the one or more particles. The particles can comprise a pathogen, optionally a virus or a bacterium, and identifying the particle type comprises identifying a type of the pathogen. The inventors have demonstrated that this approach provides particularly high performance in discriminating between different types of pathogen. Furthermore, discriminating between pathogen based on monitoring of diffusion trajectories can be implemented quickly, and the equipment used is not overly complex or expensive.

[0014] In one embodiment, one or more detectable labels are bound to each particle, the particles comprise enveloped viral particles, and each detectable label is configured to bind to an envelope (e.g. lipid membrane) of the enveloped viral particles. This approach can be particularly effectively implemented. The detectable labels can be formulated to be suitable for a wide range of enveloped viruses without the need for individual adjustments, and the labelling process can be implemented almost instantaneously, which provides a significant speed advantage over competing viral identification techniques which can require significant incubation times etc. In one particularly convenient implementation, negatively charged nucleic acid fragments are attached to the negatively charged lipid membranes of enveloped viral particles by positively charged calcium ion groups as labels.

[0015] In one embodiment, the identifying features comprises identifying a particle environment type of the particle environment of the one or more particles. Identifying the particle environment type can comprise discriminating between a plurality of predetermined regions within a biological cell. The inventors have demonstrated that this approach is sufficiently sensitive to differences in diffusion characteristics that the presence of different environments can be detected even if the nature of the particles is unchanged. This allows sensitive measurement of the location of the target particles within the cell, allowing information to be extracted about the concentration of particles in different parts of the cell, and how the concentration can change over time.

[0016] In one embodiment, the predetermined regions comprise regions within a bacterium, such as a periplasmic region and a cytoplasmic region. The inventors have demonstrated that this approach has particularly high performance in detecting whether a protein is in the periplasmic region or the cytoplasmic region of a bacterium, providing a valuable tool for studying biological mechanisms involving the passage of proteins between these regions.

[0017] In one embodiment, the identifying features comprises detecting a change in a feature of the particle environment of the particle during the detected diffusion trajectory of the particle by analysing the trajectory segments of the detected diffusion trajectories. The inventors have demonstrated that this approach is able to detect changes in the particle environment as a function of time. The inventors have demonstrated particularly high performance in the case where the change in the feature of the particle environment comprises a transition between different regions of a biological cell (in particular a transition between a periplasmic region and a cytoplasmic region of a bacterium).

[0018] In one embodiment, identifying the characteristic is performed by inputting the trajectory data to a trained machine learning model. The inventors have found that training a machine learning model to identify a link between a sequence of trajectory segments and a characteristic of a particle and / or a particle environment provides particularly high performance.

[0019] In one embodiment, the trained machine learning model is trained using a k- nearest neighbour machine learning algorithm. The inventors have demonstrated that a k- nearest neighbour machine learning algorithm provides a favourable balance between training efficiency and performance. However, other classification algorithms can also be used.

[0020] In one embodiment, the method comprises pre-processing the trajectory data by a dimensionality reduction algorithm, optionally comprising principal component analysis, prior to inputting the trajectory data to the trained machine learning model. It has been found that using a dimensionality reduction algorithm in this case can improve training efficiency and robustness.

[0021] In one embodiment, the method comprises capturing video data of a particle in a particle environment or a plurality of particle environments, and processing the video data to provide the trajectory data. Processing of the video data can comprise tracking the position of an individual particle over a sequence of frames of the video data to determine a corresponding detected diffusion trajectory. The detected diffusion trajectories can be selectively included in the trajectory data, such that only a subset of detected particles contribute to the trajectory data. Selectively including the trajectory data improves performance by allowing trajectory data of lower relevance and / or potentially ambiguous and / or erroneous to be omitted.

[0022] In one embodiment, the selection is based on a determined diffusion velocity of the detected particle. The inventors have recognised that diffusion characteristics (such as an expected diffusion velocity) of a target particle can be known in advance or calculated, and this can be exploited to exclude trajectory data which is unlikely to correspond to a target particle of interest, thereby improving performance.

[0023] In one embodiment, the selection comprises excluding a detected diffusion trajectory of a particle based on a determination that:

[0024] the change in position between temporally adjacent frames or frames separated by a predetermined number of frames of the particle is above a predetermined maximum threshold; and / or the change in position between temporally adjacent frames or frames separated by a predetermined number of frames of the particle is below a predetermined minimum threshold. Thus, particles which diffuse too quickly and / or particles which diffuse too slowly can be omitted from the trajectory data.

[0025] In one implementation, the selection includes excluding detected diffusion trajectories of a particle based on the following determination: in one or more frames of video data, the distance between the particle and detected diffusion trajectories of different particles is less than a predetermined minimum distance, optionally wherein the predetermined minimum distance is approximately equal to a predetermined maximum threshold. Excluding diffusion trajectories that are too close reduces or avoids the risk of overlapping diffusion trajectories, which could lead to errors or blurring.

[0026] According to an alternative aspect of the invention, a method for training a machine learning model is provided, comprising: receiving labeled training data including data representing a plurality of detected diffusion trajectories of particles (each detected diffusion trajectory including a sequence of trajectory fragments) and data representing a label associated with each detected diffusion trajectory, the label indicating features of the particle and / or particle environment corresponding to the diffusion trajectory; and training a machine learning model using supervised learning to identify features of the particle and / or particle environment based on the trajectory fragments.

[0027] The embodiments disclosed herein will be further described by way of example only with reference to the accompanying drawings.

[0028] Figure 1 It is a flowchart describing the elements of a method for detecting the characteristics of particles or a particulate environment.

[0029] Figure 2 Describes the use of execution Figure 1 Example system 2 for the method.

[0030] Figure 3 This is a flowchart depicting the elements of a method for training a machine learning model in a process used to detect features of particles or a particulate environment.

[0031] Figure 4 A depicts the imaging of a sample using a wide-field microscope in HILO mode.

[0032] Figure 4 B is Figure 4 An enlarged view of one of the viral particles 14 shown in Figure A.

[0033] Figure 4 C schematically depicts single-particle tracking that allows for observation of viral particle dynamics.

[0034] Figure 5 The layout of an algorithm that includes trajectory segmentation is schematically depicted. A long trajectory (spreading trajectory) is divided into smaller trajectories (containing fewer trajectory segments than the long trajectory).

[0035] Figure 6 The KNN classification is described.

[0036] Figure 7Exemplary detected diffusion trajectories are depicted: (A) SARS-CoV-2; (B) Udorn; (C) WSN; (D) IBV; (E) X31; (F) PR8. L = 1.153 pm.

[0037] Figure 8 Confusion matrices are depicted: viruses vs. negative media samples (PCA on). Diagonal percentages represent the classification performance for that class, e.g., for the true class ALLA, 94.4% were correctly predicted, 5.6% were not. Overall validation accuracy is obtained by averaging the diagonal numbers. (A) DMEM vs. Udorn, overall validation accuracy 94.3%. (B) ALLA vs. X31 vs. IBV, overall validation accuracy 94.2%. (C) MEM vs. WSN vs. SC2 vs. WSN, overall validation accuracy 78.5%.

[0038] Figure 9 Confusion matrices are depicted: all influenza A viruses. 11,000 trajectories were used for each virus. (A) PCA off, overall accuracy 96.3%. (B) PCA on, overall accuracy 89.8%.

[0039] Figure 10 Confusion matrices are depicted: all influenza A viruses and IBV. PCA on. (A) 55,000 trajectories were used for each virus, overall accuracy 87.3%. (B) 27,500 trajectories were used for each virus, overall accuracy 84.0%.

[0040] Figure 11 Colloidal electrolyte is depicted. V: virus particle. Fluorescently labeled ssDNA and calcium ions are shown. The situation in solution is complex. For example, two viruses that are close to each other at a distance can experience an attractive force and aggregate under the action of Ca ions in the middle.

[0041] Figure 12 Confusion matrices are depicted that describe the performance of a KNN model trained using TorA-halotag (uninduced) as periplasmic control and TorAKK-Halotag (induced) as cytoplasmic control. 70% of the data was used to train and validate the model.

[0042] Figures 13-16 is a graph showing the relative classification score of detected trajectories (diffusion trajectories). In each graph, the upper plot shows the relative classification score (vertical axis) as a function of subtrajectory (horizontal axis). The lower plot shows the average step size (vertical axis) as a function of subtrajectory (horizontal axis). Figure 13 is an example of a pure cytoplasmic trajectory, as expected, each subtrajectory scores close to 1.Figure 14 is an example of a pure periplasmic trajectory, as expected, each sub-trajectory score is close to -1. Figure 15 and Figure 16 depicts an example of a transition event detected in a dataset.

[0043] The present disclosure includes computer-implemented methods. Each step of such methods can be performed by a computer in the broadest sense, meaning any device capable of performing the data processing steps of the method, including special-purpose digital circuits. The computer can include various combinations of known computer elements, including, for example, CPUs, RAM, SSDs, motherboards, network connections, firmware, software, and / or other elements known in the art that allow the computer to perform the required computational operations. The required computational operations can be defined by one or more computer programs. The one or more computer programs can be provided in the form of a medium or data carrier, optionally a non-transitory medium, storing computer-readable instructions. When the computer reads the computer-readable instructions, the computer performs the required method steps. The computer can consist of a standalone unit, such as a general-purpose desktop computer, laptop, tablet, mobile phone, or other smart device. Alternatively, the computer can consist of a distributed computing system with multiple different computers connected to each other via a network, such as the Internet or an intranet.

[0044] Methods using machine learning are described below. The concept of machine learning encompasses a class of algorithms that can learn from the data provided to them. Machine learning can be divided into two broad categories: supervised learning and unsupervised learning. In supervised learning, the training set fed into the algorithm is labeled. The goal of supervised learning is to find a function that can accurately map the input data to the desired output. In contrast, in unsupervised learning, the training set is unlabeled, so the algorithm needs to learn and teach itself.

[0045] Examples of known supervised algorithms include k-Nearest Neighbour (KNN), Linear Regression, Decision Trees, and Random Forest. Supervised learning can be used to solve classification and regression problems. Classification involves assigning labels to discrete outputs, while regression gives continuous output values (e.g., predicting quantities such as house prices, temperatures, sizes, etc.). Some methods of the present disclosure aim to distinguish between different particles (e.g., pathogens), such as distinguishing between different viruses, and / or distinguishing between different particle environments, and can therefore use a classification prediction model.

[0046] Figure 1 is a flowchart depicting elements of a method of detecting features of a particle or particle environment according to the present disclosure. Figure 2An example system 2 for performing the method is depicted. The system 2 comprises a sample receiving device 4. The sample receiving device 4 is configured to receive a sample 6 to be tested. The sample receiving device 4 can comprise any suitable combination of components that allow the sample 6 to be input to the device 4 and stored in a manner that allows measurements to be made on the sample 4. The sample 6 comprises particles and / or particle environments of interest to be tested. Typically, the sample 6 will be a liquid, for example an aqueous solution.

[0047] In one embodiment, the system 2 comprises an image capture device 8. The image capture device 8 is configured to capture electromagnetic radiation, such as visible, IR or UV radiation, emitted from the sample 2. The system 2 further comprises a data processing device 10. The data processing device 10 comprises any suitable combination of data processing hardware, firmware, software and the like required to perform the required data processing functions. The data processing device 10 can comprise various combinations of known computer elements including, for example, CPUs, RAM, SSDs, motherboards, network connections, firmware, software and / or other elements known in the art that allow a computer to perform the required computational operations. The required computational operations can be defined by one or more computer programs. The one or more computer programs can be provided in the form of a medium or data carrier, optionally a non-transitory medium, storing computer readable instructions. A computer readable medium can thus be provided comprising instructions which, when the program is executed by a computer of the data processing device 10, cause the computer to carry out the method of the present disclosure. The computer readable medium is one example of a computer program product, and can be referred to as a computer program product. The data processing device 10 is configured to perform the computer-implemented method described below. The data processing system 10 can output information, such as features of detected particles and / or particle environments, to a display 12, or as an output data stream to an internal or external network or storage device.

[0048] In step S1 of the method, trajectory data is received, for example by the data processing device 10 or a data processor within the data processing device 10. The trajectory data represents diffusion trajectories of particles detected in one or more respective particle environments. The trajectory data can represent detected diffusion trajectories of multiple particles in the same particle environment or multiple particles in multiple particle environments. Each detected diffusion trajectory comprises a sequence of trajectory segments. Each trajectory segment comprises a sub-portion of a longer continuous trajectory followed by a respective particle. The continuous trajectory is thus formed from a plurality of trajectory segments. The trajectory data can be sourced from electromagnetic radiation captured by the image capture device 8. Further details of a specific example are described below with reference to Figure 4 C is given.

[0049] In one embodiment, the image capture device 8 is configured to acquire video data of particles in a granular environment or multiple granular environments. In the examples and other embodiments described below, video data is captured using a wide-field microscope, optionally configured to operate in a highly inclined and laminated optical sheet (HILO) mode.

[0050] Video data can be processed to obtain trajectory data. Video data processing can be performed by the imaging capture device 8 or by the data processing device 10, or both. Video data processing can include tracking the position of individual particles on a sequence of video data frames.

[0051] In one embodiment, a detectable tag is attached to each particle. The detectable tag can be configured, for example, to be detected by capturing electromagnetic radiation emitted from the detectable tag. The detectable tag may include a fluorescent tag.

[0052] In one embodiment, the particle includes an enveloped viral particle, and a detectable marker is configured to bind to the envelope of the enveloped viral particle.

[0053] In step S2 of the method, the characteristics of the particle and / or one or more particle environments of the particle are identified by analyzing the trajectory fragments.

[0054] In one embodiment, the identification feature includes identifying the particle type of one or more particles. In one embodiment, the particles include pathogens, and identifying the particle type includes identifying the type of pathogen. The pathogen can be a virus or bacteria.

[0055] In some implementations, feature identification is performed by feeding trajectory data into a trained machine learning model. Figure 3 It is a flowchart depicting the elements of a method for training a machine learning model.

[0056] Step M1 of this method includes receiving labeled training data. The labeled training data includes data representing multiple detected diffusion trajectories of particles. References above can be used... Figure 2 The described image capture apparatus is used to detect diffusion trajectories. Each detected diffusion trajectory includes a sequence of trajectory segments. The labeled training data also includes data representing labels associated with each detected diffusion trajectory. The labels indicate features corresponding to the particles and / or particle environment (e.g., particles diffusing along the diffusion trajectory and the particle environment in which the diffusion trajectory occurs).

[0057] Step M2 of the method includes using supervised learning to train a machine learning model to identify particles and / or features of the particle environment based on trajectory fragments.

[0058] In principle, various types of machine learning models can be used. In the example described below, the k-nearest neighbors machine learning algorithm is used to train the machine learning model. In some implementations, the trajectory data is preprocessed using a dimensionality reduction algorithm before being input into the trained machine learning model. Applying dimensionality reduction algorithms improves the speed and robustness of the training process. Dimensionality reduction algorithms may include, for example, principal component analysis.

[0059] In some implementations, as illustrated below, one or more particulate environments include an electrolyte. In this case, the electrolyte conditions input to the trajectory data of the trained machine learning model and the electrolyte conditions used during training the machine learning model can be arranged to be substantially the same in at least the following parameters: pH; temperature; chemical composition. By keeping the particulate environmental aspects affecting diffusion identical to the training data of the particles of interest, the trained machine learning model can more effectively identify aspects of the trajectory data indicative of particle properties. This can be particularly important in the context of applications for identifying pathogens, such as target bacteria or viruses.

[0060] Pathogen identification examples

[0061] To demonstrate the methods disclosed herein, two coronaviruses and four influenza A virus strains were tested:

[0062] COVID-19 (SARS-CoV-2);

[0063] IBV (Infectious Bronchitis Virus);

[0064] An H1N1 strain, PR8 (H1N1 A / Puerto Rico / 8 / 1934);

[0065] Another H1N1 strain, WSN (H1N1 A / WSN / 1933);

[0066] An H3N2 strain, Udorn (H3N2 A / Udorn / 1972); and

[0067] Another H3N2 strain, X31 (H3N2 A / X-31).

[0068] These viruses were compared with their negative growth media: allantoic fluid (ALLA), minimum essential medium Eagle (MEM), and Dulbecco modified Eagle medium (DMEM).

[0069] like Figure 4A schematic depiction of imaging a virus using a wide-field microscope in HILO (High Tilt Layered Slide) mode. Using a wide-field microscope in HILO mode illuminates more virus in solution than evanescent waves in Total Internal Reflection Fluorescence (TIR). The microscope is an example of image capturing device 8. A slide 17 is used as sample receiving device 4. Slide 17 is mounted on objective lens 16 via an oil layer 18. The schematic trajectory of the incident laser is marked 20. The laser incident angle is slightly smaller than the TIR angle, allowing light to penetrate into sample 6.

[0070] Figure 4 B is Figure 4 An enlarged view of one of the virus particles 14 shown in Figure A. Figure 4 B schematically illustrates how calcium ions are used to facilitate the interaction between the negatively charged phospholipid layer and the phosphate groups on nucleic acids. Viruses are labeled by mixing three components: a viral sample, a buffered calcium chloride (CaCl2) solution, and single-stranded DNA (ssDNA). Enveloped viruses have a net negative charge due to their lipid membrane, and ssDNA is also negatively charged due to its phosphate backbone. The positively charged calcium ion groups aggregate the negatively charged viral envelope and ssDNA together, thereby achieving... Figure 4 The label shown in B. The negatively charged nucleic acid fragment 24 thus acts as a label to attach to the lipid membrane of the enveloped virus. This labeling method is universal (effective for any enveloped virus) and instantaneous (no incubation time required).

[0071] Fluorescently labeled viruses were detected using single-particle tracking analysis software. This software locates fluorescent molecules by searching for corresponding intensity peaks significantly above the background, and then fits the signal with a Gaussian function to obtain more precise localization. Figure 4As schematically depicted in Figure C, the locations of a single viral particle across multiple frames can be connected into a continuous trajectory 26, allowing observation of particle dynamics. The continuous trajectory 26 comprises multiple trajectory segments 28A-E. Trajectory segments 28A-E are obtained by tracking the position of a single particle 14 across a sequence of video data frames. For example, each of the trajectories 28A-E can represent the diffusion motion of a particle across a predetermined number of frames in the video data (e.g., between corresponding pairs of adjacent frames). For example, in the case of consecutively numbered frame sequences, the trajectory sequence can correspond to a sequence of particle displacements representing differences in particle position: between frame 1 and frame 2; between frame 2 and frame 3; between frame 3 and frame 4, etc. Alternatively, the differences in particle position can be: between frame 1 and frame 3; between frame 3 and frame 5; between frame 5 and frame 7, etc. A range of other options are also possible. However, typically, each trajectory 28A-E occurs within the same predetermined time period, which can correspond to a time period between frames in the video data or a time period between a predetermined number of frames in the video data. Multiple particles can exist on the same frame. As described below, an exclusion radius (which can be defined, for example, based on how far a particle can travel in the time between two frames) can be used to help distinguish particles in subsequent frames. If we have 30 frames and the particle exists in all of them, we get 30 localizations. A subset of these, any localization between 2 and 29, can be defined as a trajectory segment.

[0072] The inventors have discovered that by imposing constraints on which detected particles 14 contribute to the trajectory data, the extraction of relevant trajectory data from video data can be made more efficient and / or robust. Therefore, detected diffusion trajectories can be selectively included in the trajectory data, such that only a subset of the detected particles 14 contributes to the trajectory data.

[0073] In one arrangement, the selection is based on a determined diffusion rate of the detected particle 14. For example, the selection may seek to exclude particles 14 that appear to be diffusing much faster or slower than the particle 14 of interest is expected to. In one embodiment, the selection includes excluding detected diffusion trajectories of particle 14 based on the determination that the positional change of the particle between temporally adjacent frames or between frames separated by a predetermined number of frames is greater than a predetermined maximum threshold. The predetermined maximum threshold may be referred to as the maximum step distance. Alternatively or additionally, the selection includes excluding detected diffusion trajectories of particle 14 based on the determination that the positional change of the particle between temporally adjacent frames or between frames separated by a predetermined number of frames is less than a predetermined minimum threshold. The predetermined minimum threshold may be referred to as the minimum step distance.

[0074] Applying such maximum and / or minimum step distances can increase the probability that the retained trajectory fragments correspond to the particle of interest. For example, as described below, when the particle of interest is a viral particle, applying the maximum step distance can exclude particles expected to spread faster than viral particles, such as free dyes (e.g., ssDNA). Particles 14 that appear to move beyond the maximum step distance are also more likely to be false detections, for example, corresponding to different particles detected in different video frames being mistakenly identified as the same particle.

[0075] Alternatively or additionally, this option includes excluding a detected diffusion trajectory of a particle based on the following criterion: the particle is less than a predetermined minimum distance from the detected diffusion trajectories of different particles in one or more frames of the video data. This may be referred to as the exclusion radius. Applying such an exclusion radius reduces the risk of particle trajectories overlapping each other and causing confusion about which particle belongs to which trajectory segment.

[0076] In this specific embodiment, a sample was prepared by mixing 3 μL of virus into 15 μL of 0.45 M CaCl2 buffer solution, along with 1 nM fluorescently labeled DNA, resulting in a final volume of approximately 20 μL. An intensity of 0.78 kW / cm² was used. 2 Green (532nm) and red (635nm) lasers were used to image the glass slide 17. The exposure time for each film (video) was 33.3ms.

[0077] Before running point tracking in the single-particle tracking analysis software, two values ​​are predefined for the maximum step distance and the exclusion radius. These parameters are used to ensure that the tracked particles are viral particles and not any other type of particle, such as unbound DNA. In this embodiment, the maximum step distance is used to exclude free dye (ssDNA). Free dye is smaller and diffuses faster in solution than the larger viral particles of interest. As mentioned above, the exclusion radius is used to exclude intersecting trajectories to prevent ambiguity in trajectory assignment. Particles expected to diffuse more slowly, such as cells or viral aggregates, can be excluded by applying an appropriate minimum step distance.

[0078] Appropriate values ​​for the maximum step distance and exclusion radius can be determined by calculating the virus spread step size using the Stokes-Einstein relation:

[0079]

[0080] Where D is the diffusion coefficient, k B Here, η is the Boltzmann constant, T is the temperature, η is the viscosity, r is the radius of the virus particle, Δt is the exposure time (e.g., the time between temporally adjacent frames), and Δx is the viscosity. 2 It is the mean square displacement.

[0081] Then calculate Δx using the following expression. 2 :

[0082]

[0083] Then use this expression to evaluate the maximum step size L in 2D according to the following formula:

[0084]

[0085] The exclusion radius r is estimated using the following formula. exc :

[0086]

[0087] Where N is the average number of particles per unit area. If the exclusion radius is less than the maximum distance, both parameters will use the maximum distance value. In this implementation, the exclusion radius is set to be equal to the maximum step size.

[0088] Machine learning models require large amounts of data for training. During training, the model adjusts its parameters to better map the input to the output. However, data is often only available for a limited time period, resulting in a small training set.

[0089] To ensure sufficient data is available for training, the training set can be artificially expanded during the preprocessing stage. In this embodiment, a minimum trajectory length of 10 steps (e.g., corresponding to 10 trajectory segments) is selected to ensure data standardization. Larger trajectories 51 in the original dataset 61 are divided into sub-trajectories 52 in the corresponding subset dataset 62, such as... Figure 5 As illustrated schematically, subset 62 is used to form table 70, which includes a validation set and a training set for training machine learning model 72. Each trajectory i contains n i Step (trajectory segment), then each trajectory i is divided into smaller trajectories v. i Each trajectory contains 20 locations (10 x-coordinates and 10 y-coordinates), making

[0090]

[0091] Smaller step sizes (5 and 7) were also experimented with. While these produced a large amount of data, accuracy decreased slightly. Larger step sizes (e.g., 20) also resulted in a decrease in accuracy due to the smaller dataset. A step size of 10 provided a good balance between sufficient data volume and good model accuracy. The reason for using consecutive steps was to preserve information about the step size, which is likely to be picked up as a classification feature by the machine learning algorithm. The sub-dataset was then transformed into an expanded table 70, which was later fed into the model 72.

[0092] K-Nearest Neighbors (KNN) is a supervised machine learning algorithm primarily used for classification. It involves the fundamental assumption that similar inputs lead to similar outputs, meaning that data points that are close together are more likely to be classified into the same category, such as... Figure 6 As illustrated. The algorithm can be formalized as follows: For a dataset D and a test point x, there exists a set of k nearest neighbors x, denoted as Sx, where Make:

[0093]

[0094] Where |Sx| = k and (x′, y′) are points inside Sx, while (x′′, y′′) are points outside Sx but inside D. x represents the position, and y represents the label of that point. The KNN classifier can be defined as the following function h():

[0095]

[0096] Here, `mode()` selects the tag with the highest number of votes. In summary, KNN assigns the tag that appears most frequently in its k nearest neighbors to the test point `x`. This method... Figure 6 The diagram illustrates the process. The central circle 80 represents the trajectory to be classified. When k=1, the trajectory corresponding to circle 80 receives one vote from its nearest neighbor in group 81 and is assigned a label corresponding to group 81. When k=3, the trajectory corresponding to circle 80 receives two votes from its nearest neighbor in group 82, and only one vote from its nearest neighbor in group 81, and is assigned a label corresponding to group 82.

[0097] In this embodiment, Fine KNN is used, where k=1.

[0098] Recall that the KNN classifier assumes points that are close to each other share similar labels. However, in higher dimensions, data points start to drift further apart. This problem stems fundamentally from insufficient data available for the expanded dimensions and the counterintuitive geometric properties of high-dimensional spaces. Increasing dimensionality increases the size of the data space, so the amount of data also needs to expand to maintain the same density level. However, the increase in data size to compensate for the dimensionality expansion is disproportionate. Furthermore, while the workings of the KNN classifier in 2D and 3D spaces seem obvious, it's difficult to imagine the classification process in higher dimensions. In our case, we used 20-dimensional data (10 x-coordinates, 10 y-coordinates). These problems can be addressed using dimensionality reduction. Principal component analysis (PCA) is an example implementation of dimensionality reduction, described below.

[0099] Dimensionality reduction is achieved by identifying the truly important dimensions in a dataset.

[0100] Suppose the data input to the KNN classifier has low intrinsic dimensionality: the data lies in a low-dimensional submanifold or low-dimensional subspace. For example, high-dimensional points can be fitted to a 3D plane, and then further evaluation can be performed by looking only at the projection in that 3D plane. Uniformly distributed data is not useful for classification.

[0101] Principal Component Analysis (PCA) is a method that first identifies the hyperplane that best approximates the data, and then projects the data onto it. It begins by maximizing the square of the projected length. The sum of these values ​​is used to identify the axis with the largest variance in the training set. This is equivalent to minimizing the squared distance from each data point to the projected axis. However, computers can handle projection more easily. If the training set has a total dimension d, PCA will find up to d mutually orthogonal axes. The unit vector of the i-th axis is defined as the i-th principal component (PC). To determine which axes will be used to construct the hyperplane on which points will be projected, the variance ratio needs to be calculated. The variance is calculated as follows:

[0102]

[0103] Calculate the variance for each PC, and then evaluate the proportion of variance of the dataset across each PC using the variance ratio:

[0104]

[0105] To configure the classifier to capture at least 95% of the variance, we increase the ratio from maximum to 95%. The chosen PC defines the hyperplane. Therefore, this approach does not arbitrarily choose the number of dimensions to reduce to, but rather ensures that the projection preserves as much variance as possible.

[0106] This model is used in the context of clinical testing. To understand its practicality in clinical testing, the following terminology is introduced:

[0107] True positive (TP): The patient has the disease and the test is positive;

[0108] False positive (FP): A patient does not have the disease but tests positive;

[0109] True negative (TN): The patient does not have the disease and the test is negative;

[0110] False negative (FN): The patient has a disease but the test is negative.

[0111] To evaluate clinical tests, sensitivity and specificity are used. Sensitivity represents the test's ability to correctly identify patients with the disease. It is the ratio of the number of patients successfully detected to the actual number of patients with the disease:

[0112]

[0113] The higher the sensitivity, the better the test can identify all patients with the disease.

[0114] The specificity of a clinical test refers to its ability to correctly identify patients who do not have the disease. It is the ratio of true negative cases to the total number of negative cases.

[0115]

[0116] The probability of a patient having or not having the disease and receiving a positive or negative result was assessed using positive predictive value (PPV) and negative predictive value (NPV), respectively:

[0117]

[0118] The experimental conditions were: temperature T = 33℃, and the viscosity was approximately equal to the viscosity of water at 33℃, η = The radius of the marker virus used The exposure time is Δt = 33.3 ms. Substituting these values ​​into the previously given equations related to D and L yields...

[0119]

[0120] Input these values ​​into single-particle tracking analysis software to obtain preliminary trajectory results, such as... Figure 7 As shown, the trajectories vary in length, size, and distribution. It's difficult to distinguish their types with the naked eye, but machine learning algorithms can extract useful information. Example parameter values ​​of 1.153 and 1.412 μm were used as the basis for extracting trajectories using single-particle tracking analysis software and tested with a classifier. In terms of validation accuracy, the results were similar, but slightly better for L=1.153 μm, which involves deriving more trajectories from the single-particle tracking analysis software. Therefore, all tests in the following sections utilize data obtained from the green and red channels using L=1.153 μm.

[0121] Before training the model, the dataset size needs to be adjusted to include equal contributions from all training samples. This resizing is done by identifying a suitable minimum viral data size after augmentation, then randomly selecting the same number of data points from all virus groups involved in training, and storing the remaining data for later testing. In this embodiment, due to the limited amount of SARS-CoV-2 (SC2) data, any training trials involving SC2 use 3000 trajectories per virus, while other training trials use 11000 trajectories per virus.

[0122] First, examine the model's ability to distinguish between viruses and their negative environmental results. For laboratory-cultured viruses, they are grown in different media because the nutritional requirements supporting viral growth differ (see table below).

[0123] Table: Virus and its corresponding negative control

[0124]

[0125] Comparing IBV to DMEM is unwise because DMEM is not its true negative. Therefore, the experimental organization was as follows: MEM vs. WSN vs. PR8 vs. SC2; DMEM vs. Udorn; ALLA vs. X31 vs. IBV.

[0126] After determining its ability to distinguish negative media, the model is further trained to differentiate viruses. Start with a simple pair, such as SC2 versus PR8, since they grow on the same medium MEM, and then gradually add more viruses to the training.

[0127] The KNN model with PCA enabled provided satisfactory results in distinguishing the virus from its negative counterpart. Except for the one involving SC2 (such as...). Figure 8 (As shown in the confusion matrix), the validation accuracy was above 90%. The low validation accuracy for SC2 was mainly due to its limited data volume. The training set consisted of 3,000 randomly selected data points from 3,527 SC2 data points and 3,000 selected data points from 30,680 MEM data points. The model learned from less than 10% of the data from MEM, capturing only a small number of features, resulting in low MEM recognition accuracy (69.6%). In contrast, using the majority (84.6%) of the SC2 data, a high validation accuracy of 87.2% was achieved. Two further tests were conducted comparing viruses with MEM (excluding SC2), with significantly improved results: MEM achieved 91.9% accuracy against WSN and 88.1% accuracy against PR8.

[0128] For comparisons between viruses, although the total trajectories of PR8 versus SC2 were only 6000 in a small sample, the overall results were good, yielding an overall validation accuracy of 92.3%. Then, all influenza A samples (PR8, Udorn, WSN, X31) were pooled to train the model to distinguish between subtypes and viruses. Figure 9 B). The confusion matrix shows that the model can identify X31 and PR8 better than other strains. However, it is worth noting that the validation accuracy is much higher when PCA is turned off ( Figure 9 (A) This means that some attributes may be lost during the dimensionality reduction process. The same situation occurred when comparing influenza A virus with IBV or SC2. With PCA disabled, the accuracy was 95.5% and 85%, respectively, while with PCA enabled, it was only 87.3% and 80.8%.

[0129] The information loss caused by PCA exploitation can ultimately be overcome by increasing the training sample size. To further verify this, the model was trained with different numbers of trajectories from the same type of virus. Figure 10 The reason only IBV is included here is because this combination provides the largest amount of data. Models trained on more data are generally more reliable than those trained on less data. Figure 10 The results show that model performance improves as the training set increases: when the amount of data is halved, the overall validation accuracy decreases by 3.3%.

[0130] Finally, all viruses were trained together. Again, due to the limited sample size of SC2, only 3,000 tracks were used, and all other viruses needed to maintain the same number of tracks as SC2. The overall training set was relatively small (18,000 tracks in total), resulting in lower overall accuracy: 77.2% with PCA enabled and 82.2% with PCA disabled.

[0131] Based on the results discussed above, the model appears to perform better with PCA disabled, without considering testing speed. However, testing speed is a crucial factor for practical applications.

[0132] As can be seen from the table below, compared with the case without PCA, the model with PCA has lower accuracy, but its prediction speed is much better.

[0133] Table: This table summarizes the model's performance in virus differentiation. Test data volume: All influenza A: 44,000 tracks; Influenza A vs. IBV: 55,000 tracks; All viruses: 18,000 tracks.

[0134]

[0135] The speed disadvantage may be negligible for small datasets, but the cumulative effect over thousands of tests becomes significant. Furthermore, the accuracy deficiency can be addressed by increasing the amount of data, which in turn benefits machine learning algorithms as the model becomes more reliable with more information fed in. Additionally, as can be seen from the table above, the prediction speed with PCA enabled varies greatly with the amount of training data: 2,300 object / s for all influenza A and IBV (55,000 total trajectories), and 7,800 object / s for all viruses (18,000 total trajectories). This change can be understood through the workings of KNN. A refined KNN model classifies test points by finding the nearest point and mapping the label to the test point. Searching and determining the distance to all surrounding data points in a high-dimensional space presents a greater challenge than in a low-dimensional space. It is certain that the model can produce perfect results without PCA, but the prediction speed will be severely hampered.

[0136] Therefore, a large amount of training data is unavoidable in order to have a faithful and applicable model. In this case, high prediction speed will be important, and the information loss in PCA will be compensated by more data input.

[0137] Some data was randomly excluded from training. This unseen data was used to examine how the model functioned. The table below shows the obtained test accuracy, demonstrating relatively satisfactory results, meaning the model was able to distinguish trajectories taken from the same experimental trail. Using a model with PCA enabled, the latency for all these tests was less than 1 second.

[0138] Table: The table below summarizes the test accuracy for each model.

[0139]

[0140]

[0141] Depending on the amount of data and the type of virus, the sensitivity of these models ranges from 71.6% to 91.7%, while the specificity ranges from 94.2% to 97.3% (see table below).

[0142] Table: Statistical Results. SN: Sensitivity; SP: Specificity; PPV: Positive Predictive Value; NPV: Negative Predictive Value

[0143]

[0144]

[0145]

[0146]

[0147] Strictly speaking, viruses labeled with ssDNA in calcium chloride solution are colloidal electrolytes, where normal diffusion is no longer suitable for describing their movement. Colloidal particles, or colloids, are small solid particles suspended in a fluid phase. In our case, the labeled viruses are colloids. A colloidal electrolyte is a solution of a salt in which one of the ions has been replaced by the colloid. For example, NaR is the sodium salt of an organic acid, where R is an anion, forming a colloid whose dissociation can be written as:

[0148]

[0149] Electrostatic interactions influence viral movement, making its dynamics unpredictable. The effects of electricity render the Stokes-Einstein equation inappropriate. The discrepancy between theory and reality can be seen by substituting the experimental diffusion coefficient back into the Stokes-Einstein equation, which produces radii approximately 2 to 7 times larger than typical viral sizes. This is likely due to charge interactions disrupting diffusion or viral aggregation. Furthermore, the presence of freely moving ssDNA and its interactions with surrounding charged particles create further disruptions, such as… Figure 11 Schematic illustration. Here, viral particles are labeled "V", fluorescently labeled ssDNA is labeled "90", and calcium ions are labeled "92". The situation in this solution is complex; for example, two viral particles V approaching each other at a certain distance may experience attraction and aggregate under the influence of the Ca ions 92 in the middle. To date, there is no accurate model to describe the diffusion in this colloidal electrolyte. However, approximating the situation as normal diffusion is useful for obtaining preliminary results on the trajectories used for analysis.

[0150] In the above embodiments, the identification features in the method include identifying particle type. In other embodiments, as illustrated below, the identification features include identifying particle environment type. In specific embodiments below, identifying particle environment type includes distinguishing multiple predetermined regions within a biological cell. Different regions within a biological cell may have properties that affect how particles (e.g., proteins) diffuse within them. This method utilizes these differences in properties to obtain information about the location of particles at a given time. This method is particularly useful in the case of bacteria. In this case, the predetermined regions may include different regions within the bacteria. As illustrated below, the different regions may include a periplasmic region and a cytoplasmic region. Therefore, this method can be used to detect the location of particles such as proteins within bacteria, e.g., whether the protein is in the periplasmic region or the cytoplasmic region, and to monitor the transition of the membrane separating the two regions.

[0151] Protein transmembrane transport is an important process in bacteria because it allows cells to properly divide proteins and other molecules within their cytoplasm and organelles. Protein transport can occur through the cytoplasmic membrane and the outer membrane (periplasm) in bacterial cells.

[0152] Several mechanisms exist in bacteria for transmembrane protein transport. A common mechanism is the use of specialized protein transport mechanisms, such as the Sec mechanism (general secretion pathway) or the Tat mechanism (diarginine translocation). These systems utilize energy generated from ATP hydrolysis to drive protein transmembrane transport.

[0153] Other proteins can be transported across the membrane via passive diffusion, where they simply diffuse through the lipid bilayer of the membrane. The presence of channels or pores in the membrane, or the presence of transport proteins that bind to proteins and facilitate their passage across the membrane, can promote protein movement across the membrane.

[0154] In general, transmembrane protein transport is a complex and regulated process that is essential for the normal function and organization of bacterial cells. The method disclosed herein provides a valuable new tool for obtaining information about protein transport by enabling the monitoring of protein movement with greater precision and / or in a wider range of scenarios than previously possible at a reasonable cost.

[0155] To illustrate the effectiveness of this method, the inventors focused on the diarginine translocation (Tat) pathway, a protein transport system found in many types of bacteria. It is used to transport folded proteins across the cell membrane, which separates the cell's cytoplasm from the external environment. It is named after the diarginine residues found on proteins transported via this pathway. These residues act as signals, guiding proteins into the Tat pathway for transmembrane transport.

[0156] The Tat pathway is a complex process involving multiple proteins and requiring energy in the form of ATP. The protein to be transported is first recognized by the TatC protein located in the cell membrane. TatC then recruits TatA and TatB proteins, which form a complex that crosses the cell membrane. The folded protein is then transferred to the TatA protein, which acts as a "chaperone" and helps the protein cross the membrane. Once across the membrane, the protein is released and transported to its final destination inside the cell. This pathway is an important mechanism for the transport of folded proteins across the cell membrane in bacteria, and it plays a crucial role in many cellular processes.

[0157] Diffusion in the periplasm is more complex than in the cytoplasm due to its additional complexity. The periplasm contains a fixed peptidoglycan cell wall, which is porous and non-covalently attached to thousands of OmpA molecules, outer membrane proteins found in many Gram-negative bacteria. These molecules are expected to have minimal lateral diffusion within the outer membrane. The periplasm also contains a large amount (up to 5% of the total cell weight, depending on the osmotic pressure of the growth medium), which may explain the more viscous environment compared to the cytoplasm.

[0158] The inventors have demonstrated that the method disclosed herein can extract information about the complexity of periplasmic diffusion, thereby aiding in understanding molecular behavior within the compartment. This characterization is important for classifying molecules as they move across the membrane. The ability to track single molecules and classify their compartmentalization will enable the measurement of the kinetics of protein transmembrane transport.

[0159] The bacterial strains used in this embodiment are listed in Table S1, the plasmids in Table S2, and the primers in Table S3. All constructs were verified by sequencing.

[0160] The plasmid expressing the HaloTag with an N-terminal TorA signal sequence and a C-terminal SsrA tag (pQE-80 ssTorA-Halo-SsrA, periplasmic control) was constructed as follows. sstorA-gfpmut2 was amplified from plasmid pTGS (DeLisa et al., 2002) using primers P1 and P2, and cloned between the BseRI and BamHI sites in pQE-80 (Qiagen) to generate pQETG. An AgeI restriction site was inserted between sstorA and gfpmut2 in pEQTG using Q5 site-directed mutagenesis and primers P3 and P4 to generate plasmid pQET. AgeI G. The halotag gene was amplified from the pHTC HaloTag® CMV-neo Vector (Promega) using primers P5 and P6 and used to replace pQET. AgeI The AgeI-BamHI fragment containing gfpmut2 in G was used to generate plasmid pQETH. An SsrA tag was inserted into the 3' end of the halotag gene in pQETH using Q5 site-directed mutagenesis and primers P7 and P8.

[0161] To construct plasmid pQE-80 HaloTag (cytoplasmic control), Q5 site-directed mutagenesis and primers P9 and P10 were used to remove the TorA signaling sequence coding region and the AgeI restriction site from pQETH. To generate plasmid pQE-80 HaloTag-SsrA, Q5 site-directed mutagenesis and primers P7 and P8 were used to insert the SsrA tag into the 3' end of the halotag gene of pQE-80 HaloTag.

[0162] Strains expressing the HaloTag protein construct were grown from single colonies and cultured for 16 hours at 37°C with shaking at 180 rpm in M9 basal medium containing M9 salt (Davis et al., 1986), 0.1% (v / v) tryptone, 0.2% (w / v) glucose, 2 mM MgSO4, 0.1 mM CaCl2, and 0.1 mg / mL ampicillin. 100 μl of each culture was then subcultured into 5 mL of preheated M9 basal medium and grown to mid-log (OD600 = 0.5). Janelia Fluor 646 HaloTag ligand was added to 1 mL of culture to a final concentration of 2 nM. Unreacted Halotag ligand was removed from the cells by resuspending in 1 mL of imaging buffer and centrifuging three times. Constructs containing the SsrA tag were further incubated at 37°C for 30 min without shaking and then washed twice with imaging buffer. After washing, cells were concentrated in 10–30 μl of imaging buffer and then spotted onto a 1% low-fluorescence agarose (Biorad) pad containing M9 salt and 0.2% (w / v) glucose. Glass coverslips (#1.5 thickness) (Menzel-Glaser) used for imaging were pretreated by heating in an oven to 500°C to remove any fluorescent background particles.

[0163] Fluorescence images were acquired at 25°C using a Nanoimager (Oxford Nanoimaging) equipped with a 640 nm 1W DPSS laser. Optical magnification was provided by a 100× oil immersion objective (Olympus, numerical aperture (NA) 1.4), and images were acquired using an ORCA-Flash4.0 V3 CMOS camera (Hamamatsu) with a pixel size of 117 nm. The focal plane was manually positioned to the central cross-section of the cell (for high-tilt thin-layer illumination (HiLo) imaging).

[0164] Fluorescent focal points were acquired using Nanoimager, and super-resolution fluorescence localization was performed using Nanoimager software (version 1.7.3). Trajectory preprocessing and classification were performed using the same methods described above for virus classification.

[0165] Table S1 – Escherichia coli strains used

[0166]

[0167] Table S2 – Plasmids used in this study

[0168]

[0169] Table S3 – DNA primers used in this study

[0170]

[0171] A fine-grained kNN model was trained using TorA-halotag (uninduced) as a periplasmic control and TorAKK-Halotag (induced) as a cytoplasmic control. 70% of the data was used for model training and validation. The confusion matrix based on the validation set is shown below. Figure 12 As shown, the overall accuracy of the trained model is 97.3%. Then, the model is used to classify sub-trajectories belonging to the same trajectory, and the classification score for each sub-trajectory is plotted.

[0172] Figures 13-16 This is a graph displaying the relative classification scores of the detected trajectories (diffusion trajectories). In each graph, the top graph shows the relative classification score (vertical axis) as the sub-trajectory (horizontal axis) changes. The bottom graph shows the average step size (vertical axis) as the sub-trajectory (horizontal axis) changes. Figure 13 This is an example of a pure cytoplasmic trajectory, and as expected, each sub-trajectory scores close to 1. Figure 14 This is an example of a pure periodic mass trajectory, and as expected, each sub-trajectory score is close to -1. Figure 15 and 16 Examples of transition events detected in the dataset are depicted. The signal changes from pure cytoplasm to pure periplasm.

[0173] A classification score of 1 corresponds to a complete 100% confidence level for the model where the sub-trajectory is cytoplasm (see [link]). Figure 13 ), while a -1 classification score corresponds to pure periplasmic quality (see Figure 14 The idea is that, since the model is trained on pure cytoplasmic and periplasmic trajectories, when the transition occurs, the classification score should start from 1 and decrease to 0 (corresponding to random decision), then decrease to -1, at which point the protein has completely transitioned to the periplasm (see...). Figure 15 and Figure 16This method can be used to identify precise transition points, providing information that can be used to understand the mechanisms of protein transport across the cell membrane. This function is an example of the feature identification in step S2, which involves detecting changes in the characteristics of the particle's (in this case, the protein's) granular environment during the detected diffusion trajectory by analyzing trajectory fragments (sub-trajectories) of the detected diffusion trajectory. Changes in the characteristics of the granular environment include transitions between different regions of the biological cell, in this case, different regions including the bacterial periplasmic and cytoplasmic regions based on the Tat protein export pathway. However, the method is not limited to this. Transitions between different regions of other cell types can be detected, as well as transitions using other mechanisms such as the general secretion pathway (Sec).

[0174] Therefore, this method can be used to study the dynamics of the Tat system, providing information, for example, about how certain pathogens (such as Staphylococcus aureus or Streptococcus pneumoniae) use the system to cause disease. The Tat system is also significant in biotechnology because it can be used to transport proteins across the cell membrane, which may be useful for the production of recombinant proteins. By understanding the dynamics of the Tat system, researchers can also gain insights into the mechanisms of protein transport across the cell membrane and can use these insights to develop new therapies for diseases caused by disruptions in protein transport.

[0175] The above embodiments demonstrate the use of this method to distinguish different types of particles based on diffusion trajectories, particularly in the context of distinguishing different viruses, and to detect changes in the characteristics of the particle environment, particularly in the context of detecting protein transitions between different regions of a biological cell, such as across a bacterial membrane. However, the method is not limited to these applications. Based on the same principle, the identification features in step S2 of the method can include detecting changes in the characteristics of the particle during the detected diffusion trajectory by analyzing trajectory fragments of the detected diffusion trajectory. The inventors have demonstrated that the method is sensitive to minute changes in particle properties (which are necessary for distinguishing different viruses). Therefore, changes in the particle can also be detected. For example, the method can be used to detect changes in protein folding, which can be driven by changes in the composition of the environment surrounding the protein, such as changes in pH or salt concentration. Thus, the application also provides information about changes in the particle environment. The method can also be used to detect changes in the particle composition that affect its diffusion properties. For example, the method can detect compositional changes at least in part due to material attachment to the particle. Material attachment to the particle can, for example, include viral aggregation into composite particles. Thus, the monitored particle may initially include a single virus, and the method can detect when viruses aggregate with single viruses to form composite particles comprising multiple viruses. Alternatively or additionally, this method can be used to detect changes in particle composition caused at least in part by material detachment from the particle, such as a change from a composite particle containing multiple viruses to a particle composed of a single virus or fewer viruses than those present in the original composite particle.

[0176] Therefore, the methods disclosed herein provide a tool that can be used to obtain new insights into fundamental biological processes, thereby contributing to the development of new therapies and biotechnology products.

Claims

1. A computer-implemented method for detecting features of particles or a particulate environment, comprising: Receive trajectory data representing the diffusion trajectories of particles detected in one or more corresponding particulate environments, each detected diffusion trajectory comprising a sequence of trajectory segments; and The characteristics of the particles and / or one or more particle environments are identified by analyzing the trajectory segments.

2. The method according to claim 1, wherein, The one or more particulate environments include an electrolyte.

3. The method according to claim 1 or 2, wherein, The identification features include identifying the particle type of one or more of the particles.

4. The method according to claim 3, wherein, The particles include pathogens, optionally viruses or bacteria, and the particle type identification includes the type that identifies the pathogen.

5. The method according to any one of the preceding claims, wherein, The identification features include identifying one or more particle environment types of the particle environment.

6. The method according to claim 5, wherein, The identification of particle environment types includes distinguishing multiple predetermined regions within biological cells.

7. The method according to claim 6, wherein, The particles contain proteins.

8. The method according to claim 6 or 7, wherein, The predetermined area includes the region within the bacteria.

9. The method according to claim 8, wherein, The regions within the bacteria include the periplasmic region and the cytoplasmic region.

10. The method according to any one of the preceding claims, wherein, The identification features include detecting changes in the characteristics of the particle during the detected diffusion trajectory by analyzing the trajectory segments of the detected diffusion trajectory.

11. The method according to claim 10, wherein, The changes in the characteristics of the particles include alterations in protein folding.

12. The method according to claim 10 or 11, wherein, The changes in the characteristics of the particles include changes in the composition of the particles.

13. The method according to claim 12, wherein, The change in composition is at least partly due to the attachment of material to the particles, optionally wherein the particles include viruses and the attachment of material to the particles includes the aggregation of viruses into composite particles.

14. The method according to claim 12 or 13, wherein, The change in composition is at least partly due to the material detaching from the particles.

15. The method according to any one of the preceding claims, wherein, The identification features include detecting changes in the characteristics of the particle's particle environment during the detected diffusion trajectory of the particle by analyzing the trajectory segments of the detected diffusion trajectory.

16. The method according to claim 15, wherein, The alteration of the characteristics of the particulate environment includes transitions between different regions of a biological cell, wherein, optionally: The particles comprise proteins; The biological cells include bacteria; and / or The different regions include the periplasmic region and the cytoplasmic region of bacteria.

17. The method according to claim 16, wherein, The aforementioned transitions between different regions occur via the diarginine translocation Tat protein export pathway or the general secretory pathway Sec.

18. The method according to any one of the preceding claims, wherein, The identification features are performed by inputting the trajectory data into a trained machine learning model.

19. The method of claim 18, wherein: The one or more particulate environments include an electrolyte; and The electrolyte conditions for the trajectory data and the electrolyte conditions during training the machine learning model are substantially the same in terms of at least the following parameters: pH; temperature; chemical composition.

20. The method according to claim 18 or 19, wherein, The trained machine learning model was trained using the k-nearest neighbor machine learning algorithm.

21. The method according to any one of claims 18-20, comprising: The trajectory data is preprocessed using a dimensionality reduction algorithm, optionally including principal component analysis, before being input into the trained machine learning model.

22. The method according to any one of the preceding claims, comprising: The video data is processed to provide the trajectory data.

23. The method according to any one of the preceding claims, comprising: Video data of the particles in the particle environment or multiple particle environments is captured and processed to provide the trajectory data.

24. The method according to claim 22 or 23, wherein, The processing of the video data includes tracking the position of individual particles on a frame sequence of the video data to determine the corresponding detected diffusion trajectories.

25. The method according to claim 24, wherein, The detected diffusion trajectories are selectively included in the trajectory data such that only a subset of the detected particles contributes to the trajectory data.

26. The method of claim 25, wherein, The selection is based on the determined diffusion rate of the detected particles.

27. The method according to claim 26, wherein, The selection includes excluding detected diffusion trajectories of particles based on the following criteria: The position change of the particle between temporally adjacent frames or between frames separated by a predetermined number of frames exceeds a predetermined maximum threshold; and / or The positional change of the particle between temporally adjacent frames or between frames separated by a predetermined number of frames is below a predetermined minimum threshold.

28. The method according to any one of claims 25-27, wherein, The selection includes excluding detected diffusion trajectories of particles based on the following determination: in one or more frames of the video data, the distance between the particle and the detected diffusion trajectories of different particles is less than a predetermined minimum distance, optionally wherein the predetermined minimum distance is approximately equal to the predetermined maximum threshold.

29. The method according to any one of claims 25-28, wherein, The video data was captured using a wide field-of-view microscope, optionally configured to operate in HILO mode with a highly tilted layered light sheet.

30. The method according to any one of the preceding claims, wherein, One or more detectable markers are bound to each particle.

31. The method according to claim 30, wherein, Each detectable tag is configured to be detected by capturing electromagnetic radiation emitted from the detectable tag, which optionally includes a fluorescent tag.

32. The method according to claim 30 or 31, wherein, The particles include enveloped viral particles, and the detectable marker is configured to bind to the envelope of the enveloped viral particles.

33. A method for training a machine learning model, comprising: The system receives training data for markers, the training data including data representing multiple detected diffusion trajectories of particles and data representing markers associated with each detected diffusion trajectory, each detected diffusion trajectory including a sequence of trajectory segments, the markers indicating characteristics of the particle corresponding to the diffusion trajectory and / or characteristics of the particle environment; and The machine learning model is trained using supervised learning to identify features of particles and / or features of the particle's environment based on the trajectory fragments.

34. A system for detecting characteristics of particles or a particulate environment, the system comprising a data processing device configured to perform the method according to any one of claims 1-32.

35. The system according to claim 34, further comprising: A sample receiving device configured to receive a sample to be tested; as well as An image capturing device configured to capture electromagnetic radiation emitted from the sample, and a system configured to derive trajectory data from the captured electromagnetic radiation.

36. A computer program or computer-readable medium comprising instructions that, when executed by a computer, cause the computer to perform the method according to any one of claims 1-32.