Combining microscopy with further analysis for adaptation
Through the method of combining microscopy and neural network, the sample preparation and analysis steps are adapted to the inefficiency problem in the prior art and achieve more efficient and accurate sample processing and analysis.
Patent Information
- Application Number
- CN202380085147.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2022-12-15
- Filing Date
- 2023-12-12
- Publication Date
- 2025-07-08
AI Technical Summary
The existing methods of microscopy combined with other sample analysis techniques are inefficient and poorly performed, resulting in inaccurate positioning and identification of areas of interest and causing experimental losses.
Sample images are captured by microscope, targets or regions of interest are identified using neural networks, and sample preparation, image capture and neural networks are adapted based on the analysis results, combined with a variety of analysis technologies such as laser microscissification, fluorescence activated cell sorting, etc., to generate training data to improve sample processing and analysis.
It improves the efficiency and accuracy of sample processing and analysis, and enables more precise identification and extraction of specific targets or areas of interest, reducing experimental errors.
Smart Images

Figure CN120283189A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to microscopy. In particular, the present invention relates to the combination of microscopy with other sample analysis techniques. Background Art
[0002] Life science research involves the scientific study of living organisms and their life processes. This includes studying the structure of cells and tissues, and how these structures affect the function of these cells and tissues. Sequencing methods provide information about cell types and their heterogeneity in complex tissues.
[0003] Spatial transcriptomics can elucidate single-cell heterogeneity and define cell types while retaining spatial information. The preservation of this spatial context is crucial for understanding key aspects in the fields of cell biology, developmental biology, neurobiology, tumor biology, etc., because specific cell types and their specific structures are closely related to biological activities, and specific cell types and their specific structures at the scale of whole tissues and organisms have not been fully explored. Spatial transcriptomics techniques can be used to generate complete maps of large complex tissues (such as the human brain).
[0004] Spatial transcriptomics studies are generally carried out in two ways. One is to read the transcriptome via in situ sequencing or multiplex fluorescence in situ hybridization (FISH) through microscopy. The other way is to capture RNA in a way that retains spatial information while using non-in situ RNA sequencing for sequencing. These methods are generally complementary and differ in terms of target coverage, spatial resolution, and throughput.
[0005] The state-of-the-art technologies currently include the combination of microscopy for phenotypic assessment, such as via laser microdissection, with technologies such as targeted molecular biology. These technologies are expensive and time-consuming, and inaccurate positioning or identification of the region of interest in laser cutting experiments can result in serious losses.
[0006] Therefore, there is a need to combine other technologies to address the problems of low efficiency and poor performance in methods for analyzing samples using microscopy. Summary of the Invention
[0007] To address the challenges in the prior art, new methods and systems are disclosed that are adapted to adjust at least one step in a series of steps for processing and analyzing samples.
[0008] In an embodiment, the method includes: capturing at least one image of a first prepared sample through a microscope; identifying at least one target or region of interest in the at least one image using a neural network; analyzing at least a portion of the first prepared sample that corresponds to the at least one identified target or region of interest; and adapting (110) at least one of the following based on a comparison between the analysis result of at least a portion of the first prepared sample and data obtained from the at least one image of the first prepared sample: the preparation of a second sample, the capturing of at least one image of a second prepared sample, and the neural network.
[0009] Adapting the preparation of a second sample, the capturing of at least one image of a second prepared sample, and the neural network based on a comparison between the analysis result of at least a portion of a sample and data obtained from the at least one image of the first prepared sample provides an improved method for enhancing the efficiency and accuracy of processing and analyzing samples.
[0010] According to some aspects, analyzing at least a portion of a sample includes extracting that portion of the sample, wherein analyzing at least a portion of the sample includes analyzing the extracted portion of the sample. Analyzing at least a portion of the sample may include using at least one of the following: laser microdissection (LMD), laser capture microdissection (LCM), tissue lysis and subsequent fluorescence-activated cell sorting (FACS), physical local tissue extraction (e.g., by scalpel or needle), extraction of whole cells or cytoplasm with a micropipette / membrane patch pipette (electrophysiology), protein extraction methods, tissue lysis and fluorescence-activated cell sorting (FACS sorting), immunomagnetic bead cell sorting, spatial transcriptomics, spatial multi-omics, indirect methods, next-generation sequencing (NGS) methods, ribonucleic acid sequencing (RNAseq) methods, microarrays, real-time polymerase chain reaction (qPCR) methods, blotting methods, and mass spectrometry (MS).
[0011] In one aspect, the microscope includes one of an epifluorescence microscope and a confocal microscope, and at least one of an electron microscope, a scanning electron microscope (SEM), a focused ion beam scanning electron microscope (FIB-SEM), a transmission electron microscope (TEM), a correlative light electron microscope (CLEM), and a cryogenic correlative light electron microscope (CryoCLEM) is used to analyze at least a portion of the first prepared sample.
[0012] According to some aspects, at least one of the first sample and the second sample is prepared to identify at least one component in the corresponding sample using at least one of a stain, an ion label, an antibody dilution, a fluorescence in situ hybridization (FISH) probe, a fluorophore, and a detergent treatment. The component of the corresponding sample may be one of a single cell nucleus, a cell compartment, an intracellular compartment, a pathogen, a virus, a bacterium, a fungus, a single cell, and a cell cluster.
[0013] According to some aspects, the method further includes: obtaining data associated with at least one of the preparation of the first sample, the capture of at least one image of the first prepared sample, at least one image of the first prepared sample, and the analysis result of at least a portion of the sample; and storing the data as a single data point in a repository. The repository may include multiple data points corresponding to at least one of different parameters for preparing the sample, different parameters for capturing at least one image of the prepared sample, different targets or regions of interest, different samples, and different types of samples. The preparation of the second sample, the capture of at least one image of the second prepared sample, and the adaptation of at least one of the neural networks may be based on the analysis of the multiple data points.
[0014] By generating the repository and utilizing the data in the repository to perform the claimed adaptation of at least one step in the series of steps for processing and analyzing the sample, the measurement accuracy can be improved. For example, the data in the repository can be analyzed to obtain information on how to improve the processing and analysis of the sample for a specific target.
[0015] Additionally or alternatively, the method may further include generating training data for at least one neural network based on the multiple data points for the step of identifying at least one target or region of interest in at least one image. The method may further include fine-tuning the neural network based on the training data for the step of identifying at least one target or region of interest in at least one image. The method may further include training a set of neural networks based on the training data for the step of identifying at least one target or region of interest in at least one image.
[0016] This enables the improvement of the neural network to accurately identify at least one target or region of interest in at least one image captured by the microscope. Additionally, new neural networks can be developed to accurately detect and identify specific structures or components in the sample.
[0017] According to some aspects, the comparison between the analysis result of at least a portion of the first prepared sample and the data obtained from at least one image of the first prepared sample includes: correlating the signal intensity obtained from at least one image of the first prepared sample and related to one of the stain, ion label, antibody dilution, fluorescence in situ hybridization (FISH) probe, fluorophore, and detergent treatment applied to the first sample with the molecular content of the analyzed portion of the first prepared sample obtained according to the analysis result of at least a portion of the first prepared sample. The comparison between the analysis result of at least a portion of the first prepared sample and the data obtained from at least one image of the first prepared sample may include: combining phenotypic assessment and genotypic assessment.
[0018] In an embodiment, a system for generating neural network training data is provided. The system includes: a microscope configured to capture at least one image of a sample; at least one processor configured to identify at least one target or a region of interest associated with the at least one target in the at least one image; and an extractor configured to extract a portion of the sample corresponding to the identified at least one target or region of interest. The system is configured to analyze the extracted portion of the sample, and the system is configured to generate training data for training a neural network based on the at least one image of the object and the analysis result of the extracted portion of the sample. The training data may associate phenotypic information obtainable from one or more images with at least one of genotypic information, proteotypic information, and transcriptomic profile information obtainable from the analysis of the extracted portion of the sample. The region of interest may correspond to the identified target. The analysis result of the extracted portion may include information about the molecular content of the extracted portion. The target in the at least one target may be one of a single cell nucleus, a cell compartment, an intracellular compartment, a pathogen, a virus, a bacterium, a fungus, a single cell, and a cell cluster. The sample may include tissue.
[0019] By generating the claimed training data, a neural network can be trained to achieve the technical objective of accurately identifying targets in images, where the labels of the training data can be obtained by analyzing only a portion of the sample, thereby accurately and reliably identifying the components of the sample, such as specific cells or single cell nuclei. For the analysis of this portion of the sample, non-image-based methods can be used. For example, this portion of the sample can be analyzed to obtain information about the molecular content of the extracted portion, such as by using mass spectrometry. This allows the use of labels that cannot be obtained from the sample images to label the training data and can accelerate measurements or experiments.
[0020] The microscope may include at least one of a fluorescence microscope, a confocal microscope, a wide-field microscope, a light-sheet microscope, a super-resolution microscope, an X-ray microscope, an electron microscope, a scanning probe microscope, and a transmission electron microscope.
[0021] The extractor may use at least one of laser microdissection LMD, fluorescence-activated cell sorting FACS, a scalpel, a needle, a pipette, and a focused ion beam scanning electron microscope FIB-SEM to extract a portion of the sample corresponding to the identified at least one target or region of interest.
[0022] The system may use at least one of methods such as quantitative polymerase chain reaction (qPCR) based on polymerase chain reaction (PCR), microarray, mass spectrometry MS, and next-generation sequencing NGS to analyze the extracted portion of the sample.
[0023] In another embodiment, a computer-readable storage medium storing computer-executable instructions is provided. When executed by one or more processors, the computer-executable instructions perform at least a portion of the adaptation method of at least one step of the above-described series of steps for processing and analyzing a sample.
[0024] The following detailed description and the accompanying drawings provide a more detailed understanding of the nature and advantages of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS
[0025] The drawings are incorporated into and form a part of this specification to explain the principles of the embodiments. The drawings should not be construed as limiting the embodiments to the illustrated and described embodiments of how to make and use them. More features and advantages will become apparent from the following, particularly from the description of the embodiments shown in the drawings, in which: Figure 1 A process flow diagram of an adaptation method of at least one step of a series of steps for processing and analyzing a sample according to at least one embodiment is shown, Figure 2 An example architecture of a system according to an embodiment is shown, and Figure 3 A process flow diagram of a method for generating training data for a neural network according to an embodiment is shown. DETAILED DESCRIPTION
[0026] Systems and methods for adapting at least one step of a series of steps for processing and analyzing a sample, and systems and methods for generating training data for a neural network are described herein. For ease of explanation, numerous examples and specific details are set forth herein to provide a thorough understanding of the described embodiments. The embodiments defined by the claims may include some or all of the features of these examples, either individually or in combination with other features hereinafter, and may also include modifications and equivalents of the features and concepts herein. Exemplary embodiments will be described with reference to the accompanying drawings, in which elements and structures are indicated by reference numerals. Additionally, if the embodiment is a method, the steps and elements of the method may be performed in parallel or in sequential combination. As long as they are not mutually contradictory, all the embodiments described hereinafter may be combined with each other.
[0027] As used herein, the term "and / or" includes any and all combinations of one or more of the associated listed items and may be abbreviated as " / ".
[0028] Although certain aspects are described in the context of an apparatus, it is apparent that these aspects also represent a description of the corresponding method, where a block or device corresponds to a method step or a feature of a method step. Similarly, aspects described in the context of a method step also represent a description of the corresponding block or item or feature of the corresponding apparatus.
[0029] Some embodiments relate to microscopes that include systems related to one or more of Figures 1 to 3 Or, the microscope can be part of or connected to a system related to one or more of Figures 1 to 3 FIG. 6 shows a schematic diagram of a system 200 configured to perform the methods herein. System 200 includes a microscope 210 and a computer system 220. Microscope 210 is configured to capture images and is connected to computer system 220. Computer system 220 is configured to perform at least a portion of the methods herein. Computer system 220 can be configured to perform machine learning algorithms. Computer system 220 and microscope 210 can be separate entities, but can also be integrated within a common housing. Computer system 220 can be part of the central processing system of microscope 210, and / or computer system 220 can be part of a sub-component of microscope 210, such as a sensor, actuator, camera, or illumination unit of microscope 210, etc. Figure 2
[0030] The computer system 220 can be a local computer device (such as a personal computer, laptop, tablet, or mobile phone) having one or more processors and one or more storage devices, or can be a distributed computer system (such as a cloud computing system, having one or more processors and one or more storage devices, distributed at different locations, e.g., at a local client and / or one or more remote server farms and / or data centers). The computer system 220 can include any circuit or combination of circuits. In one embodiment, the computer system 220 can include one or more processors, which can be of any type. As used herein, a processor can represent any type of computing circuit, such as for example a microscope or microscope component (e.g., a camera) or any other type of processor or processing circuit, e.g., but not limited to, a microprocessor, a microcontroller, a complex instruction set computing (CISC) microprocessor, a reduced instruction set computing (RISC) microprocessor, a very long instruction word (VLIW) microprocessor, a graphics processor, a digital signal processor (DSP), a multi-core processor, a field programmable gate array (FPGA). Other types of circuits that can be included in the computer system X20 can be custom circuits, application specific integrated circuits (ASICs), etc., such as one or more circuits (e.g., communication circuits) for wireless devices (such as mobile phones, tablets, laptops, two-way radios, and similar electronic systems). The computer system 220 can include one or more storage devices, and the storage devices can include one or more storage elements suitable for a particular application, such as main memory in the form of random access memory (RAM), one or more hard disk drives, and / or one or more drives for handling removable media (such as compact discs (CDs), flash cards, digital video discs (DVDs), etc.). The computer system 220 can also include a display device, one or more speakers, and a keyboard and / or controller, which can include a mouse, trackball, touch screen, voice recognition device, or any other device that allows a system user to input information to and receive information from the computer system 220.
[0031] Some or all of the method steps can be performed by (or using) a hardware device, such as a processor, microprocessor, programmable computer, or electronic circuit. In some embodiments, one or more of the most important method steps can be performed by such a device.
[0032] According to certain implementation requirements, embodiments of the present invention can be implemented in hardware or software form. It can be implemented using a non-transitory storage medium (such as a digital storage medium, such as a floppy disk, DVD, Blu-ray, CD, ROM, PROM, EPROM, EEPROM or FLASH memory), in which electronically readable control signals are stored, and these control signals cooperate with (or are capable of cooperating with) a programmable computer system to perform corresponding methods. Therefore, the digital storage medium can be computer-readable.
[0033] Some embodiments according to the present invention include a data carrier having electronically readable control signals, and the data carrier is capable of cooperating with a programmable computer system to thereby execute one of the methods herein.
[0034] Generally, embodiments of the present invention can be implemented as a computer program product having program code, and when the computer program product runs on a computer, the program code is used to execute one of the methods. The program code can be stored, for example, on a machine-readable carrier.
[0035] Other embodiments include a computer program for executing one of the methods described herein, and the computer program is stored on a machine-readable carrier.
[0036] In other words, therefore, one embodiment of the present invention is a computer program having program code, and the program code is used to execute one of the methods described herein when the computer program runs on a computer.
[0037] Therefore, another embodiment of the present invention is a storage medium (or data carrier or computer-readable medium) on which a computer program for executing one of the methods described herein when executed by a processor is stored. The data carrier, digital storage medium or recording medium is generally tangible and / or non-transitory. Another embodiment of the present invention is the device described herein, including a processor and a storage medium.
[0038] Therefore, another embodiment of the present invention is a data stream or signal sequence representing a computer program for executing one of the methods described herein. For example, the data stream or signal sequence can be configured to be transmitted through a data communication connection (such as through the Internet).
[0039] Further embodiments include a processing device, such as a computer or a programmable logic device, which is configured or adapted to execute one of the methods described herein.
[0040] Further embodiments include a computer on which a computer program for executing one of the methods herein is installed.
[0041] Another embodiment of the present invention includes an apparatus or system configured to transmit (e.g., electronically or optically) a computer program for performing one of the methods described herein to a receiver. The receiver can be, for example, a computer, a mobile device, a storage device, etc. The apparatus or system can include, for example, a file server for transmitting the computer program to the receiver.
[0042] In some embodiments, a programmable logic device (e.g., a field programmable gate array) can be used to perform some or all of the functions of the methods described herein. In some embodiments, a field programmable gate array can cooperate with a microprocessor to perform one of the methods described herein. Generally, these methods are preferably performed by any hardware device.
[0043] Embodiments can be based on the use of machine learning models or machine learning algorithms, such as neural networks. Machine learning can refer to algorithms and statistical models that a computer system can use, which do not require the use of explicit instructions, but rely on models and inferences to perform specific tasks. For example, in machine learning, data transformations inferred from the analysis of historical and / or training data can be used instead of rule-based data transformations. For example, a machine learning model or a machine learning algorithm can be used to analyze the content of an image. To enable a machine learning model to analyze the content of an image, training images can be used as input and training content information can be used as output to train the machine learning model. By training the machine learning model using a large number of training images and / or training sequences (e.g., words or sentences) and associated training content information (e.g., labels or annotations), the machine learning model "learns" to recognize the content of the image, and thus can be used to recognize the content of an image not included in the training data. The same principle also applies to other types of sensor data: by training the machine learning model using training sensor data and expected outputs, the machine learning model "learns" the transformation between the sensor data and the outputs, and can be used to provide outputs based on non-training sensor data provided to the machine learning model. The data provided (e.g., sensor data, metadata, and / or image data) can be preprocessed to obtain a feature vector, which is used as input to the machine learning model.
[0044] A machine learning model can be trained using training input data. The above example uses a training method called "supervised learning". In supervised learning, the machine learning model is trained using multiple training samples, where each sample can include multiple input data values and multiple expected output values, i.e., each training sample is associated with an expected output value. By specifying the training samples and the expected output values, the machine learning model "learns" which output value to provide based on input samples similar to those provided during training. In addition to supervised learning, semi-supervised learning can also be used. In semi-supervised learning, some of the training samples lack corresponding expected output values. Supervised learning can be based on supervised learning algorithms (e.g., classification algorithms, regression algorithms, or similarity learning algorithms). When the output is restricted to a finite set of values (categorical variables), a classification algorithm can be used, i.e., classifying the input into one of the finite set of values. When the output can have any numerical value (within a certain range), a regression algorithm can be used. Similarity learning algorithms can be similar to classification and regression algorithms, but are based on learning from examples using a similarity function that measures the similarity or relatedness of two objects. In addition to supervised or semi-supervised learning, unsupervised learning can also be used to train a machine learning model. In unsupervised learning, (only) input data can be provided, and unsupervised learning algorithms can be used to find structures in the input data (e.g., by grouping or clustering the input data to find commonalities in the data). Clustering is the assignment of input data including multiple input values to subsets (clusters) such that the input values within the same cluster are similar according to one or more (predefined) similarity criteria, but dissimilar to the input values included in other clusters.
[0045] Reinforcement learning is a third group of machine learning algorithms. In other words, reinforcement learning can be used to train a machine learning model. In reinforcement learning, one or more software actors (called "software agents") are trained to take actions in an environment. Based on the actions taken, a reward is calculated. Reinforcement learning is based on training one or more software agents to select actions to increase the cumulative reward, so that the software agents perform better on a given task (an increase in reward is evidence).
[0046] In addition, some techniques can be applied to some machine learning algorithms. For example, feature learning can be used. In other words, a machine learning model can be trained at least in part using feature learning, and / or a machine learning algorithm can include a feature learning component. Feature learning algorithms (which can be called representation learning algorithms) can retain the information in their input, but can also transform it in a useful way, typically as a preprocessing step before performing classification or prediction. For example, feature learning can be based on principal component analysis or cluster analysis.
[0047] In some examples, anomaly detection (i.e., outlier detection) can be used, which aims to identify input values that are significantly different from most of the input or training data and thus raise suspicion. In other words, a machine learning model can be trained at least in part using anomaly detection, and / or a machine learning algorithm can include an anomaly detection component.
[0048] In some examples, a machine learning algorithm can use a decision tree as a predictive model. In other words, a machine learning model can be based on a decision tree. In a decision tree, an observation about an item (e.g., a set of input values) can be represented by a branch of the decision tree, while the output value corresponding to that item can be represented by a leaf of the decision tree. A decision tree can support both discrete and continuous values as output values. If discrete values are used, the decision tree can be represented as a classification tree, and if continuous values are used, the decision tree can be represented as a regression tree.
[0049] Association rules are another technique that can be used in a machine learning algorithm. In other words, a machine learning model can be based on one or more association rules. Association rules are created by identifying relationships between variables in a large amount of data. A machine learning algorithm can identify and / or utilize one or more relationship rules, which represent knowledge derived from the data. These rules can be used to store, manipulate, or apply knowledge.
[0050] A machine learning algorithm is generally based on a machine learning model. In other words, the term "machine learning algorithm" can represent a set of instructions that can be used to create, train, or use a machine learning model. The term "machine learning model" can represent a data structure and / or set of rules that represents the learned knowledge (e.g., based on training performed by a machine learning algorithm). In an embodiment, the use of a machine learning algorithm can imply the use of an underlying machine learning model (or underlying machine learning models). The use of a machine learning model can imply that the machine learning model and / or the data structure / set of rules that is the machine learning model is trained by a machine learning algorithm.
[0051] For example, the machine learning model can be an artificial neural network (ANN). An ANN is a system inspired by biological neural networks, such as those in the retina or the brain. An ANN includes multiple interconnected nodes and multiple connections between the nodes, i.e., the so-called edges. There are generally three types of nodes: input nodes that receive input values, hidden nodes that are (only) connected to other nodes, and output nodes that provide output values. Each node can represent an artificial neuron. Each edge can transmit information from one node to another. The output of a node can be defined as a (non-linear) function of its inputs (e.g., the sum of its inputs). The inputs of a node can be used in the function based on the "weights" of the edges or the nodes providing the inputs. The weights of the nodes and / or edges can be adjusted during the learning process. In other words, the training of an artificial neural network can include adjusting the weights of the nodes and / or edges of the artificial neural network, i.e., achieving the desired output for a given input.
[0052] Alternatively, the machine learning model can be a support vector machine, a random forest model, or a gradient boosting model. A support vector machine (i.e., a support vector network) is a supervised learning model with an associated learning algorithm that can be used to analyze data (e.g., in classification or regression analysis). A support vector machine can be trained by providing an input with multiple training input values belonging to one of two classes. A support vector machine can be trained to assign new input values to one of the two classes. Or, the machine learning model can be a Bayesian network, which is a probabilistic directed acyclic graph model. A Bayesian network can represent a set of random variables and their conditional dependencies using a directed acyclic graph. Or, the machine learning model can be based on a genetic algorithm, which is a search algorithm and heuristic technique that mimics the process of natural selection.
[0053] Spatial biology is defined as the study of tissues in a two-dimensional or three-dimensional environment. For example, the spatial structure of cells can be mapped, and the interactions between cells and their surrounding environment can be determined. Spatial biology enables single-cell analysis. Through spatial biology tools, information that cannot be obtained by sequencing can be obtained. Spatial biology can be used to study oncology, immuno-oncology, neurobiology, and even COVID research.
[0054] By using immunofluorescence and next-generation sequencing technologies together, users can determine how transcriptional dynamics change in a spatial context. Spatial information can be obtained at various scales, including the tissue, single-cell, and subcellular levels.
[0055] Figure 1 is a processing flowchart of an exemplary adaptation method 100 for at least one step in a series of steps for processing and analyzing a sample according to an embodiment.
[0056] In step 102, method 100 includes preparing a first sample. The first sample can be prepared to identify at least one component in the first sample using at least one of a stain, an ion label, an antibody diluent, a fluorescence in situ hybridization (FISH) probe, a fluorophore, and a detergent treatment. For example, staining techniques such as immunogold labeling or immunogold staining (IGS) can be used to prepare the first sample.
[0057] In step 104, at least one image of the first prepared sample is captured by a microscope. The microscope can include at least one of a fluorescence microscope, a confocal microscope, a wide-field microscope, a light-sheet microscope, a super-resolution microscope, an X-ray microscope, an electron microscope, a scanning probe microscope, and a transmission electron microscope.
[0058] In step 106, a neural network is used to identify at least one target or region of interest in the at least one image. The neural network can be a convolutional neural network (CNN), such as a deep CNN. The target can correspond to the region of interest. The neural network can be a machine learning algorithm. The neural network can be trained to identify the target or the regions of interest corresponding to the respective targets. The target can be one or more of a single cell, a single cell nucleus, a detection event, etc. The neural network can be trained based on an image-based dataset.
[0059] In step 108, at least a portion of the first prepared sample is analyzed. This portion of the first prepared sample can correspond to the at least one target or region of interest identified. Steps 104 and 106 can be regarded as a first analysis method, while step 108 can correspond to a second analysis method different from the first analysis method. Analyzing at least a portion of the sample can include using at least one of the following: laser microdissection (LMD), laser capture microdissection (LCM), tissue lysis and subsequent fluorescence-activated cell sorting (FACS), physical local tissue extraction (e.g., by means of a scalpel or a needle), whole cell or cytoplasm extraction with a micropipette / patch pipette (electrophysiology), protein extraction methods, tissue lysis and fluorescence-activated cell sorting (FACS sorting), immunomagnetic bead cell sorting, spatial transcriptomics, spatial multi-omics, indirect methods, next-generation sequencing (NGS) methods, ribonucleic acid sequencing (RNAseq) methods, microarrays, quantitative polymerase chain reaction (qPCR) methods, blotting methods, and mass spectrometry (MS).
[0060] In an example, the microscope includes one of an epi-fluorescence microscope and a confocal microscope, and at least one of an electron microscope, a scanning electron microscope SEM, a focused ion beam scanning electron microscope FIB-SEM, a transmission electron microscope TEM, a correlative light electron microscope CLEM, and a cryogenic correlative light electron microscope CryoCLEM is used to analyze at least a portion of the first prepared sample.
[0061] Step 108 may include extracting a portion of the first prepared sample from the first prepared sample. For example, a portion of the first prepared sample can be extracted by performing laser microdissection. Alternatively, a portion of the first prepared sample can be extracted by using FIB-SEM to make a thin section. Analyzing at least a portion of the first prepared sample includes analyzing the extracted portion of the first prepared sample. For example, the extracted portion of the first prepared sample can be analyzed by mass spectrometry. Alternatively, at least one suitable downstream method (such as one or more of NGS, MS, microarray, etc.) can be used to analyze the extracted portion of the sample to generate biomolecular data in a spatial context, or multiple downstream techniques can be combined (for example, subsequent TEM imaging, and one or more of NGS, MS, and microarray) for analysis.
[0062] In step 110, at least one step of method 100 is adapted for processing and analyzing a second sample. This adaptation is based on a comparison between the analysis result of at least a portion of the first prepared sample in step 108 and the data obtained from at least one image of the first prepared sample. This adaptation includes adapting at least one of step 102, step 104, step 106, and step 108. For example, at least one of the following can be adapted: the preparation of the second sample, the capture of at least one image of the second prepared sample, the neural network, and the analysis method for analyzing at least a portion of the second sample. Adapting the neural network can include: fine-tuning the neural network based on training data for the step of identifying at least one target or region of interest in at least one image; alternatively, replacing the neural network with a different neural network for the step of identifying at least one target or region of interest in at least one image. Different neural networks can be used to identify different targets.
[0063] The comparison between the analysis result of at least a portion of the first prepared sample and the data obtained from the at least one image of the first prepared sample can include: correlating the signal intensity related to one of the stain, ion label, antibody diluent, fluorescence in situ hybridization FISH probe, fluorophore, and detergent treatment applied to the first sample and obtained from at least one image of the first prepared sample, with the molecular content of the analyzed portion of the first prepared sample obtained according to the analysis result of the at least a portion of the first prepared sample.
[0064] The comparison between the analysis result of at least a portion of the first prepared sample and the data obtained from the at least one image of the first prepared sample includes: combining phenotypic evaluation and genotypic evaluation.
[0065] In an example, data obtained from a captured image or an identified target or region of interest is compared and / or correlated with data obtained from the analysis in step 108. The comparison can include comparing the signal of a labeled protein or nucleic acid (e.g., RNA) with the relative expression level from a downstream method used to analyze the sample or a portion of the sample. This comparison can balance the number of labels (e.g., antibodies or FISH probes) and select different fluorophores for the next experiment. Additionally, new potential markers, such as biomarkers, can be discovered and labeled for upcoming experiments. Furthermore, the type of labeled fluorophore can be adjusted: for example, if a certain marker is relatively highly expressed, a less bright label or a label that does not fully match the imaging device properties of the next sample can be used; while for a target with lower expression, a brighter and more easily detectable fluorophore can be used for labeling for the next sample. In cases where the signal obtained from at least one image does not match the analysis result (e.g., the determined expression level), the method of processing and analyzing the sample can be improved to achieve a match between the fluorescence signal obtained from at least one image and the true molecular expression obtained from the analysis. This can also improve experiments without downstream processing (purely imaging-based), thereby achieving a reliable correlation result between the signal intensity and the true molecular expression.
[0066] By incorporating all available information in the sample into a reliable feedback loop, method 100 can be improved by adapting the sample preparation, imaging, and region of interest localization of subsequent experiments. This can be achieved, for example, by selecting an intelligent combination of expression levels and smart tags based on the binding characteristics of the markers and the way of detecting the labels using the corresponding available systems.
[0067] In step 120, data related to at least one of the preparation of the first sample, the capture of at least one image of the first prepared sample, the at least one image of the first prepared sample, and the analysis result of at least a portion of the sample can be obtained. In step 122, the data can be stored as a single data point in a repository. The repository can include multiple data points corresponding to at least one of the following: different parameters for preparing the sample, different parameters for capturing at least one image of the prepared sample, different targets or regions of interest, different samples, and different types of samples. For example, all sample preparation steps can be tracked to obtain information about one or more of embedding, fixation, stain, stain concentration, antibody dilution, FISH probe, fluorophore, detergent treatment, etc. of the sample. The data associated with capturing at least one image of the first prepared sample can include information about the microscope, such as one or more of the available filter sets, illumination sources, laser wavelengths, detector efficiencies, etc.
[0068] The data in the repository can be used for or analyzed to find patterns to modify step 102, modify sample preparation; modify step 104, modify the image acquisition method and / or parameters (such as exposure time), and expand the data in the repository.
[0069] The data in the repository can be used to improve the training of the neural network or specify the neural network for more specific target identification.
[0070] For example, data from multiple similar experiments in the repository can be used for further research on specific pathways, such as for the proteins and / or nucleic acids involved and their modifications, such as to determine the intracellular spatial environment and pathway validation, such as by high-resolution imaging, to verify known pathways and potential new pathways or the interactions of the molecules involved using STED or EM.
[0071] As indicated by the arrow from step 122 to step 110, the preparation of the second sample, the capture of at least one image of the second prepared sample, and the adaptation of at least one of the neural networks can also additionally be based on the analysis of multiple data points. For example, combining the image data and the results of subsequent measurements can be used to improve the targeting in laser microdissection for subsequent experiments. In-depth analysis related to antibody staining and the true (e.g., verified by MS (or CryoTEM)) protein content of single cells can be obtained from the data in the repository. This can also overall improve antibody labeling and FISH (fluorescence in situ hybridization) or related techniques or combined techniques to obtain more reliable quantitative results. The data in the repository can be used to precisely improve the processing itself.
[0072] Using the repository (such as data alignment results), the information generated downstream of LMD (such as genotype, "proteotype", etc.) can be correlated with the spatial image information to identify special cell clusters or single cells for the next round or next experiment. For example, by analyzing the data in the repository or combining the analysis results in step 108 with at least one image obtained in step 104, it can be determined whether the analysis results (such as showing the activity or absence of signal cascade proteins in certain single cells or cell clusters) match the spatial attributes (such as shape parameters, distance from other cells or cell clusters, and / or changes in the appearance of the cell or the fluorescence pattern of FISH immunofluorescence) in the spatial imaging environment. These associations between imaging and the spatial environment generated downstream of LMD can be used for more specific AI training options and further improve sample preparation and target labeling.
[0073] In step 124, training data for at least one neural network can be generated based on multiple data points in the repository for the step of identifying at least one target or region of interest in at least one image.
[0074] In step 128, the training data can be used to fine-tune a neural network based on the training data for the step of identifying at least one target or region of interest in at least one image. Fine-tuning involves using the weights of the trained network as the starting values for training a new network. Fine-tuning can include freezing one or more layers and / or one or more weights such that not all weights can be adjusted. Alternatively or additionally, the learning rate can be decreased during fine-tuning.
[0075] In step 126, the training data can be used to train a set of neural networks based on the training data for the step of identifying at least one target or region of interest in at least one image. The set of neural networks can include one or more neural networks.
[0076] In an example, in a first step, a tissue is imaged using a high-resolution microscope. In a second step, machine learning is applied to find regions of interest in one or more images of this step for extraction. The part of the sample corresponding to the region of interest is processed using laser microdissection technology. Other methods can also be used to cut the sample part. The sample part can be analyzed using ultra-high sensitivity mass spectrometry. The analysis results can be further analyzed using bioinformatics and provided to researchers and clinicians as an improved data resource. The bioinformatics data resource can be used to train a neural network to improve the specificity of the algorithm, so as to more specifically identify the target regions of interest, for example, by identifying through a suitable combination of biomarkers.
[0077] The embodiments combine microscopy (phenotypic assessment) with underlying molecular biology patterns (genotypic assessment). Examples of such combinations include using fluorescence (epifluorescence) microscopy, confocal microscopy, light sheet microscopy, and / or wide-field microscopy (e.g., slide scanning technology, bright-field microscopy, and laser microdissection) for assessment to extract specific homogeneous microscopic target regions from heterogeneous tissues, such as single cell nuclei, single cells, cell clusters, etc., for use in downstream molecular analysis. The microscope can be configured to record and store images, tile scan, and / or metadata (e.g., relative positioning, sample type, sample carrier type), microscope system settings (e.g., one or more of exposure, filter set, wavelength, etc.).
[0078] Downstream molecular biology analysis can be next-generation sequencing, microarray, real-time qPCR, mass spectrometry, blotting, etc. Other methods that do not use laser microdissection can include tissue lysis / dissolution and FACS sorting dedicated slides.
[0079] In another embodiment, different microscopes are used in combination, including two or more of an optical microscope (such as an epi-fluorescence microscope and / or a confocal microscope), an electron microscope (such as an SEM or a FIB-SEM), and a transmission electron microscope (TEM). CryoCLEM can be used to capture images under cryogenic conditions.
[0080] Figure 2 System 200 is shown, which includes a microscope 210, a computer system 220 including at least one processor, and an extractor 230. System 200 can be configured to perform all or part of method 100.
[0081] In an embodiment, microscope 210 is configured to capture at least one image of a sample. Computer system 220 can be configured to identify at least one target or a region of interest associated with the at least one target in the at least one image. Extractor 230 can be configured to extract a portion of the sample corresponding to the identified at least one target or region of interest. System 200 can be configured to analyze the extracted portion of the sample. The analysis result of the extracted portion can include information about the molecular content of the extracted portion. The target in the at least one target includes one of a single cell nucleus, a cell compartment, an intracellular compartment, a pathogen, a virus, a bacterium, a fungus, a single cell, and a cell cluster.
[0082] System 200 can be configured to generate training data for a neural network based on at least one image of the sample and the analysis result of the extracted portion of the sample. The training data can associate phenotypic information that can be obtained from one or more images with at least one of genotype information, proteotype information, and transcriptome profile information that can be obtained from the analysis of the extracted portion of the sample.
[0083] Microscope 210 can include at least one of a fluorescence microscope, a confocal microscope, a wide-field microscope, a light-sheet microscope, a super-resolution microscope, an X-ray microscope, an electron microscope, a scanning probe microscope, and a transmission electron microscope.
[0084] Extractor 230 can use at least one of laser microdissection, fluorescence-activated cell sorting, a scalpel, a needle, a pipette, and a focused ion beam scanning electron microscope to extract a portion of the sample corresponding to the identified at least one target or region of interest.
[0085] System 200 can use at least one of polymerase chain reaction (PCR)-based methods such as qPCR, microarrays, mass spectrometry (MS), and next-generation sequencing (NGS) methods to analyze the extracted portion of the sample.
[0086] Figure 3It is a process flow diagram of an exemplary method 300 for generating training data for training a neural network according to an embodiment. The method 300 can be executed by the system 200.
[0087] In step 302, at least one image of the sample is captured. The at least one image of the sample can be captured by a microscope (such as microscope 210).
[0088] In step 304, at least one target or a region of interest associated with the at least one target is identified in the at least one image.
[0089] In step 306, the part of the sample corresponding to the identified at least one target or region of interest is extracted. The part of the sample can be extracted by the extractor 230.
[0090] In step 308, the extracted part of the sample can be analyzed.
[0091] In step 310, training data for training a neural network can be generated based on the at least one image of the sample and the analysis result of the extracted part of the sample. For example, the neural network can be a convolutional neural network. The training data can include images and corresponding labels. The labels can be obtained from the analysis result of the extracted part of the sample. The training data can combine imaging (phenotype) information and measured molecular content (genotype) information. The training data can be stored in the cloud or on a computing device. The neural network can be trained using the training data. The training can be performed in the cloud or on a computing device. The trained neural network can be deployed into a computing system (such as computing system 220).
[0092] The embodiment allows a series of steps for processing and analyzing a sample to be improved over time in order to efficiently conduct experiments. For example, by adapting the preparation of a second sample, the capture of at least one image of the second prepared sample, and a neural network based on a comparison between the data obtained from at least one image of a first prepared sample and the analysis result of at least a part of the first prepared sample, the identification of a target or a region of interest can be improved compared to an initial experiment. In addition, the embodiment allows known or new biomarkers to be located based on the obtained imaging knowledge (phenotype) and measured molecular content (genotype).
[0093] List of reference numerals 100, 300 Methods 102 - 128; 302 - 310 Method steps 210 Microscope 220 Computer system 230 Extractor
Claims
1. An adaptation method for at least one step in a series of steps for processing and analyzing a sample, the method comprising: Capturing at least one image of a first prepared sample by a microscope; Identifying at least one target or region of interest in the at least one image using a neural network; Analyzing at least a portion of the first prepared sample, the portion of the first prepared sample corresponding to the identified at least one target or region of interest; and Based on a comparison between the analysis result of the at least a portion of the first prepared sample and the data obtained from the at least one image of the first prepared sample, adapting at least one of the following: The preparation of a second sample, The capture of at least one image of the second prepared sample, and The neural network.
2. The method according to claim 1, wherein The analyzing at least a portion of the first prepared sample includes extracting the portion of the first prepared sample, wherein the analyzing at least a portion of the first prepared sample includes analyzing the extracted portion of the first prepared sample.
3. The method according to claim 1 or 2, wherein The analyzing at least a portion of the first prepared sample includes using at least one of the following: Laser microdissection LMD, Laser capture microdissection LCM, Tissue lysis and subsequent fluorescence-activated cell sorting FACS, Physical local tissue extraction, Extracting whole cells or cytoplasm with a micropipette or patch pipette, Protein extraction methods, Tissue lysis or tissue dissolution and fluorescence-activated cell sorting, i.e., FACS sorting, Immunomagnetic bead cell sorting, Spatial transcriptomics, Spatial multi-omics, Indirect methods, Next-generation sequencing NGS methods, RNA sequencing RNAseq methods, Microarrays, Real-time polymerase chain reaction qPCR methods, Blotting methods, and Mass spectrometry MS.
4. The method according to any one of claims 1 to 3, wherein The microscope includes one of an epifluorescence microscope and a confocal microscope, and Wherein, at least one of an electron microscope, a scanning electron microscope SEM, a focused ion beam scanning electron microscope FIB-SEM, a transmission electron microscope TEM, a correlative light electron microscope CLEM, and a cryogenic correlative light electron microscope CryoCLEM is used to analyze the at least a portion of the first prepared sample.
5. The method according to one of claims 1 to 4, wherein Preparing at least one of the first sample and the second sample to identify at least one component in at least one of the first sample and the second sample using at least one of a stain, an ion label, an antibody dilution, a fluorescence in situ hybridization FISH probe, a fluorophore, and a detergent.
6. The method according to one of claims 1 to 5, further comprising: Obtaining data associated with at least one of the preparation of the first sample, the capture of at least one image of the first prepared sample, the at least one image of the first prepared sample, and the analysis result of the at least a portion of the sample; and Storing the data as a single data point in a repository, Wherein, the repository includes a plurality of data points corresponding to at least one of different parameters for preparing a sample, different parameters for capturing at least one image of a prepared sample, different targets or regions of interest, different samples, and different types of samples.
7. The method according to claim 6, wherein, The preparation of the second sample, the capture of at least one image of the second prepared sample, and the adaptation of at least one of the neural networks are based on the analysis of the plurality of data points.
8. The method according to claim 6, further comprising generating training data for at least one neural network based on the plurality of data points, the at least one neural network being for identifying at least one target or region of interest in at least one image.
9. The method according to claim 8, further comprising at least one of the following: Fine-tuning the neural network based on the training data, the neural network being for identifying at least one target or region of interest in at least one image; and Training a set of neural networks based on the training data, the set of neural networks being for identifying at least one target or region of interest in at least one image.
10. The method according to one of claims 1 to 9, wherein, Comprising at least one of the following: The comparison between the analysis result of at least a part of the first prepared sample and the data obtained from at least one image of the first prepared sample includes: correlating the signal intensity related to one of the stain, ion label, antibody diluent, fluorescence in situ hybridization FISH probe, fluorophore, and detergent treatment applied to the first sample and obtained from at least one image of the first prepared sample with the molecular content of the analyzed part of the first prepared sample obtained according to the analysis result of at least a part of the first prepared sample; and The comparison between the analysis result of at least a part of the first prepared sample and the data obtained from at least one image of the first prepared sample includes: combining phenotypic evaluation and genotypic evaluation.
11. A system for generating training data for a neural network, the system comprising: A microscope configured to capture at least one image of a sample; At least one processor configured to identify at least one target or region of interest associated with the at least one target in at least one image; And An extractor configured to extract a part of the sample corresponding to the identified at least one target or region of interest, Wherein the system is configured to analyze the extracted part of the sample, and Wherein the system is configured to generate training data for the neural network based on at least one image of the sample and the analysis result of the extracted part of the sample.
12. The system according to claim 11, wherein, The training data associates phenotypic information obtainable from one or more images with at least one of genotypic information, proteotypic information, and transcriptomic profile information obtainable from the analysis of the extracted part of the sample.
13. The system according to claim 11 or 12, wherein, The analysis result of the extracted part includes information about the molecular content of the extracted part.
14. The system according to any one of claims 11 to 13, comprising at least one of the following: The target in the at least one target is one of a single cell nucleus, cell compartment, intracellular compartment, pathogen, virus, bacterium, fungus, single cell, and cell cluster, and The sample includes tissue.
15. The system according to any one of claims 11 to 14, comprising at least one of the following: The microscope includes at least one of a fluorescence microscope, a confocal microscope, a wide-field microscope, a light-sheet microscope, a super-resolution microscope, an X-ray microscope, an electron microscope, a scanning probe microscope, and a transmission electron microscope, The extractor uses at least one of laser microdissection LMD, fluorescence-activated cell sorting FACS, a scalpel, a needle, a pipette, and a focused ion beam scanning electron microscope FIB-SEM to extract the part of the sample corresponding to the at least one identified target or region of interest, and The system uses at least one of polymerase chain reaction PCR-based methods such as qPCR, microarrays, mass spectrometry MS, and next-generation sequencing NGS methods to analyze the extracted part of the sample.