Synthetic image generation and machine learning analysis of biological substances
By generating synthetic microscopic images and training machine learning models, the problems of low resolution and high spot density of microscopic images are solved, and the accuracy of intensity values and base judgments are improved, which is suitable for high-resolution image generation and sequencing of biological matter.
Patent Information
- Application Number
- CN202510070000.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2024-01-29
- Filing Date
- 2025-01-16
- Publication Date
- 2025-07-29
AI Technical Summary
In the prior art In the process of biological matter sequencing, the resolution of microscopic images is low or the spot density is high, resulting in inaccurate extraction of intensity values, and biochemical processes may cause artifacts, affecting the accuracy of base judgment.
By generating synthetic microscopic images and training machine learning models, using point diffusion functions and seed intensity to simulate real microscopic images, high-resolution synthetic images are generated, and intensity values are extracted using machine learning models, and base judgment is performed in combination with residual channel attention network.
It improves the accuracy of resolution and intensity values of microscopic images, improves the accuracy of base judgment, and is especially suitable for low-resolution and high spot density images, improving the accuracy of the sequencing process.
Smart Images

Figure CN120388615A_ABST
Abstract
Description
[0001] Cross - Reference to Related Applications
[0002] This application claims the benefit and priority of U.S. Provisional Application No. 63 / 626,168, filed on January 29, 2024, the entire content of which is incorporated herein by reference in its entirety for all purposes.
[0003] Sequence Listing
[0004] This application contains a Sequence Listing that has been submitted herewith and is hereby incorporated by reference in its entirety. The.xml copy created on December 12, 2024, is named 106340-1473828-5123-CN and is 2,037 bytes in size. Technical Field
[0005] Imaging systems and methods, such as imaging systems and methods for synthetic image generation and machine learning analysis of biological materials. Background Art
[0006] Various techniques have been developed to analyze biological materials, such as tissue cells, DNA, biological fluids, etc. One such technique is sequencing, which itself includes different techniques distinguished according to the research objective (such as a single strain (bacteria, virus, etc.) or a complete population, DNA or RNA, and all present materials or only specific targeted regions). In genetics, the term sequencing refers to methods for determining the primary structure or sequence of a biological polymer, including nucleic acids (e.g., DNA, RNA, etc.). More specifically, DNA and RNA sequencing are processes for determining the order of nucleotide bases (e.g., adenine, guanine, cytosine, and thymine) in a given nucleic acid fragment. Such sequencing methods generally include determining the base at a certain position in a nucleic acid fragment, and the determined base is used to determine the sequence of the nucleic acid.
[0007] For example, when sequencing a target nucleic acid, the process generally includes extracting and fragmenting the target nucleic acid from a sample. The fragmented nucleic acid is used to generate a target nucleic acid template, which typically includes one or more adaptors. The target nucleic acid template can be subjected to an amplification method, such as bridge amplification, to replicate and generate copies of the target nucleic acid. Then, sequencing applications are performed on the copied target nucleic acid.
[0008] In traditional Sanger sequencing, each type of nucleotide (e.g., A, T, C, and G) is labeled with a different fluorescent dye. In next-generation sequencing (NGS) technologies (e.g., DNA nanoball sequencing), fluorescently labeled reversible terminators are used. Then the nucleic acid fragments are subjected to a sequencing reaction in which a new nucleic acid strand is synthesized. During this synthesis, the fluorescently labeled bases are incorporated into the growing nucleic acid strand. In Sanger sequencing, gel electrophoresis is used to separate the nucleic acid fragments by size. In NGS, the fragments are typically attached to a solid surface (such as a flow cell), and the sequencing reactions occur in a massively parallel manner. After the sequencing reaction, the fluorescent signals from the labeled bases are detected by an imaging system. Each base emits light of a characteristic wavelength due to its specific fluorescent label. The order of the bases in the sequence is determined by analyzing the fluorescent signals in the acquired image. The color and intensity of the fluorescence at each position correspond to the type of base incorporated into the sequence.
[0009] An intensity value (e.g., a fluorescent signal) corresponding to the base incorporated at a specific position in a nucleic acid sequence can indicate the base at that position. For example, four different types of fluorescence can be used, which correspond to the four types of bases to be identified. The nucleic acid sequence is amenable to relatively inexpensive and efficient imaging techniques, in which the nucleic acid sequence is captured in four color images, one for each type of fluorescence used. These four images can then be processed by software to extract intensity information. The intensity value of the target nucleic acid template can correspond to one pixel or multiple pixels of the image, or a pixel can have multiple templates (i.e., more than one template per pixel). In any case, the intensity value of each of the four bases can be assigned to the template. The images and templates are processed by specialized software to convert the fluorescent signals into a readable nucleic acid sequence. The pixel intensity at each position corresponds to the type and quantity of each labeled base, thus allowing the identification of the nucleic acid sequence. However, the resolution of the microscopic image may be low or the spot density may be high, resulting in inaccurate extraction of intensity values. Additionally, the biochemistry of the sequencing process may cause artifacts, and the intensity signals may vary significantly from one position and template to another and from sample to sample.
[0010] Accordingly, there is a desire to provide improved methods and systems for intensity extraction and base determination. SUMMARY OF THE INVENTION
[0011] In one example, a method involves: receiving a set of real microscopic images that represent a plurality of objects of biological material, each object of the plurality of objects corresponding to one or more pixels of the set of real microscopic images; obtaining a plurality of features of the set of real microscopic images and the plurality of objects; and generating one or more synthetic microscopic images based on the plurality of features, the synthetic microscopic images representing the plurality of objects of the biological material.
[0012] In some embodiments, the method further involves simulating a sequencing biochemical process performed on the biological material. The simulation is configured to receive the plurality of features of the real microscopic images in the set of real microscopic images as input. The method further involves determining, based on the input, a seed intensity for each object of the plurality of objects in the real microscopic images. The seed intensity corresponds to the signal volume of the object.
[0013] In some embodiments, the method further involves generating a seed image based on the seed intensity for each object of the plurality of objects. Each pixel in the seed image represents the signal volume of the object.
[0014] In some embodiments, generating the one or more synthetic microscopic images involves: generating a point spread function for the plurality of objects of the real microscopic images based on the plurality of features; determining a signal distribution over a plurality of pixels by pooling the point spread function and the seed intensity of each object; and generating a synthetic image of the one or more synthetic microscopic images based on the signal distribution over the plurality of pixels and the plurality of features.
[0015] In some embodiments, the method further involves training a machine learning model by using the one or more synthetic microscopic images and a corresponding set of seed images as training data to generate a trained machine learning model for generating intensity values of additional real microscopic images.
[0016] In some embodiments, the biological material includes a DNA array, an oligonucleotide array, biological tissue, or a cell array.
[0017] In some embodiments, the biological material includes a DNA array, and the plurality of objects include a plurality of DNA nanospheres.
[0018] In some embodiments, the one or more synthetic microscopic images have features that are substantially similar to the features of the set of real microscopic images.
[0019] In one example, a method involves: using seed intensities and features extracted from or known in a set of real microscopic images to generate a set of synthetic microscopic images, each synthetic microscopic image representing multiple objects of biological material; generating a set of seed images from the seed intensities, where each seed image corresponds to a synthetic microscopic image in the set of synthetic microscopic images, and where each pixel in the seed image represents the signal volume of one of the multiple objects; and training a machine learning model using the set of synthetic microscopic images and the set of seed images as training data to generate a trained machine learning model, thereby generating intensity values of additional real microscopic images.
[0020] In some embodiments, the method further involves inputting a real microscopic image into the trained machine learning model. The real microscopic image depicts additional multiple objects. The method may further involve receiving an output from the trained machine learning model, the output representing the seed intensity of each of the additional multiple objects in the real microscopic image, and generating a simulated microscopic image corresponding to the real microscopic image based on the output.
[0021] In some embodiments, the method further involves determining a difference between the real microscopic image and the simulated microscopic image.
[0022] In some embodiments, the method further involves, in response to determining the difference, determining a set of features for generating a subsequent simulated microscopic image.
[0023] In some embodiments, the trained machine learning model is a first trained machine learning model, and the method further involves, in response to determining the difference, inputting the simulated microscopic image into a second trained machine learning model and receiving, from the second trained machine learning model, a result of an adjusted simulated microscopic image corresponding to the real microscopic image.
[0024] In some embodiments, generating the set of synthetic microscopic images involves: receiving a set of real microscopic images, the set of real microscopic images representing multiple objects of the biological material, each of the multiple objects corresponding to one or more pixels of the set of real microscopic images; obtaining multiple features of the set of real microscopic images and the multiple objects; and generating one or more synthetic microscopic images based on the multiple features, the synthetic microscopic images representing the multiple objects of the biological material.
[0025] In some embodiments, the method involves simulating a sequencing biochemical process on the biological material. The simulation is configured to receive the plurality of features of the real microscopic images in the real microscopic image set as input. The method further involves determining a seed intensity for each of the plurality of objects in the real microscopic image based on the input. The seed intensity corresponds to the signal volume of the object.
[0026] In some embodiments, the method further involves generating a seed image based on the seed intensity for each of the plurality of objects. Each pixel in the seed image represents the signal volume of the object.
[0027] In some embodiments, generating the one or more synthetic microscopic images involves: generating a point spread function for the plurality of objects in the real microscopic image based on the plurality of features; determining a signal distribution over a plurality of pixels by pooling the point spread function and the seed intensity for each object; and generating a synthetic image of the one or more synthetic microscopic images based on the signal distribution over the plurality of pixels and the plurality of features.
[0028] In some embodiments, the biological material includes a DNA array, an oligonucleotide array, a biological tissue, or a cell array.
[0029] In some embodiments, the biological material includes a DNA array, and the plurality of objects includes a plurality of DNA nanoballs.
[0030] In some embodiments, the machine learning model includes a residual channel attention network.
[0031] In one example, a method for training a machine learning model involves: providing a set of real microscopic sequencing images of DNA or sequencing based on a DNA nanoball array of a reference sequence, the set of real microscopic images representing a sequencing process; providing a plurality of base calls for the set of real microscopic sequencing images determined using mapping of the reference sequence or a plurality of sequence barcodes; and generating a trained machine learning model that generates a base call probability from a sequencing image by using the set of real microscopic sequencing images and the plurality of base calls as training data.
[0032] In one example, a method involves: receiving a set of ground truth microscopic images of a plurality of objects depicting biological material, the set of ground truth microscopic images corresponding to different labeled bases in a DNA sequence; converting the set of ground truth microscopic images into an array of base determination probabilities for each of the plurality of objects by a first trained machine learning model; and determining, by a second trained machine learning model, a base determination for each of the plurality of objects in each of the ground truth microscopic images of the set of ground truth microscopic images based on the array of base determination probabilities and the sequence context of the base determination for each object.
[0033] In some embodiments, the first trained machine learning model and the second trained machine learning model are the same or different machine learning models.
[0034] In some implementations, a computer program product tangibly embodied in a non-transitory machine-readable medium is provided, the computer program product including instructions configured to cause one or more data processors to perform a computer-implemented method or operation according to any one of the preceding claims.
[0035] In some implementations, a system is provided that includes: one or more data processors; and a non-transitory computer-readable medium storing instructions that, when executed on the one or more data processors, cause the one or more data processors to perform a computer-implemented method or operation according to any one of the preceding claims. BRIEF DESCRIPTION OF THE DRAWINGS
[0036] Figure 1 is a block diagram showing an example system for synthetic image generation and machine learning analysis according to various embodiments.
[0037] Figure 2 (SEQ ID NO:1) shows an example of sampling a ground truth sequence from a reference genome according to various embodiments.
[0038] Figure 3 (SEQ ID NO:1) shows an example of a simulation of a phasing event that occurs during sequencing according to various embodiments.
[0039] Figure 4 (SEQ ID NO:1) shows an example of a simulation of an incorporation event that occurs during sequencing according to various embodiments.
[0040] Figure 5 (SEQ ID NO:1) shows an example of phasing through multiple cycles according to various embodiments.
[0041] Figure 6 Shows an example point spread function table according to various embodiments.
[0042] Figure 7 Shows an example point spread array according to various embodiments.
[0043] Figure 8 Shows an example background distribution of a point spread array according to various embodiments.
[0044] Figure 9 Shows an additional example background distribution of a point spread array according to various embodiments.
[0045] Figure 10 Shows an example synthetic microscopic image according to various embodiments.
[0046] Figure 11 Shows an example of an image through synthetic image generation and machine learning analysis according to various embodiments.
[0047] Figure 12 Shows an example of a machine learning model 800 for predicting base determination of DNA nanospheres according to various embodiments.
[0048] Figure 13 Is a flowchart of a method for generating a synthetic microscopic image according to various embodiments.
[0049] Figure 14 Is a flowchart of a method for developing an intensity extraction model according to various embodiments.
[0050] Figure 15 Is a flowchart of a method for using a base determination model according to various embodiments.
[0051] Figure 16 Shows a block diagram of an example computer system that can be used with the systems and methods according to various embodiments.
[0052] Figure 17 Shows an example result of intensity extraction using a machine learning model according to various embodiments. Detailed Description
[0053] Introduction
[0054] The embodiments described herein generally relate to synthetically generating single-pixel "images" using machine learning, where each pixel corresponds to the accurate density of DNA spots from a real digital image. These pixels are the final intensities in an array (such as a DNA or DNA nanoball (DNB) array) without signal leakage from adjacent DNA spots and background. Thus, there may be no wobbling of DNA or DNBs. Machine learning-based techniques allow for the extraction of accurate fluorescence intensities for each nucleic acid strand or DNB. Algorithm methods and / or additional machine learning-based techniques can then be used to determine base calling and calculate quality scores from the final intensities. The machine learning-based techniques described herein are particularly advantageous for low-resolution and / or high-spot-density images, such as can be observed in sequencing techniques like DNA nanoball sequencing. In such cases, a machine learning model can be trained to generate a higher-resolution image or a single-pixel image with sub-pixel resolution relative to the original image to improve intensity extraction and final base calling determination. Although the machine learning-based techniques are described herein specifically for DNA / DNB sequencing, it should be understood that these techniques are generally applicable to any type of technique for generating images (e.g., whole slide images, tissue sections, or random cell microscopy).
[0055] However, these machine learning-based techniques require single-pixel "images" with accurate intensities or high-resolution images for other applications to train the machine learning model. Two solutions are disclosed herein for generating these accurate intensity or high-resolution images for training the machine learning model.
[0056] The first solution involves high-resolution imaging and includes generating high-resolution images by using an advanced high-NA imager or by using super-resolution-based imaging that uses multiple low-resolution images of the same array to reconstruct a high-resolution image. The final intensities extracted from these images are then used as ground truth for training the machine learning model together with the real low-resolution images.
[0057] The second solution involves realistic simulated images and includes simulating realistic images as an array of overlapping signal distributions. The spacing between the centers of the distributions in the array reflects the spacing between DNBs in the real image. The signal distributions are also simulated at the nanoscale, allowing the pixel size of the image to be scaled and reflecting the pixel size of the real image. The shape of these signal distributions is determined by factors measured from the real image. Given a unique value for each distribution, the signal volume of each distribution is generated according to a simulation of the sequencing biochemistry, and the value corresponds to the distribution of DNB signals typically seen in the real data on the image. The signal volume of each DNB is saved as a seed image, where each pixel in the seed image represents the corresponding DNB signal volume. The seed image is then used as the ground truth for training a machine learning model on the simulated image.
[0058] Once a machine learning model is trained using either solution to generate a single-pixel "image" where each pixel corresponds to the accurate intensity of a DNA spot from a real digital image, one or more additional machine learning models (e.g., convolutional neural networks) can be implemented to directly analyze and predict the base calls for each DNB from the single-pixel "image". In some cases, these techniques utilize two machine learning models trained on the same data to produce highly accurate base calls. The first machine learning model takes four images corresponding to different labeled bases and converts them into an array of probabilities for the bases that each DNB might be called. The second machine learning model uses multiple cycles of base call probabilities to further improve the accuracy of the base calls by providing the sequence context for each DNB's base call. These machine learning models can be combined into a single pipeline that reconfigures the output of multiple cycles from the first model to act as the input to the second model. The result is an array of improved base call probabilities that can be used to calculate the correct base call and quality score metric for each DNB. In another embodiment, a single machine learning model (e.g., a single CNN) that performs both steps of the pipeline can be trained.
[0059] Using Machine Learning to Extract Intensities from DNA / DNB Arrays and Their Analysis
[0060] The embodiments described herein provide for synthetic image generation and intensity extraction of images (e.g., microscopic images) during sequencing. A machine learning model can be developed to determine the intensity values of objects depicted in a real image. The machine learning model can be trained on synthetic images, where the seed intensity values of the objects depicted in the synthetic images are known. The synthetic images are generated by performing mathematical modeling calculations on compiled data, rather than by using a more traditional photographic process of generating real images by focusing light waves through a camera or other optical instrument. Conventionally, due to low resolution of the image or high density of objects in the image, it may be difficult to determine the intensity values in the real image. However, since the synthetic images are generated to substantially simulate real images, the machine learning model learns during training how to interpret images with low resolution or high density objects, which enables accurate extraction of intensity values from real microscopic images during inference.
[0061] The embodiments can be applied to various microscopic images, such as images depicting tissue cells, DNA arrays, oligonucleotide arrays, etc. In an example of applying the embodiments to the sequencing technology of DNBs, the intensity value at each position in the image associated with a DNB corresponds to the base of the DNB for that cycle. Thus, in the synthetic image, each pixel corresponds to the accurate intensity of a DNA spot from the real image. These pixels are the final intensities with minimal signal leakage from adjacent DNA spots and the background in a perfect array without DNA wobbling.
[0062] The embodiments also relate to a machine learning model that can directly predict the base determination of DNA nanoballs. The machine learning model receives four input images at a first machine learning model, and the first machine learning model generates an array of probabilities of bases for each DNA nanoball. A second machine learning model receives an array of multiple cycles and adjusts the probabilities of the bases of each DNA nanoball. In this case, the machine learning model still uses microscopic images to generate base determination predictions, but bypasses the calculation of the seed intensity of the images.
[0063] Figure 1 FIG. 10 is a block diagram showing an example system 100 for synthetic image generation and machine learning analysis according to an embodiment. In this embodiment, the system 100 can include a sequencing instrument 110 and a computer system 130. The computer system 130 is connected to the sequencing instrument 110 through a direct wired or wireless connection or through a high-speed local area network (not shown). The sequencing instrument 110 includes main subsystems, such as a substrate 112 for accommodating biological substances and an imaging system 114.
[0064] Computer system 130 includes typical hardware components (not shown), which include one or more processors, input devices (e.g., keyboard, pointing device, etc.), and output devices (e.g., display device, etc.). The computer system further includes computer-readable / writable media, such as memory and storage devices (e.g., flash memory, hard disk drive, optical disk drive, magnetic disk drive, etc.), which contain computer instructions that, when executed by the processor, implement the disclosed functions. Computer system 130 may further include software and / or hardware for controlling sequencing instrument 110.
[0065] Sequencing can operate on input biological material, which can be obtained by extracting a sample from a target organism. For example, the biological material can be a DNA array, an oligonucleotide array, an array of cells, biological tissue, biological fluid, etc. In some cases, the sequencing performed is DNB sequencing, which is a high-throughput sequencing technology that can be used to determine the entire genomic sequence of a target organism. The compact DNBs generated during library preparation are captured on one or more substrates 112, and then the substrate 112 is inserted into sequencing instrument 110. Substrate 112 can be unpatterned or patterned. In the unpatterned embodiment, samples of the biological material can each be deposited at discrete locations on substrate 112, but these locations do not need to be fixed. Then substrate 112 is loaded into sequencing instrument 110, and the sequence is read with a fluorescent probe that identifies DNA or RNA bases.
[0066] More specifically, DNB sequencing involves separating the DNA to be sequenced and fragmenting it into small fragments of 100 - 350 base pairs (bp). Separation includes lysing cells and extracting DNA from the cell lysate. The high-molecular-weight DNA (usually several megabase pairs long) is then fragmented by physical or enzymatic methods to break the DNA double-strands at random intervals. For small RNA sequencing, the ideal fragment length for sequencing can be selected by gel electrophoresis, while for larger fragment DNA sequencing, the DNA fragments can be separated by bead-based size selection.
[0067] DNB sequencing further involves ligating adapter sequences to the fragments and circularizing the fragments. The adapter sequences can attach to the unknown DNA such that DNA fragments with known sequences flank the unknown DNA. In the first round of adapter ligation, right (Ad153_right) and left (Ad153_left) adapters can attach to the left and right flanks of the fragmented DNA, and the DNA is amplified by PCR. Then, a splint oligonucleotide can hybridize to the ends of the fragments that are ligated to form a circle. Exonuclease is added to remove all remaining linear single-stranded and double-stranded DNA products. The result is a complete single-stranded circular DNA template.
[0068] Once a single-stranded circular DNA template (containing sample DNA ligated to two unique adaptor sequences) is generated, the complete sequence is amplified into a long string of DNA. This is achieved by rolling circle replication using a polymerase such as Phi 29 DNA polymerase, which binds to and replicates the DNA template. The newly synthesized strand is released from the circular template, resulting in a long single-stranded DNA containing several head-to-tail copies of the circular template. The resulting nanoparticles self-assemble into tight spheres of DNA (DNBs) approximately 300 nanometers (nm) in diameter. The DNBs remain separated from each other because they are negatively charged and naturally repel each other, thus reducing tangling between different single-stranded DNA lengths.
[0069] To obtain the sequence of the nucleic acid, the DNBs are attached to a patterned array flow cell (e.g., substrate 112). The flow cell can be a silicon wafer coated with silica, titanium, hexamethyldisilazane (HMDS), and a photoresist material. The DNA nanospheres are added to the flow cell and selectively bind to the positively charged aminosilane in a highly ordered pattern, allowing for sequencing of very high-density DNA nanospheres. From there, the flow cell is loaded into the sequencing instrument 110, and combinatorial probe-anchor synthesis (cPAS) chemistry is used to hybridize sequencing primers to the DNBs, and fluorescently labeled reversible terminator probes are incorporated into successive sequencing cycles by DNA polymerase.
[0070] After each DNA nucleotide incorporation step, the fluorescent probes are excited by a laser, and the imaging system 114 captures an image of the biological material on the substrate 112. In the specific case where the biological material is DNA organized in an array of DNBs, the image represents the DNBs. From the image, the genomic sequence of each DNA nanosphere can be determined. The intensity values (e.g., fluorescence signals) corresponding to the bases incorporated into the nucleic acid at a specific position can indicate the base at that position. Since the intensity values may be difficult to extract from the image, a machine learning model 134 can be developed to predict the intensity values of real microscopic images, as described in further detail below.
[0071] To assist in developing the machine learning model 134, synthetic microscopic images with known intensity values can be generated based on real microscopic images. Each synthetic microscopic image can mimic a real microscopic image. Additionally, multiple synthetic microscopic images can be generated from a single real microscopic image. Then, the synthetic microscopic images and the known intensity values can be used to train the machine learning model 134 to extract intensity values from real microscopic images.
[0072] Generally, synthetic images are simulated with a number of quality - defining features predefined, measured, or estimated from real images such that the synthetic images are a realistic representation of the real images. The feature set includes, but is not limited to: the spacing between objects (DNB), the nanoscale size of pixels, the emission wavelength of fluorophores, the effective NA of the combination of the imaging instrument and the slide solution, the attenuation of the signal propagating along the axis due to the imaging process, the variation in the number of photons emitted by each fluorophore, the variation in the number of photons recorded by each pixel, and the background average and variation, both signal - independent and signal - related, as well as other reproducible non - uniformities and artifacts. Thus, each synthetic microscopic image has features that are substantially similar to those of a realistic microscopic image. As used herein, the terms "similar", "substantially", "about", and "approximately" are defined as mostly but not necessarily fully the specified content as understood by a person of ordinary skill in the art (and fully includes the specified content). In any disclosed embodiment, the terms "similar", "substantially", "about", or "approximately" can be replaced with "within [percentage]" of the specified content, where the percentages include 0.1%, 1%, 5%, and 10%.
[0073] In one embodiment, to generate a synthetic microscopic image, computer system 130 generates a high - resolution image by using an advanced high - numerical - aperture imager or by using super - resolution - based imaging that uses multiple low - resolution images of the same substrate to reconstruct the image at high resolution. The extracted final intensities from these images can be used as ground truth for training purposes together with the real low - resolution images.
[0074] As another example, to generate a synthetic microscopic image, computer system 130 can receive a set of real microscopic images of objects (e.g., DNA nanospheres, tissue cells, etc.) representing biological material. The feature extractor 132 of computer system 130 can then acquire the features of the real microscopic images and the objects. Acquiring the features can involve actively determining or importing the features from a stored location. These features can include the physical properties of the objects and the imaging properties of the real microscopic images. Physical properties can include the size of the objects (e.g., DNA nanosphere diameter), the distance between adjacent objects (nm spacing), the distribution of the number of photons emitted from each dye, etc. Imaging properties can include the pixel distance between objects (pixel spacing) or the pixels per object, pixel - photon registration (shot noise), signal - independent background and variation, signal - related background and differences, optical aberrations, focus accuracy n position effects, illumination non - uniformity, etc.
[0075] Additionally, computer system 130 can perform a simulation of the sequencing biochemical process to model the biochemical events that may occur during sequencing, and ultimately generate a synthetic microscopic image. The simulation can receive the features of the real microscopic images as input. AsFigure 2 As shown, the simulation can involve generating a distribution of copy numbers for different objects of real sequences sampled from a reference genome. The computer system 130 can assign copy numbers to objects from a realistic distribution by selecting copy numbers from a normal distribution or by calculating copy numbers from a combination of random insert sizes and random object qualities (e.g., in thousands). As an example, the random sequence can be ACTGTTACGAGTCGAT (SEQ ID NO:1). Additionally, the simulation can further involve modeling random phasing events (e.g., lag, consecutive, termination, and exonuclease activity), as Figure 3 shown. The read position is at position three, corresponding to the first T. Figure 3 The copy resolution is fourteen in-phase, two plus one consecutive, two minus one lag, one minus two exonuclease activities, and one termination. As Figure 4 shown, the simulation can also model incorporation events (e.g., antibody incorporation, dyes per antibody distribution, and unwashed dyes). The Poisson distribution of dyes per antibody is set as a list, where each element is the average number of dyes per antibody for the corresponding channel (e.g., A, C, T, and G). The percentage distribution has key-value pairs (antibody percentage: number of dyes per antibody). For example, (0.25:0, 0.5:1, 0.25:2) corresponds to 25% of the antibodies having no dyes, 50% of the antibodies having one dye, and 25% of the antibodies having two dyes. In Figure 4 , 50% do not incorporate, 12.5% of the antibodies have no dyes, 25% of the antibodies have one dye, and 12.5% of the antibodies have two dyes. For minus one, minus two, or termination events, no antibody binds. Phasing is repeated by antibody incorporation for a specified number of cycles. As Figure 5 shown, the number of dyes present in each channel for each cycle can be aggregated, and these values can represent the seed intensity of the object. Thus, the output of the sequencing biochemistry is the seed intensity of the objects of the real microscopic image.
[0076] Generating a synthetic microscopic image can also involve the computer system 130 generating a point spread function for the objects of the real microscopic image based on the characteristics of the real microscopic image. The point spread function can be generated using a Bessel function (e.g., Airy Disk) or a Gaussian approximation of the Airy Disk. The point spread function can produce a table of point spread functions at the nanoscale. The point spread function of the complete object can contain the aggregation of multiple point spread functions within the object diameter region. The Gaussian function is expressed as:
[0077]
[0078] Where I0 is the intensity value of the object, NA is the numerical aperture and is inversely correlated with the spread of the function, λ is the emission wavelength of the light and is directly correlated with the spread, x0 is the center point of the point spread function, and x is the point in the real microscopic image.
[0079] As Figure 6 shown, the computer system 130 can integrate the nanoscale point spread function table into an appropriate pixel size to generate a pixel-level point spread function table. The integration can be performed at the 0.01 sub-pixel read frame. In Figure 6 it, NA is 0.9, the wavelength is 500 nm, there is 150 nm / pixel, and the object size is 150 nm. The pixel-level point spread function table can be normalized to 1 such that the seed intensity of a particular object corresponds to the volume of the point spread function. The seed intensity corresponds to the signal volume of the object. The computer system 130 can then multiply the seed intensity by a scalar from a normal distribution that represents the variation in the quantum yield of the dye. Then, these values (or the seed intensity itself) can be pooled with the pixel-level point spread function table to generate the signal distribution of the object on the pixel. Pooling can involve multiplying the seed intensity of each object by the point spread function.
[0080] As Figure 6 shown, the computer system 130 can place multiple signal distributions on the image in the form of an array. The signal distributions can be spaced as specified by the characteristics of the real microscopic image. The spacing (pitch) between objects can be predefined. In Figure 7 it, the pitch is 300 nm and 1.9 pixels, where NA is 0.9. Then a background distribution can be applied to the object point spread function array, as Figure 8 shown. In one example, a Poisson distribution can be applied to the pixels of the image to represent shot noise. The Poisson distribution corresponds to the number of photons recorded by each pixel. As Figure 9 shown, two additional normal distributions can also be applied to the pixels in the image to represent signal-independent background and signal-dependent background. The signal-independent background is a normal distribution based on a scalar value that is not proportional to the signal. The signal-dependent background is a normal distribution based on a scalar value that is proportional to the 80th percentile value of the pre-background image. Then, the computer system 130 uses the characteristics to generate one or more synthetic microscopic images of the objects representing the biological material. The synthetic microscopic images can substantially simulate the real microscopic image, but the intensity values of the objects in the synthetic microscopic images are known. Figure 10 Examples of synthetic microscopic images are shown in
[0081] Thus, generally speaking, realistic images can be simulated as an array of overlapping signal distributions. The spacing between the centers of the distributions in the array can reflect the spacing between objects in a real microscopic image. The signal distributions can also be simulated at the nanoscale, allowing the pixel size of the synthetic image to be scaled to reflect the pixel size of the real image. Signal distribution parameters can include NA, wavelength, etc. The shape of these signal distributions (e.g., circular, elliptical, where the X and Y positions have different NA values, etc.) can be determined by factors measured from the real image. Additionally, the signal volume of each distribution can be generated based on the simulation of the sequencing biochemistry, giving each distribution a unique value that corresponds to the distribution of the object signal in the real data on the image. A seed image can be generated based on the signal volume (seed intensity) of each object, where each pixel in the seed image represents the corresponding object signal volume. Relative to the real microscopic image, the single-pixel intensity can be sub-pixel resolution.
[0082] Synthetic images 122 can be generated for various real microscopic images. The synthetic images 122 and their corresponding seed images 124 can be used to train a machine learning model 134 to determine the intensity values of the real microscopic images. The machine learning model 134 can use a residual channel attention network (RCAN) architecture with convolutional neural network (CNN) layers. The machine learning model 134 can also be any other suitable machine learning model trained to provide predictions, such as a generative machine learning model.
[0083] To train the machine learning model 134, the computer system 130 can obtain training data from the data repository 120. The training data can include synthetic images 122 and seed images 124. In some cases, portions of the training data (e.g., real images) can be augmented with synthetic images 122 and seed images 124. In other cases, the training data consists entirely of synthetic images 122 and seed images 124. The computer system 130 uses a trainer and a validator as part of a machine learning operationalization framework that includes hardware such as one or more processors (e.g., CPU, GPU, TPU, FPGA, etc., or any combination thereof), memory, and storage devices that are used to operate software or computer program instructions (e.g., TensorFlow, PyTorch, Keras, etc.) to perform arithmetic, logical, input, and output commands of the machine learning model 134. Specifically, the trainer performs iterative operations of training that involve inputting portions of the training data into the machine learning model 134 to find a set of model parameters (e.g., weights and / or biases) that minimize or maximize an objective function (e.g., loss function, cost function, contrastive loss function, etc.). The objective function can be constructed to measure the difference between the output inferred using the machine learning model 134 and the ground truth values annotated with labels to the images. For example, for a supervised learning-based model, the goal of training is to learn a function "h()" (sometimes also called the hypothesis function) that maps the training input space X to the target value space Y, h:X→Y, such that h(x) is a good predictor of the corresponding value of y. Various different techniques can be used to learn such a hypothesis function. In some machine learning algorithms (such as neural networks), this is done using backpropagation. The current error is typically propagated backward to the previous layer, where it is used to modify the weights and biases in a way that minimizes or maximizes the error. The weights are modified using an optimization function. The optimization function typically computes the error gradient, i.e., the partial derivative of the objective function with respect to the weights, and the weights are modified in the opposite direction of the computed error gradient. For example, techniques such as backpropagation, stochastic feedback, direct feedback alignment (DFA), indirect feedback alignment (IFA), Hebbian learning, etc., are used to update the model parameters in a way that minimizes or maximizes the objective function. This cycle is repeated until the minimum or maximum of the objective function is reached.
[0084] The trainer also uses an optimization algorithm to perform the process of selecting hyperparameters to find the parameters corresponding to the best fit between the prediction and the actual output. Example optimization algorithms include the stochastic gradient descent algorithm or its variants such as batch gradient descent or mini-batch gradient descent. Hyperparameters are settings that can be adjusted or optimized to control the behavior of the machine learning model 134. Most models have well-defined hyperparameters that control different characteristics of the model, such as memory or execution cost. However, additional hyperparameters can be defined to adapt the machine learning model 134 to a specific scenario. For example, hyperparameters can include the number of hidden units in the model, the learning rate of the model, the width of the convolutional kernel, the number of kernels in the model, the maximum depth of the trees in a random forest, the minimum sample split, the maximum number of leaf nodes, the minimum number of leaf nodes, etc.
[0085] Once a set of model parameters has been identified, the machine learning model 134 has been trained, and the computer system 130 can perform additional testing or validation processes using a subset of the training data. The validation process includes iterative operations of inputting a validation dataset into the machine learning model 134 using validation techniques such as K-Fold Cross-Validation, Leave-one-out Cross-Validation, Leave-one-group-out Cross-Validation, Nested Cross-Validation, etc. to fine-tune the hyperparameters and ultimately find the optimal set of hyperparameters. Once the optimal set of hyperparameters is obtained, a holdout set of test data from the initial split of the training data is input into the machine learning model 134 to obtain an output (in this instance, a prediction regarding intensity v), and relevant techniques such as the Bland-Altman method and Spearman's rank correlation coefficient, as well as performance metrics such as error, accuracy, precision, recall, receiver operating characteristic curve (ROC), etc. are used to evaluate the output against the ground truth. These metrics can be used to analyze the performance of the machine learning model 134 to provide recommendations.
[0086] Once trained, the computer system 130 can input a real microscopic image into the machine learning model 134, which generates an output representing the seed intensity of each object in the real microscopic image. Based on the output, the image simulator 135 can generate a simulated image 136, where each pixel of the simulated image 136 corresponds to the seed intensity of an object in the real microscopic image. That is, if the real microscopic image belongs to an array of DNA nanospheres, the simulated image 136 can have pixels corresponding to each DNA nanosphere, and the pixel values can correspond to the seed intensity of the DNA nanospheres. Thus, the machine learning model 134 allows the determination of the intensity values of objects from real microscopic images. The seed intensity can be used for base calling.
[0087] Figure 11 An example of generating an image for machine learning analysis throughout a synthetic image according to one embodiment is shown. The synthetic image generated for the real microscopic image shows two cycles of DNA sequencing (cycle 1 and cycle 500). Additionally, the seed intensity image related to the synthetic image is also shown. The synthetic image and the seed intensity image are generated as previously described. A machine learning model (e.g., Figure 1 the machine learning model 134 in Figure 11 can receive the real microscopic image of each synthetic image in the synthetic image and generate a simulated image showing intensity values. Then, an image representing the absolute error of each pixel can be generated to evaluate the performance of the model. As
[0088] shown, the machine learning model can accurately determine the seed intensity of objects from real microscopic images.
[0089] Additionally or alternatively, the image simulator 135 can include another trained machine learning model (e.g., a CNN), which serves as the last step in the image simulation. The simulated image 136 can be input into the machine learning model to generate a result of an adjusted simulated microscopic image corresponding to the real microscopic image. The adjustments made to the simulated image 136 may be changes that cannot be explained by the current set of parameters used by the machine learning model 134. To train the machine learning model, the simulated images can be used as inputs and the real microscopic images can be used as the ground truth. Both image sets can use the same seed intensity as the ground truth from the machine learning model 134, which is extracted from the real microscopic image and used as an input for generating the simulated images.
[0090] Figure 12 An example of a machine learning model 1200 for predicting base calls of DNA / DNBs according to one embodiment is shown. The machine learning model 1200 can be Figure 1 another instance of the machine learning model 134 in. The machine learning model 1200 can involve a machine learning model that can be a combination of machine learning models (e.g., CNNs) trained on the same data to produce accurate base calls. The machine learning models can be the same or different machine learning models. Once trained, the first machine learning model can receive a set of images (e.g., four images) corresponding to different labeled bases in a DNA sequence. Each image can depict a DNA array of a DNA nanoball. The first machine learning model can convert the images into an array of probabilities of the bases that may be called for each DNA nanoball. The second machine learning model can use the multiple rounds of base call probabilities to further improve the accuracy of the base calls by providing sequence context for the base calls for each DNA nanoball. The machine learning models can be combined into a single pipeline that reconfigures the outputs of multiple rounds from the first machine learning model to serve as inputs to the second machine learning model. The result is an array of improved base call probabilities that can be used (e.g., by Figure 1 the computer system 130 in) to calculate the correct base calls and quality score metrics for each DNA nanoball. In another embodiment, a model that performs both steps can be trained.
[0091] Training a machine learning model can involve providing a set of real microscopic sequencing images of DNA or DNB array-based sequencing of a reference sequence. The set of real microscopic images represents a given sequencing process. That is, the set of real images represents the particular machine or process that was used to generate them. For adequate machine and / or process representation, images from multiple different instruments of the same type, from different DNA sequencing libraries, and repeated runs of the same library on each instrument can be used. This allows for specialized training for a variety of different machines and processes.
[0092] Base calls can also be provided that are associated with the set of real microscopic sequencing images. A mapping to a reference sequence or sequence barcode can be used to determine the base calls. Then, the set of real microscopic images and the base calls can be used as training data to generate a trained machine learning model. The DNA array and associated images can have spots that are occupied by multiple DNBs or by no DNBs (e.g., DNBs that are not bound chemically or not loaded, or DNBs with several template copies). For spots with no DNBs, these spots can be trained by setting the value of the training data to zero, and for spots with a mixture of DNBs, these spots can be trained by a combination of base probabilities. This enables more accurate base calls, as well as a more accurate estimation of their probabilities. After training, the residual values of the output of the machine learning model for empty DNB spots can then be used to determine information about the background of a particular imaging system.
[0093] In one example, the training data can involve realistic simulations, and known sequences can be used as ground truth. The realistic simulations can be synthetic microscopic images generated as described Figure 1 above. Alternatively, real images can be used in the training data. The real images can be a PCR-free library from a controlled genome, or can be optionally barcoded synthetic DNA from a known DNA nanosphere genome sequence. In another example, images generated from a fully synthetic library can be used, where their sequences are fully known. The sequences generated using traditional base calling methods can be mapped, and the library sequences can be used for the mapped sequences to generate an accurate base call output array for the corresponding real images.
[0094] Training using an accurate (e.g., 99.1%, 99.3%, 99.5%, 99.7%, 99.9% or higher) sequence of DNA nanospheres determined by mapping or barcoding can eliminate the need to determine accurate sequencing intensities in real low-resolution images. A machine learning model 1200 for calling bases from images can provide computational efficiency and adaptability by adequately training on different inconsistencies and artifacts in real data, including specific training for each sequencing system, instead of separate image processing and base calling steps.
[0095] Figure 13 is a flowchart of a method 1300 for generating synthetic microscopic images according to an embodiment of the present disclosure. Figure 9 The processes depicted therein may be implemented in software (e.g., code, instructions, programs) executed by one or more processing units (e.g., processors, cores) of the corresponding systems, hardware, or combinations thereof (e.g., intelligent selection machines) described herein. The software may be stored on a non-transitory storage medium (e.g., on a memory device). In Figure 13 The methods presented and described below are intended to be illustrative and non-limiting. Although Figure 13 depicts individual processing steps occurring in a particular sequence or order, this is not intended to be restrictive. In certain alternative embodiments, these steps may be performed in some different order, or some steps may also be performed in parallel.
[0096] At block 1302, a set of real microscopic images of a plurality of objects representing biological material is received. The real microscopic images may show an array of DNA nanospheres.
[0097] At block 1304, a plurality of features of the real microscopic images and the plurality of objects are obtained. These features may include physical characteristics of the objects and imaging characteristics of the real microscopic images. A simulation of a sequencing biochemical process may be performed on the biological material to generate a distribution associated with the objects based on the features. Additionally, a point spread function may be generated based on the features such that a seed intensity corresponding to the signal volume of each object can be determined.
[0098] At block 1306, one or more synthetic microscopic images of the plurality of objects representing biological material are generated based on the plurality of features. The synthetic microscopic images may substantially simulate the real microscopic images. The synthetic microscopic images may be generated based on the features and a seed image may be generated based on the seed intensity of the objects.
[0099] Figure 14 is a flowchart of a method 1400 for developing an intensity extraction model according to an embodiment of the present disclosure. Figure 14 The processes depicted therein may be implemented in software (e.g., code, instructions, programs) executed by one or more processing units (e.g., processors, cores) of the corresponding systems, hardware, or combinations thereof (e.g., intelligent selection machines) described herein. The software may be stored on a non-transitory storage medium (e.g., on a memory device). In Figure 14 The methods presented and described below are intended to be illustrative and non-limiting. Although Figure 14 depicts individual processing steps occurring in a particular sequence or order, this is not intended to be restrictive. In certain alternative embodiments, these steps may be performed in some different order, or some steps may also be performed in parallel.
[0100] At block 1402, a synthetic microscopic image set is generated using seed intensities and features extracted from or known in a real microscopic image set. Synthetic microscopic images can be generated to simulate the real microscopic image set. Synthetic microscopic images can be generated based on the extracted features and simulations of the sequencing biology that generates the seed intensities.
[0101] At block 1404, a seed image set is generated from the seed intensities. Each seed image can correspond to a synthetic microscopic image of the synthetic microscopic image set. Each pixel in the seed image can represent the signal volume of an object.
[0102] At block 1406, the synthetic microscopic image set and the seed image set are used as training data to generate a trained machine learning model to generate intensity values of additional real microscopic images. The synthetic microscopic image set is used as input, and the seed images are used as ground truth, such that the machine learning model learns to generate intensity values of microscopic images. Then, a real microscopic image can be input into the trained machine learning model, and the intensity value of the real microscopic image can be output.
[0103] Figure 15 is a flowchart of a method 1500 for using a base calling model according to an embodiment of the present disclosure. Figure 15 The processes depicted in can be implemented in software (e.g., code, instructions, programs) executed by one or more processing units (e.g., processors, cores) of the corresponding systems, hardware, or combinations thereof (e.g., intelligent selection machines) described herein. The software can be stored on a non-transitory storage medium (e.g., on a memory device). In Figure 15 The methods presented and described below are intended to be illustrative and non-limiting. Although Figure 15 depicts individual processing steps occurring in a particular sequence or order, this is not intended to be restrictive. In certain alternative embodiments, these steps can be performed in some different order, or some steps can also be performed in parallel.
[0104] At block 1502, a real microscopic image set depicting a plurality of objects of biological material is received. The real microscopic image set corresponds to different labeled bases in a DNA sequence. Thus, the real microscopic image set can include four images. The first image can correspond to adenine, the second image can correspond to guanine, the third image can correspond to cytosine, and the fourth image can correspond to thymine.
[0105] At block 1504, a first machine learning model converts a set of ground truth microscopic images into an array of base determination probabilities for each of a plurality of objects. The first machine learning model can be a CNN that receives the set of ground truth microscopic images as input. The set of ground truth microscopic images for multiple cycles can be input into the first machine learning model. The arrays of base determination probabilities for each channel can be combined into a combined array of base determination probabilities.
[0106] At block 1506, a second machine learning model determines the base determination for each of the plurality of objects in each ground truth microscopic image of the set of ground truth microscopic images, based on the array of base determination probabilities and the sequence context of the base determination for each object (e.g., cycle number, gene localization, etc.). The second machine learning model can receive the combined array of base determination probabilities as input. The second machine learning model can output an improved array of base determination probabilities from which the base determination for each object can be determined.
[0107] Any computer system among the computer systems mentioned herein can utilize any suitable number of subsystems. Examples of such subsystems are Figure 16 shown in computer system 1600. In some embodiments, the computer system includes a single computer device, where the subsystems can be components of the computer device. In other embodiments, the computer system can include multiple computer devices with internal components, and each computer device is a subsystem.
[0108] Figure 16 The subsystems shown therein are interconnected by a system bus 1675. Additional subsystems are shown, such as printer 1674, keyboard 1678, storage device 1679, monitor 1676 coupled to a display adapter 1682, etc. Peripheral devices and input / output (I / O) devices coupled to an I / O controller 1671 can be connected to the computer system by any number of means known in the art, such as a serial port 1677. For example, the serial port 1677 or an external interface 1681 (e.g., Ethernet, Wi-Fi, etc.) can be used to connect the computer system 1600 to a wide area network such as the Internet, a mouse input device, or a scanner. The interconnection by the system bus 1675 allows the central processing unit 1673 to communicate with each subsystem and control the execution of instructions from the system memory 1672 or one or more storage devices 1679 (e.g., fixed disks such as hard disk drives or optical disks), as well as the exchange of information between subsystems. The system memory 1672 and / or the storage device 1679 can be embodied as a computer-readable medium. Any data among the data mentioned herein can be output from one component to another and can be output to a user.
[0109] A computer system may include, for example, multiple identical components or subsystems connected together by an external interface 1681 or by an internal interface. In some embodiments, the computer system, subsystem, or device may communicate via a network. In such cases, one computer may be considered a client and another computer may be considered a server, where each computer may be part of the same computer system. The client and the server may each include multiple systems, subsystems, or components.
[0110] Examples
[0111] Figure 17 Example results of intensity extraction using a machine learning model according to various embodiments are shown. For analysis, synthetic images were generated with a feature set similar to real data. Two machine learning models were trained for different DNA nanosphere spacings of 715 nm pitch and 500 nm pitch. The base calling performance of the intensities of the top 95% DNA nanospheres sorted by size in 500 cycles extracted using the machine learning model was evaluated. The base calling performance is expressed as the mapping and mismatch percentages. The base calling performance of the machine learning model was compared with the base calling performance from the seed intensity, which corresponds to perfect intensity extraction. Overall, it can be seen that the base calling performance between the seed intensity and the intensity extracted using the machine learning model is similar, as for the 500 nm pitch, the difference in the mismatch percentage between the seed intensity and the machine learning model is only 0.025%. This demonstrates the effectiveness of using the machine learning model for intensity extraction.
[0112] Additional Considerations
[0113] It should be understood that any embodiment of the present disclosure may be implemented in a modular or integrated manner using hardware (e.g., an application specific integrated circuit or a field programmable gate array) and / or using computer software with a general programmable processor in the form of control logic. As used herein, a processor includes a multi-core processor located on the same integrated chip, or multiple processing units located on a single circuit board or networked. Based on the disclosure and teachings provided herein, those of ordinary skill in the art will know and understand other ways and / or methods of implementing the embodiments of the present disclosure using hardware as well as combinations of hardware and software.
[0114] Any of the software components or functions described in this application can be implemented as software code executed by a processor using any suitable computer language (e.g., Java, C, C++, C#, or scripting languages such as Perl or Python) using, for example, conventional or object-oriented techniques. The software code can be stored as a series of instructions or commands on a computer-readable medium for storage and / or transmission, and suitable media include random access memory (RAM), read-only memory (ROM), magnetic media such as hard disk drives or floppy disks, or optical media such as compact discs (CDs) or digital versatile discs (DVDs), flash memory, etc. The computer-readable medium can be any combination of such storage or transmission devices.
[0115] Such programs can also be encoded and transmitted using carrier signals suitable for transmission via wired, optical, and / or wireless networks (including the Internet) in accordance with various protocols. Thus, a computer-readable medium according to embodiments of the present disclosure can be created using data signals encoded with such programs. The computer-readable medium encoded with program code can be packaged together with a compatible device or provided separately from other devices (e.g., downloaded via the Internet). Any such computer-readable medium can reside on or within a single computer product (e.g., a hard disk drive, a CD, or an entire computer system) and can exist on or within different computer products within a system or network. The computer system can include a monitor, a printer, or other suitable display for providing any of the results mentioned herein to a user.
[0116] A computer system including one or more processors can be used to perform, in whole or in part, any of the methods described herein, and the one or more processors can be configured to execute these steps. Thus, embodiments can be directed to a computer system configured to execute the steps of any of the methods described herein, and the computer system may have different components for executing the corresponding steps or groups of corresponding steps. Although presented as numbered steps, the steps of the methods herein can be executed at the same time or in a different order. Additionally, portions of these steps can be used in conjunction with portions of other steps from other methods. Moreover, all or part of these steps can be optional. Additionally, any of the steps of any method can be performed by a module, a circuit, or other means for executing these steps.
[0117] Without departing from the spirit and scope of the embodiments of the present disclosure, the specific details of particular embodiments can be combined in any suitable manner. However, other embodiments of the present disclosure can relate to particular embodiments related to each individual aspect or specific combinations of these individual aspects.
[0118] For purposes of illustration and description, the foregoing description of the exemplary embodiments of the present disclosure has been presented. The foregoing description is not intended to be exhaustive or to limit the disclosure to the precise forms described, and many modifications and variations are possible in light of the above teachings. For the best explanation of the principles of the disclosure and its practical application, the described embodiments are chosen and described so that others skilled in the art can best utilize the disclosure in various embodiments and utilize various embodiments with various modifications suitable for the particular purposes contemplated.
[0119] As used herein, when an action is “based on” something, this means that the action is based at least in part on at least a portion of something. As used herein, the terms “similar,” “substantially,” “about,” and “approximately” are defined as being mostly but not necessarily entirely what a person of ordinary skill in the art would understand as the specified content (and completely includes the specified content). Additionally, the terms “similar,” “substantially,” “about,” or “approximately” may be replaced with “within [percentage]” of the specified content, where the percentage includes 0.1%, 1%, 5%, and 10%. Further, as used herein, the recitation of “a / an” or “the” is intended to mean “one or more” unless specifically stated to the contrary.
Claims
1. A method, the method comprising: Receiving a set of real microscopic images, the set of real microscopic images representing a plurality of objects of biological material, each object of the plurality of objects corresponding to one or more pixels of the set of real microscopic images; Obtaining a plurality of features of the set of real microscopic images and the plurality of objects; And Generating one or more synthetic microscopic images based on the plurality of features, the synthetic microscopic images representing the plurality of objects of the biological material.
2. The method according to claim 1, further comprising: Performing a simulation of a sequencing biochemical process on the biological material, wherein the simulation is configured to receive the plurality of features of the real microscopic images in the set of real microscopic images as input; and Determining a seed intensity for each object of the plurality of objects of the real microscopic images based on the input, wherein the seed intensity corresponds to the signal volume of the object.
3. The method according to claim 2, further comprising: Generating a seed image based on the seed intensity for each object of the plurality of objects, wherein each pixel in the seed image represents the signal volume of the object.
4. The method according to claim 2, wherein generating the one or more synthetic microscopic images comprises: Generating a point spread function for the plurality of objects of the real microscopic images based on the plurality of features; Determining a signal distribution over a plurality of pixels by pooling the point spread function and the seed intensity for each object; and Generating a synthetic image of the one or more synthetic microscopic images based on the signal distribution over the plurality of pixels and the plurality of features.
5. The method according to claim 1, further comprising: Training a machine learning model using the one or more synthetic microscopic images and a corresponding set of seed images as training data to generate a trained machine learning model for generating intensity values of additional real microscopic images.
6. The method according to claim 1, wherein the biological material comprises a DNA array, an oligonucleotide array, a biological tissue, or a cell array.
7. The method according to claim 1, wherein the biological material comprises a DNA array, and the plurality of objects comprise a plurality of DNA nanospheres.
8. The method according to claim 1, wherein the one or more synthetic microscopic images have features that are substantially similar to the features of the set of real microscopic images.
9. A computer program product tangibly embodied in a non-transitory machine-readable medium, the computer program product comprising instructions configured to cause one or more data processors: Generate a set of synthetic microscopic images using seed intensities and features extracted from or known in a set of real microscopic images, each synthetic microscopic image representing a plurality of objects of biological material; Generate a set of seed images from the seed intensities, wherein each seed image corresponds to a synthetic microscopic image in the set of synthetic microscopic images, and wherein each pixel in the seed image represents the signal volume of one of the plurality of objects; and Training a machine learning model using the synthetic microscopic image set and the seed image set as training data to generate a trained machine learning model, thereby generating intensity values of additional real microscopic images.
10. The computer program product according to claim 9, further comprising instructions configured to cause the one or more data processors to: Input a real microscopic image into the trained machine learning model, the real microscopic image depicting an additional plurality of objects; Receive an output from the trained machine learning model, the output representing the seed intensity of each of the additional plurality of objects in the real microscopic image; and Generate a simulated microscopic image corresponding to the real microscopic image based on the output.
11. The computer program product according to claim 10, further comprising instructions configured to cause the one or more data processors to: Determine the difference between the real microscopic image and the simulated microscopic image.
12. The computer program product according to claim 11, further comprising instructions configured to cause the one or more data processors to: In response to determining the difference, determine a feature set for generating a subsequent simulated microscopic image.
13. The computer program product according to claim 11, wherein the trained machine learning model is a first trained machine learning model, and wherein the computer program product further comprises instructions configured to cause the one or more data processors to: In response to determining the difference, input the simulated microscopic image into a second trained machine learning model; and Receive a result of an adjusted simulated microscopic image corresponding to the real microscopic image from the second trained machine learning model.
14. The computer program product according to claim 9, further comprising instructions configured to cause the one or more data processors to: Generate the synthetic microscopic image set by: Receiving the real microscopic image set, the real microscopic image set representing the plurality of objects of the biological material, each of the plurality of objects corresponding to one or more pixels of the real microscopic image set; Obtaining a plurality of features of the real microscopic image set and the plurality of objects; And Generating one or more synthetic microscopic images based on the plurality of features, the synthetic microscopic images representing the plurality of objects of the biological material.
15. The computer program product according to claim 14, further comprising instructions configured to cause the one or more data processors to: Perform a simulation of a sequencing biochemical process on the biological material, wherein the simulation is configured to receive the plurality of features of the real microscopic images in the real microscopic image set as input; and Determine the seed intensity of each of the plurality of objects in the real microscopic image based on the input, wherein the seed intensity corresponds to the signal volume of the object.
16. The computer program product according to claim 15, further comprising instructions configured to cause the one or more data processors: Generate a seed image based on the seed intensity of each of the plurality of objects, wherein each pixel in the seed image represents the signal volume of the object.
17. The computer program product according to claim 15, wherein generating the one or more synthetic microscopic images comprises: Generating a point spread function of the plurality of objects of the real microscopic image based on the plurality of features; Determining a signal distribution over a plurality of pixels by pooling the point spread function and the seed intensity of each object; and Generating a synthetic image of the one or more synthetic microscopic images based on the signal distribution over the plurality of pixels and the plurality of features.
18. The computer program product according to claim 9, wherein the biological material comprises a DNA array, an oligonucleotide array, a biological tissue or a cell array.
19. The computer program product according to claim 9, wherein the biological material comprises a DNA array, and the plurality of objects comprise a plurality of DNA nanospheres.
20. A system comprising: One or more data processors; and A non-transitory computer-readable medium storing instructions that, when executed on the one or more data processors, cause the one or more data processors to: Generate a set of synthetic microscopic images using seed intensities and features extracted from or known in a set of real microscopic images, each synthetic microscopic image representing a plurality of objects of biological material; Generate a set of seed images from the seed intensities, wherein each seed image corresponds to a synthetic microscopic image in the set of synthetic microscopic images, and wherein each pixel in the seed image represents the signal volume of one of the plurality of objects; and Train a machine learning model using the set of synthetic microscopic images and the set of seed images as training data to generate a trained machine learning model for generating intensity values of additional real microscopic images.