Methods and systems for integrated sequencing, analysis, and dignostic reporting

Un-patterned flowcells and edge-based LLMs in a single apparatus address inefficiencies in NGS by enabling high-density cluster formation and localized data analysis, improving sequencing efficiency and reducing diagnostic times.

WO2026112861A1PCT designated stage Publication Date: 2026-06-04GENESENSE HONG KONG LTD

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
GENESENSE HONG KONG LTD
Filing Date
2024-11-28
Publication Date
2026-06-04

AI Technical Summary

Technical Problem

Existing sequencing technologies using patterned flowcells in Next-Generation Sequencing (NGS) face limitations in cluster density and spatial arrangement due to optical device constraints, leading to inefficient sequencing performance and altered surface properties, while current deep learning models are less specialized and require remote access, prolonging data analysis times and increasing costs.

Method used

The use of un-patterned flowcells for cluster formation and edge-based large-language models (LLMs) integrated within a single apparatus for sequencing and diagnostic reporting, enabling high-density, spatially close clusters detectable by high-resolution imaging, and performing integrated data analysis offline.

Benefits of technology

This approach enhances sequencing efficiency and accuracy, reduces manual clinical detection errors, and shortens diagnostic report generation to hours from days, providing cost-effective, specialized, and localized data analysis.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024135118_04062026_PF_FP_ABST
    Figure CN2024135118_04062026_PF_FP_ABST
Patent Text Reader

Abstract

An apparatus for performing an integrated process of sequencing a biological sample and providing diagnostic outputs based on an edge-based large-language model (LLM) is provided. The apparatus comprises: a fluidic sub-system configured to generate fluorescent signals based on analysis of the biological sample; an imaging sub-system coupled to the fluidic sub-system to sense the fluorescent signals and obtain images of the fluorescent signals; one or more processors of a computing device; and memory storing one or more instructions. The instructions, when executed by the one or more processors, cause the computing device to perform: basecalling using the images of the fluorescent signals to obtain sequencing data of the biological sample; and analyzing the sequencing data based on the edge-based LLM stored at the apparatus. The edge-based LLM is operatable at the apparatus in an offline mode.
Need to check novelty before this filing date? Find Prior Art

Description

METHODS AND SYSTEMS FOR INTEGRATED SEQUENCING, ANALYSIS, AND DIGNOSTIC REPORTINGFIELD OF TECHNOLOGY

[0001] The present disclosure relates generally to biological material sequencing, analysis, and diagnostic reporting, and more specifically to systems, devices, and methods for performing an integrated process of sequencing of a biological sample and providing diagnostic outputs based on an edge-based large-language model (LLM) .BACKGROUND

[0002] Sequencing-by-synthesis is a method used to identify sequences of segments (also referred to as strands) of nucleic acid (e.g., DNA) molecules. Sanger sequencing is a first-generation sequencing technique that uses the sequencing-by-synthesis method. Historically, Sanger sequencing has a high degree of accuracy but is low in sequencing throughput. Second-generation sequencing techniques (also referred to as next generation sequencing or NGS techniques) massively increase the throughput of the synthesizing process by parallelizing many reactions similar to those in Sanger sequencing. Third-generation sequencing techniques allow direct sequencing of single nucleic acid molecules. Regardless of the sequencing technologies used, basecalling can be performed and diagnostic report is provided manually by a specialist or clinician. Basecalling is an essential process by which an order of the nucleotide bases in a template strand is inferred during or after a sequencing readout.SUMMARY

[0003] Next-Generation Sequencing (NGS) is a genomic tool that began to rapidly develop in recent years. NGS, in contrast to traditional Sanger sequencing methods, offers higher efficiency, larger throughput, and lower costs, enabling the simultaneous sequencing of hundreds of thousands to millions of DNA molecules. In NGS, the sample, after being processed into a library, is immobilized on a sequencing chip, and DNA clusters are formed via methods such as bridge polymerase chain reaction (PCR) or emulsion PCR. The process begins with anchoring DNA fragments to a flowcell, followed by clonal amplification to generate clusters of identical DNA molecules. These clusters are essential for producing fluorescent signals that can be detected by sequencing instruments. In particular, for example, in NGS, patterned flowcells have been used to facilitate the growth of DNA clusters. One problem with the existing technologies of using patterned flowcells to grow clusters is that the density of the cluster and spatial arrangements of the clusters are relatively fixed in order to satisfy the separation requirements due to the limitation of the conventional optical device (e.g., an optical microscope) or imaging device (e.g., a low resolution / low throughput CCD camera) .

[0004] Another problem with using the existing patterned flowcells in the NGS relates to the alteration of the surface properties of the nanoscale wells. The nanoscale wells are formed on the planar surface of the flowcell and therefore the surface properties may be altered, impacted, or even destroyed by the nanoscale wells. In turn, the sequencing performance may be negatively affected due to the change of the surface properties around the nanoscale wells. In the present disclosure, un-patterned flowcells are used for forming clusters. Un-patterned flowcells have no wells and the surface is flat, thereby eliminating the aforementioned problems. The clusters of strands can also be controlled intelligently in one or more aspects such as density, spatial arrangements, and sizes. The clusters formed using technologies disclosed herein can have higher density and can be spatially closer to each other, and yet still be detectable and distinguishable by using high-throughput and high resolution image devices.

[0005] In NGS, after clonal amplification to form clusters on the patterned flowcells, Sequencing by Synthesis (SBS) is initiated. The SBS is performed by recording the addition of each base through the fluorescent or luminescent signals produced in each synthesis cycle. During this phase, the fluorescent emissions from the clusters are captured by an optical device or an imaging device as raw imaging data. And the basecalling algorithm converts raw imaging data from a sequencer into a sequence of nucleotide bases-Adenine (A) , Cytosine (C) , Guanine (G) , and Thymine (T) . The process analyzes the fluorescent signals to determine the nucleotides present at each position in the DNA strand being sequenced. The basecalling algorithm processes and interprets the raw sequencing image data to generate sequencing data in a certain file format (e.g., a FastQ file) .

[0006] In NGS, the sequencing data can be further processed for obtaining genetic sequences, performing alignment to a reference genome, performing variant detection and annotation, and performing other analyses based on various algorithms. Due to its high-throughput nature, NGS can generate a large amount of data in a short time period. Therefore, due to the vast amount of data and the complexity of the analysis process, there are challenges related to the currently existing bioinformatics analysis algorithms and methods. For example, the existing algorithms and methods may be time consuming, inaccurate, inconsistent, inefficient, and not cost-effective.

[0007] Accompanying the flourishing of Artificial Intelligence (AI) and Big Data, AI applications in all life science fields are growing. Deep learning models have shown potential in life science areas such as sequence alignment, variant detection, gene expression analysis, epigenetic studies, protein structure prediction, and clinical applications. Some of these deep learning models are large language models (LLMs) . Existing LLMs, however, may be less specialized in knowledge and may not be used if the device operates in an offline mode. The present disclosure provides LLMs that can overcome the aforementioned deficiencies. The LLMs used in the present disclosure can be customized based on the user selection, and can obtain expert knowledge automatically from a network (e.g., the Internet) or another data source. The LLMs can be edge-based LLMs stored locally at the apparatus for performing the integrated process of sequencing of a biological sample and providing diagnostic outputs. Such an apparatus is also referred to as an all-in-one apparatus. Thus, the LLMs disclosed herein can improve the overall performance and efficiency.

[0008] The present disclosure further provides an apparatus that can perform an integrated process of sequencing of a biological sample and providing diagnostic outputs based on an edge-based large-language model (LLM) stored locally at the apparatus. The apparatus described herein can integrate all aforementioned processes including performing next generation sequencing, SBS, secondary data analysis including variant calling, alignments to reference genomes, and tertiary data analysis including diagnostic report generation, report analysis, and providing clinical advice. In some embodiments, all the processes can be performed by a single machine. In some embodiments, the apparatus can perform these processes in an offline mode without requiring remote access to an LLM model.

[0009] The apparatus and methods described herein can also reduce the efforts of manual clinical detections and diagnostics, which may be prone to error and inefficient. The apparatus and methods described herein can further provide useful suggestions to the medical progressional (e.g., a doctor) and shorten the time of generating clinical detection and diagnosis report and making therapeutic regimen. In addition, the apparatus described herein can further receive user feedback (e.g., the efficacy of the treatment) and adjust the one or more processes described above to improve the overall performance.

[0010] In some embodiments of the invention, an apparatus for performing an integrated process of sequencing a biological sample and providing diagnostic outputs based on an edge-based large-language model (LLM) is provided. The apparatus comprises: a fluidic sub-system configured to generate fluorescent signals based on analysis of the biological sample; an imaging sub-system coupled to the fluidic sub-system to sense the fluorescent signals and obtain images of the fluorescent signals; one or more processors of a computing device; and memory storing one or more instructions. The instructions, when executed by the one or more processors, cause the computing device to perform: basecalling using the images of the fluorescent signals to obtain the sequencing data of the biological sample; analyzing the sequencing data based on the edge-based LLM stored at the apparatus; and providing, via a user interface, the diagnostic outputs based on analysis results of the sequencing data. The edge-based LLM is operatable at the apparatus in an offline mode.

[0011] These and other embodiments are described more fully below.BRIEF DESCRIPTION OF THE DRAWINGS

[0012] FIG. 1A illustrates an example apparatus for performing an integrated process of sequencing a biological sample and providing diagnostic outputs in accordance with an embodiment of the present invention;

[0013] FIG. 1B illustrates an example apparatus for performing an integrated process based on an edge-based large language model (LLM) in an offline mode, in accordance with an embodiment of this present invention;

[0014] FIG. 1C is a flowchart illustrating an example method of performing an integrated process of sequencing a biological sample and providing diagnostic outputs based on an edge-based large-language model (LLM) stored locally at the apparatus, in accordance with an embodiment of this present invention;

[0015] FIG. 2 illustrates an exemplary sequencing-by-synthesis process using an NGS system in accordance with an embodiment of the present invention;

[0016] FIG. 3 is a diagram illustrating prior art patterned flowcells having many wells for growing clusters of strands;

[0017] FIG. 4A is a block diagram illustrating a method performed by an apparatus for controlling the formation of clusters on the planar surfaces of un-patterned flowcells in accordance with an embodiment of the present invention;

[0018] FIG. 4B is a block diagram illustrating various aspects of clusters controllably formed on un-patterned flowcells, in accordance with an embodiment of the present invention;

[0019] FIG. 4C is a flowchart illustrating a method of controlling cluster formation on un-patterned flowcells, in accordance with an embodiment of the present invention;

[0020] FIG. 5 is a block diagram illustrating an example image sub-system configured to sense fluorescent signals and obtain images of the fluorescent signals, in accordance with an embodiment of the present invention;

[0021] FIG. 6A is a flowchart illustrating a method of performing basecalling based on a self-attention transformer-based neural network, in accordance with an embodiment of the present invention;

[0022] FIG. 6B is a block diagram illustrating an encoder-decoder based convolutional neural network (CNN) , in accordance with an embodiment of the present invention;

[0023] FIG. 6C is a block diagram illustrating a transformer-based neural network, in accordance with an embodiment of the present invention;

[0024] FIG. 7A is a flowchart illustrating a method of performing sequencing data analysis, in accordance with an embodiment of the present invention;

[0025] FIG. 7B illustrates an example method of performing sequencing data analysis including variant calling and pathogenic detection, in accordance with an embodiment of the present invention;

[0026] FIG. 7C is a diagram illustrating a data flow deployment of an integrated process of sequencing a biological sample and providing diagnostic outputs based on an edge-based large-language model (LLM) stored locally at an apparatus, in accordance with an embodiment of the present invention;

[0027] FIG. 8 is a flowchart illustrating a method for receiving user feedback and making adjustment of the integrated process, in accordance with an embodiment of the present invention; and

[0028] FIG. 9 illustrates a block diagram of an exemplary computing device that may incorporate embodiments of the present invention.

[0029] While the embodiments of the present invention are described with reference to the above drawings, the drawings are intended to be illustrative, and other embodiments are consistent with the spirit, and within the scope, of the invention.DETAILED DESCRIPTION

[0030] The various embodiments now will be described more fully hereinafter with reference to the accompanying drawings, which form a part hereof, and which show, by way of illustration, specific examples of practicing the embodiments. This specification may, however, be embodied in many different forms and should not be construed as limited to the embodiments set forth herein; rather, these embodiments are provided so that this specification will be thorough and complete, and will fully convey the scope of the invention to those skilled in the art. Among other things, this specification may be embodied as methods or devices. Accordingly, any of the various embodiments herein may take the form of an entirely hardware embodiment, an entirely software embodiment or an embodiment combining software and hardware aspects. The following specification is, therefore, not to be taken in a limiting sense.

[0031] As described above, Next-Generation Sequencing (NGS) is a genomic tool that began to rapidly develop in the recent years. NGS, in contrast to traditional Sanger sequencing methods, offers higher efficiency, larger throughput, and lower costs, enabling the simultaneous sequencing of hundreds of thousands to millions of DNA molecules. In NGS, the sample, after being processed into a library, is immobilized on a sequencing chip, and DNA clusters are formed via methods such as bridge polymerase chain reaction (PCR) or emulsion PCR. The process begins with anchoring DNA fragments to the flow cell, followed by clonal amplification to generate clusters of identical DNA molecules. These clusters are essential for producing fluorescent signals that can be detected by sequencing instruments. In particular, for example, in NGS, patterned flowcells have been used to facilitate the growth of DNA clusters. Each of the patterned flowcells has an array of nanoscale wells for growing the clusters of strands. These nanoscale wells are designed to hold individual DNA NanoBalls during the sequencing process, ensuring that each of these NanoBalls are spatially separated and identifiable by the imaging device. Each of these DNA NanoBalls includes thousands of identical DNA molecules. The nanoscale wells are coated with primers and reagents necessary for DNA amplification and sequencing. One problem with the existing technologies of using patterned flowcells to grow clusters is that the density of the clusters and spatial arrangements of the clusters are relatively fixed in order to satisfy the separation requirements due to the limitation of the conventional optical device (e.g., an optical microscope) or imaging device (e.g., a low resolution / low throughput camera) . For example, if the cluster density is too high, the conventional optical device and / or imaging device may not be able to separate fluorescent signals from different clusters. Thus, using patterned flowcells limits the capability and efficiency of the overall sequencing process.

[0032] Another problem with using the existing patterned flowcells in the NGS relates to the altered surface properties of the nanoscale wells. The nanoscale wells are formed on the planar surface and therefore the surface properties may be altered, impacted, or even destroyed by the nanoscale wells. Such surface properties may include, for example, surface chemical concentrations in different parts of the nanoscale wells, surface curvatures, surface roughness, surface tension, etc. In turn, the sequencing performance may be negatively affected due to the change of the surface properties around the nanoscale wells. In the present disclosure, un-patterned flowcells are used for forming clusters. Un-patterned flowcells have no wells and the surface is flat, thereby eliminating the aforementioned problems. The clusters of strands can also be controlled intelligently in one or more aspects such as density, spatial arrangements, and sizes. As described in more detail below, the clusters formed using technologies disclosed herein can have a higher density and can be spatially closer to each other, and yet still be detectable and distinguishable by high-throughput and high-resolution imaging sensors.

[0033] In NGS, after clonal amplification to form clusters on the flowcells, Sequencing by Synthesis (SBS) is initiated. The SBS is performed by recording the addition of each base through the fluorescent or luminescent signals produced in each synthesis cycle. During this phase, the fluorescent emissions from the clusters are captured by an optical device or an imaging device as raw imaging data. And the basecalling algorithm converts raw imaging data from a sequencer into a sequence of nucleotide bases-Adenine (A) , Cytosine (C) , Guanine (G) , and Thymine (T) . The process analyzes the fluorescent signals to determine the nucleotides present at each position in the DNA strand being sequenced. The basecalling algorithm processes and interprets the raw sequencing image data to generate sequencing data in a certain file format (e.g., a FastQ file) . This file format is used in bioinformatics to store nucleotide sequences along with their corresponding quality scores. An example SBS process is illustrated more completely using FIG. 2 below.

[0034] In NGS, the sequencing data can be further processed for obtaining genetic sequences, performing alignment to a reference genome, performing variant detection and annotation, and performing other analyses based on various algorithms. Due to its high-throughput nature, NGS can generate a large amount of data in a short time period. Therefore, due to the vast amount of data and the complexity of the analysis process, there are challenges related to the currently existing bioinformatics analysis algorithms and methods. For example, the FastQ files containing the sequencing data may undergo analysis to detect variations or identify pathogens. Such a process is also referred to as the secondary data analysis. Afterwards, a human specialist, who has an intimate knowledge of interpretation of the data in the context of existing medical literature and databases, makes further data analysis manually (also referred to as the tertiary data analysis) and generates a comprehensive report referring to the secondary data analysis. Typically, a clinician reviews the report and offers clinical interpretation and decision-making, involving correlating genetic variants with disease phenotypes and determining appropriate treatment plans or follow-up actions. These secondary and tertiary data analyses have been conventionally manual or semi-automated processes. Therefore, they may be time consuming, inaccurate, inconsistent, inefficient, and not cost-effective.

[0035] Accompanying the flourishing of Artificial Intelligence (AI) and Big Data, AI applications in all life science fields are growing. In contrast to traditional algorithms, deep learning models can enhance analytical efficiency through automatically learning features from raw data, thereby reducing the demand of manual feature engineering. Deep learning models have improved capabilities of handling high-dimensional data and effectively capturing complex patterns and correlations, and have demonstrated higher predictive accuracy compared to traditional models. Therefore, deep learning models have shown potential in life science areas such as sequence alignment, variant detection, gene expression analysis, epigenetic studies, protein structure prediction, and clinical applications.

[0036] Some of these deep learning models are large language models (LLMs) . In various fields, LLMs have significantly increased processing efficiency and showed improved capabilities. In the fields of medical and healthcare, for example, LLMs demonstrated capabilities like analyzing vast amounts of data and providing unprecedented support in areas such as diagnostics, treatment planning, and personalized medicine. LLMs’ capabilities to understand and generate human-like text accelerates advancements, making healthcare delivery more efficient, accurate, and accessible. Certain LLMs, however, may be only accessible remotely. For example, certain models may be stored in cloud devices and therefore network is required. Remote access caused time delays and sometimes could be significant delays. In addition, the LLMs may be trained only with general knowledge and thus may not be specialized in certain medical or technology area (e.g., pathogenic detection for certain diseases) . Therefore, the existing LLMs may be less specialized in knowledge and may be incapable of performing offline model operation. The present disclosure provides LLMs that can overcome the aforementioned deficiencies of LLMs. The LLMs used in the present disclosure can be customized based on selection of the user, and can obtain expert knowledge automatically from a network (e.g., the Internet) or another data source. The LLMs can be edge-based LLMs stored locally at the apparatus for performing the integrated process of sequencing of a biological sample and providing diagnostic outputs. Such an apparatus is also referred to as an all-in-one apparatus. Thus, the LLMs disclosed here in can improve the efficiency.

[0037] As described above, secondary and tertiary data analyses may be performed after the SBS process. The secondary data analysis and tertiary data analysis, however, are usually dispersed. They may involve several places, instruments, and specialized people to generate and analyze report, and to make clinical decision. It generally takes a day to several weeks to accomplish these analyses. Therefore, it is time consuming, inefficient, and not cost-effective. The present disclosure provides an apparatus that can perform an integrated process of sequencing a biological sample and providing diagnostic outputs based on an edge-based large-language model (LLM) stored locally at the apparatus. The apparatus may be a single machine as shown in FIG. 1B and is described in more detail below. The apparatus described herein can integrate all aforementioned processes including performing next generation sequencing; SBS; secondary data analysis including variant calling, alignments to reference genomes; and tertiary data analysis including diagnostic report generation, report analysis, and providing clinical advice. In some embodiments, all the processes can be performed by a single machine. In some embodiments, the apparatus can perform these processes in an offline mode without requiring remote access to an LLM model.

[0038] The apparatus and methods described herein can also reduce the efforts of manual clinical detections and diagnostics, which may be prone to error and inefficiency. The apparatus and methods described herein can further provide useful suggestions to the medical progressional (e.g., a doctor) and shorten the time of generating clinical detection and diagnosis report and making therapeutic regimen. In addition, the apparatus described herein can further receive user feedback (e.g., the efficacy of the treatment) and adjust the one or more processes described herein to improve the overall performance.

[0039] Embodiments of the present invention discussed herein provide an apparatus for performing an integrated process of sequencing a biological sample and providing diagnostic outputs based on an edge-based large-language model (LLM) . The apparatus comprises: a fluidic sub-system configured to generate fluorescent signals based on analysis of the biological sample; an imaging sub-system coupled to the fluidic sub-system to sense the fluorescent signals and obtain images of the fluorescent signals; one or more processors of a computing device; and memory storing one or more instructions. The instructions, when executed by the one or more processors, cause the computing device to perform: basecalling using the images of the fluorescent signals to obtain the sequencing data of the biological sample; analyzing the sequencing data based on the edge-based LLM stored at the apparatus; and providing, via a user interface, the diagnostic outputs based on analysis results of the sequencing data. The edge-based LLM is operatable at the apparatus in an offline mode. Details of the embodiments of the present invention are described below. Structure of an apparatus for performing an integrated process of  sequencing of  a biological sample and providing diagnostic outputs based on an edge-based LLM

[0040] FIG. 1A is a block diagram illustrating an exemplary apparatus 100 for performing an integrated process of sequencing a biological sample and providing diagnostic outputs. Apparatus 100 may include an analytical system 110. As illustrated in FIG. 1, analytical system 110 includes an optical sub-system 120, an imaging sub-system 118, a fluidic sub-system 112, a control sub-system 114, sensors 116, and a power sub-system 122. Analytical system 110 can be used to perform next-generation sequencing (NGS) reactions and produce fluorescence images 140 captured during multiple synthesis cycles. Fluorescence images 140 are images of fluorescence signals. Images 140 are provided to computing device (s) 103 for further processing, such as basecalling, secondary data analysis, tertiary data analysis, and diagnostic report generation. As described above, the secondary data analysis includes, for example, variant detection, genome alignments, and pathogenic detection. The tertiary data analysis includes, for example, analyzing diagnostic report, offering clinical interpretation and decision-making, correlating genetic variants with disease phenotypes and determining appropriate treatment plans or follow-up actions, and providing any clinical advice.

[0041] Referencing FIG. 1A, one or more flowcell (s) 132 are provided to analytical system 110. A flowcell is a slide where the sequencing reactions occur. In some embodiments of the present invention, flowcells 132 are un-patterned flowcells. Each of the flowcells 132 has a planar surface without any wells or tiles. Numerous clusters are controllably generated on the surface of a flowcell and forms a logical unit for imaging and data processing. FIG. 2 illustrates a flowcell 132 and also illustrates a portion 208 of flowcell 132. Clusters of strands are formed on the portion 208, and similarly many other portions of flowcell 132. The synthesis process occurs on flowcell 132 and is described below in more detail.

[0042] Referencing back to FIG. 1A, optical sub-system 120, imaging sub-system 118, and sensors 116 are configured to perform various functions including providing an excitation light, guiding or directing the excitation light (e.g., using an optical waveguide or an optical fiber) , detecting light emitted from samples as a result of the excitation light, and converting photons of the detected light to electrical signals. For example, optical sub-system 120 includes an excitation optical module and one or more light sources, an optical waveguide, and / or one or more filters. In some embodiments, the excitation optical module and the light source (s) include laser (s) and / or light-emitting diode (LED) based light source (s) that generate and emit excitation light. The excitation light can have a single wavelength, a plurality of wavelengths, or a wavelength range (e.g., wavelengths between 200 nm to 1600 nm) . For instance, if analytical system 110 has a four-fluorescence channel configuration, optical sub-system 120 uses four different fluorescent lights having different wavelengths to excite four different corresponding fluorescent dyes (one for each of the bases A, G, T, C) .

[0043] In some embodiments, the excitation optical module can include further optical components such as beam shaping optics to form uniform collimated light. The excitation optical module can be optically coupled to an optical waveguide. For example, one or more of grating (s) , mirror (s) , prism (s) , diffuser (s) , and other optical coupling devices can be used to direct the excitation lights from the excitation optical module toward the optical waveguide.

[0044] In some embodiments, the optical waveguide can include three parts or three layers-a first light-guiding layer, a fluidic reaction channel, and a second light-guiding layer. The fluidic reaction channel may be bounded by the first light-guiding layer on one side (e.g., the top side) and bounded by the second light-guiding layer on the other side (e.g., the bottom side) . The fluidic reaction channel can be used to dispose flowcell (s) 132 bearing the biological sample. The fluidic reaction channel can be coupled to, for example, fluidic pipelines in fluidic sub-system 112 to receive and / or exchange liquid reagent. A fluidic reaction channel can be further coupled to other fluidic pipelines to deliver liquid reagent to the next fluidic reaction channel or a pump / waste container.

[0045] In some embodiments, fluorescent lights are delivered to flowcell (s) 132 without using an optical waveguide. For example, the fluorescent lights can be directed from the excitation optical module to flowcell (s) 132 using free-space optical components such as lens, grating (s) , lens (es) , mirror (s) , prism (s) , diffuser (s) , and other optical coupling devices.

[0046] As described above, fluidic sub-system 112 delivers various reagents to flowcell (s) 132 directly or through a fluidic reaction channel using fluidic pipelines. Fluidic sub-system 112 performs reagent exchange or mixing, and dispose waste generated from the liquid photonic system. One embodiment of fluidic sub-system 112 is a microfluidics sub-system (e.g., a digital microfluidic device) , which can process small amount of fluidics using channels measuring from tens to hundreds of micrometers. A microfluidics sub-system allows accelerating PCR processes, reducing reagent consumption, reaching high throughput assays, and integrating pre-or post-PCR assays on-chip. In some embodiments, fluidic sub-system 112 can include one or more reagents, one or more multi-port rotary valves, one or more pumps, and one or more waste containers.

[0047] The one or more reagents can be sequencing reagents in which sequencing samples are disposed. Different reagents can include the same or different chemicals or solutions (e.g., nucleic acid primers) for analyzing different samples. Biological samples that can be analyzed using the systems described in this application include, for example, fluorescent or fluorescently-labeled biomolecules such as nucleic acids, nucleotides, deoxyribonucleic acid (DNA) , ribonucleic acid (RNA) , peptide, or proteins. In some embodiments, fluorescent or fluorescently-labeled biomolecules include fluorescent markers capable of emitting light in one, two, three, or four wavelength ranges (e.g., emitting red and yellow lights) when the biomolecules are provided with an excitation light. The emitted light from the fluidic sub-system 112 can be further processed (e.g., filtered) before they reach the image sensors.

[0048] With reference to FIG. 1A, analytical system 110 further includes a control sub-system 114 and a power sub-system 122. Control sub-system 114 can be configured (e.g., via software) to control various aspects of the analytical system 110 of apparatus 100. For example, control sub-system 114 can include hardware and software to control the operation of optical sub-system 120 (e.g., control the excitation light generation) , fluidic sub-system 112 (e.g., control the multi-port rotary valve and pump, or control the digital microfluidic device to move droplets) , and power sub-system 122 (e.g., control the power supply of the various systems shown in FIG. 1A) . It is understood that various sub-systems of analytical system 110 of apparatus 100 shown in FIG. 1A are for illustration only. Analytical system 110 can include more or fewer sub-systems than shown in FIG. 1A. Moreover, one or more sub-systems included in analytical system 110 can be combined, integrated, or divided in any manner that is desired.

[0049] Referencing FIG. 1A, analytical system 110 of apparatus 100 includes an imaging sub-system 118. In some embodiments, imaging sub-system 118 has one or more image sensor (s) 116. Sensor (s) 116 detect photons of light emitted from the biological sample and convert the photons to electrical signals. Sensor (s) 116 are also referred to as image sensor (s) . An image sensor can be a semiconductor-based image sensor (e.g., silicon-based CMOS sensor) or a charge-coupled device (CCD) image sensor. A semiconductor-based image sensor can be a backside illumination (BSI) based image sensor or a front side illumination (FSI) based image sensor. In some embodiments, sensor (s) 116 may include one or more filters to remove scattered light or leakage light while allowing a substantial portion of the light emitted from the biological sample to pass. Filters can thus improve an image sensor’s signal-to-noise ratio. An example high-throughput image sensor 116 is described in more detail below in connection with FIG. 5.

[0050] The photons detected by sensor (s) 116 are processed by a signal processing circuitry 117 of imaging sub-system 118. An imaging sub-system 118 also includes a signal processing circuitry 117, which is electrically coupled to sensor (s) 116 to receive electrical signals generated by sensor (s) 116. In some embodiments, signal processing circuitry 117 can include one or more charge storage elements, an analog signal readout circuitry, and a digital control circuitry. In some embodiments, the charge storage elements receive or read out electrical signals generated in parallel based on substantially all photosensitive elements of an image sensor 116 (e.g., using a global shutter) ; and transmit the electrical signals to the analog signal read-out circuitry. The analog signal read-out circuitry may include, for example, an analog-to-digital converter (ADC) , which converts analog electrical signals to digital signals.

[0051] In some embodiments, the signal processing circuitry of imaging sub-system 118 converts analog electrical signals to digital signals, and transmits the digital signals to a data processing system to produce digital images such as fluorescence images 140 (also referred to as images 140 of fluorescence signals) . For example, the data processing system can perform various digital signal processing (DSP) algorithms (e.g., compression) for high-speed data processing. In some embodiments, at least a part of the data processing system can be integrated with the signal processing circuitry on a same semiconductor die or chip. In some embodiments, at least a part of the data processing system can be implemented separately from the signal processing circuitry (e.g., using a separate DSP chip or cloud computing resources) . Thus, data can be processed and shared efficiently to improve the performance of the sample analytical system 110. It is appreciated that at least a portion of the signal processing circuitry and data processing system in imaging sub-system 118 can be implemented using, for example, CMOS-based application specific integrated circuits (ASIC) , field programmable gate array (FPGA) , discrete IC technologies, and / or any other desired circuit techniques. One such example of circuitry for implementing the data processing system and / or the signal processing circuitry is shown in FIG. 9 and described below in greater detail.

[0052] It is further appreciated that power sub-system 122, optical sub-system 120, imaging sub-system 118, sensor (s) 116, signal processing circuitry 117, control sub-system 114, and fluidic sub-system 112 may be separate systems or components or may be integrated with one another. The combination of at least a portion of optical sub-system 120, imaging sub-system 118, and sensors 116 is sometimes also referred to as a liquid photonic system.

[0053] Referencing FIG. 1A, analytical system 110 of apparatus 100 provides fluorescence images 140 and / or other data to computing device (s) 103 to perform further processes including image preprocessing, cluster detection, feature extraction, basecalling, secondary data analysis, tertiary data analysis, and diagnostic reporting. Instructions for implementing one or more deep learning neural networks 102 reside on computing device (s) 103 in computer program product 104 which is stored in storage 105 and those instructions are executable by processor 106. One or more deep learning neural networks 102 can be used for performing various processes described below. When processor 106 is executing the instructions of computer program product 104, the instructions, or a portion thereof, are typically loaded into working memory 109 from which the instructions are readily accessed by processor 106. In one embodiment, computer program product 104 is stored in storage 105 or another non-transitory computer readable medium (which may include being distributed across media on different devices and different locations) . In alternative embodiments, the storage medium is transitory.

[0054] In one example, processor 106 in fact comprises multiple processors which may comprise additional working memories (additional processors and memories not individually illustrated) including a graphics processing unit (GPU) comprising at least thousands of arithmetic logic units supporting parallel computations on a large scale. Other embodiments comprise one or more specialized processing units comprising systolic arrays and / or other hardware arrangements that support efficient parallel processing (e.g., tensor processing units or TPUs, neural processing units or NPUs, programmable logic devices or PLDs) . In some embodiments, such specialized hardware works in conjunction with a CPU and / or GPU to carry out the various processing described herein. In some embodiments, such specialized hardware comprises application specific integrated circuits and the like (which may refer to a portion of an integrated circuit that is application-specific) , field programmable gate arrays and the like, or combinations thereof. In some embodiments, however, a processor such as processor 106 may be implemented as one or more general purpose processors (preferably having multiple cores) without necessarily departing from the spirit and scope of the present invention. As described below, the processor 106 can enable apparatus 100 to perform an integrated process of performing basecalling, analyzing the sequencing data, and providing the diagnostic outputs based on the biological sample.

[0055] User device 107 incudes a user interface 108 for displaying results of processing carried out by the one or more deep learning neural networks 102. The results may be, for example, a diagnostic report. In alternative embodiments, a neural network such as neural network 102, or a portion of it, may be stored in storage devices and executed by one or more processors residing on analytical system 110 and / or user device 107. Such alternatives do not depart from the scope of the invention.

[0056] In FIG. 1A, some or all of the sub-systems can be included in a single housing, thereby making apparatus 100 a standalone all-in-one machine, as shown in FIG. 1B. For example, the fluidic sub-system 112, the imaging sub-system 118, the processor 106, the user interface 108, and optionally one or more other sub-systems (e.g., the control sub-system 114, the optical sub-system 120, etc. ) can all be mounted inside a same housing. With this all-in-one machine, the user can perform sample analyzing, sequencing, basecalling, secondary data analysis, tertiary data analysis, and diagnosis reporting using the same apparatus, thereby significantly cutting down the processing time, complexity, and cost.

[0057] FIG. 1B illustrates an example apparatus 100 in another way. In some embodiments, apparatus 100 is an all-in-one machine that is configured to perform an integrated process based on an edge-based large language model (LLM) in an offline mode. As described above, currently, it is a great challenge to rapidly and accurately generate, based on a patient’s DNA sample and sequencing results, a clinical-level report including advice on disease diagnosis results and therapeutic regimens, because the whole data analysis procedure is dispersed, requiring multiple institutions, specialized equipment, and skilled personnel. In detail, the present diagnostic and reporting process needs collaboration among molecular biology laboratories, bioinformatics teams, and clinical departments. The whole process typically involves several laboratories. And these institutions are sometimes located in different locations. As a result, the processing time is very long, and the whole procedure often needs a day to several weeks to complete. In addition, the integration and organization of data, along with generating reports annotated with disease information, require expert personnel to review tremendous literature and databases, which can be very time-consuming. Furthermore, lower-level facilities, such as community clinics or rural hospitals, often lack experience in interpreting reports and making clinical decisions, necessitating intervention from higher-level hospitals to reach diagnostic conclusions and formulate treatment plans.

[0058] As shown in FIG. 1B, apparatus 100 in the present disclosure can provide an all-in-one machine configured to perform an integrated process including receiving the biological sample (e.g., a patient’s DNA sample) , performing the sequencing process, obtaining the images of fluorescence signals, basecalling to obtain sequencing data, performing secondary and tertiary data analyses, and providing a diagnostic report at the user interface 108 (or by another way like saving a file) . In some examples, the integrated process is performed by apparatus 100 in an offline mode without the need to remotely access a neural network model. For example, a trained and customized LLM model can be pre-installed to apparatus 100 for performing the secondary and / or tertiary data analyses. As such, the time required for performing the entire process can be shorted to a few hours to 1 day, rather than several days to weeks. Apparatus 100 can also be very cost-effective because it reduces or eliminates the requirements for involving multiple facilities, experts, and clinicians. In some examples, the locally stored LLM model can be updated when there is further training, further development of the model, and / or further customization based on user feedback. The update can be performed remotely (e.g., by downloading the updated LLM model from a remote server) or locally.

[0059] FIG. 1B also illustrates an example layer configuration of apparatus 100. As shown in FIG. 1B, apparatus 100 can be implemented to include three layers, i.e., the hardware layer 134, the middleware layer 136, and the application layer 138. Apparatus 100 receives a biological sample 101 and prepare the sample for analysis. The sample preparation process is described in more detail below using FIG. 2. The prepared biological sample 131 (e.g., a library) is analyzed by apparatus 100 (e.g., using the fluidic sub-system 112 and control sub-system 114 shown in FIG. 1A) . During analysis, fluorescent signals are generated and sensed by an imaging sub-system (e.g., sub-system 118) of apparatus 100. Images of fluorescent signals are thus obtained for performing further processes.

[0060] Next, the images of fluorescent signals are further analyzed by apparatus 100 based on its layer configuration shown in FIG. 1B. The hardware layer 134 include one or more types of processors, memory, and / or other hardware components. Hardware layer 134 provides computing power for performing various processes to analyze the images of the fluorescent signals and other intermediate data. As described above, in NGS, the throughput of generating the sequencing data can be very high. Therefore, apparatus 100 is configured to provide a hardware layer 134 that has sufficient computing power to process a vast amount of sequencing data generated in a short time period. For doing so, hardware layer 134 may include multiple GPUs, CPUs, TPUs, NPUs, FPGAs, ASICs, etc. and memory devices.

[0061] The second layer in the layer configuration is the middleware layer 136. In one example, middleware layer 136 illustrated in FIG. 1B includes neural network models for performing primary data analysis. For example, deep neural network models like transformer-based neural network models and encoder-decoder structures may be used for performing the primary data analysis (e.g., basecalling and quality filtering) . The primary data analysis such as basecalling is described in more detail using FIG. 6. The results of the primary data analysis include the sequencing data in a certain file format (e.g., a FastQ file) .

[0062] FIG. 1B further shows that the middleware layer 136 may also include neural network models for performing secondary data analysis and tertiary data analysis. As described above, secondary data analysis may include detecting variations or identifying pathogens. And the tertiary data analysis may include analyzing the secondary data analysis results, correlating genetic variants with disease phenotypes, and determining appropriate treatment plans or follow-up actions. In some embodiments, the secondary data analysis and / or the tertiary data analysis can be performed using an LLM stored locally at apparatus 100. Such an LLM may be referred to as an edge-based LLM (sometimes also called an end-based LLM) . The edge-based LLM refers to an artificial intelligence mode deployed directly at the data collection points or on user devices, compared to being deployed on remote cloud servers. In the example of FIG. 1B, the edge LLM is deployed directly at apparatus 100. In other examples, the edge LLM may be deployed at a local server or another computing device that is located near apparatus 100 (e.g., in the same room, building, or location) . Edge-based LLM moves the model operation and data storage closer to the source of the data. This allows apparatus 100 to process data locally. An edge-based LLM thus has advantages in terms of real time processing, data privacy protection, latency reduction, reduced bandwidth usage, and decreased dependence on network connectivity.

[0063] The analyses of sequencing data by using the middleware layer 136 may also use a vectorized biomedical knowledge database. The database comprises vectors representing nucleotide sequencing information, disease knowledge, and corresponding treatment protocols. The vectors are inputs to the LLM to enhance the background biomedical knowledge of the LLM.

[0064] In the layer configuration shown in FIG. 1B, the apparatus 100 may be implemented or configured with an application layer 138. The application layer 138 may include one or more applications such as rare disease detection, pathogenic detection, skin disease detection, or any other detections. These applications can be customized or configured in any way desired. For example, if the user of the apparatus 100 is specialized in skin diseases, apparatus 100 can be configured to have the skin disease detection application. If the user is specialized in cancer, apparatus 100 can be configured to have a cancer detection or diagnostic application, and so on. The applications included in the application layer 138 can take the outputs of the middleware layer 136 as inputs and generate, for example, disease diagnosis results, therapeutic regimens, recommended medicines or treatments, etc.

[0065] In some embodiments, the apparatus 100 can provide the outputs of the application layer 138 through a user interface 108 (e.g., an interactive graphical user interface, a file exchange interface, etc. ) . Via user interface 108, apparatus 100 can also receive user inputs and control the various sub-systems (e.g., fluidic sub-system, imaging sub-system, control sub-system, etc. ) and processes (e.g., primary data analysis, secondary data analysis, and tertiary data analysis) based on the user inputs. Thus apparatus 100 can be configured to operate in a closed-loop manner to analyze samples, perform sequencing data analysis, provide diagnostic outputs based on edge-based LLMs, receive user feedback, and optimize to provide more accurate or desirable diagnostic outputs. The feedback flow is described in more detail using FIG. 8.

[0066] FIG. 1C is a flowchart illustrating an example method 150 of performing an integrated process of sequencing of a biological sample and providing diagnostic outputs based on an edge-based large-language model (LLM) stored locally at the apparatus 100. Method 150 can be performed by the computer device (s) 103 of apparatus 100 using the three layers configuration shown in FIG. 1B. As shown in FIG. 1C, in step 152, apparatus 100 performs basecalling using the images 140 of fluorescent signals to obtain sequencing data 153. The sequencing data 153 can be stored in a certain file format (e.g., a FastQ file) . In step 154, the apparatus 100 can analyze the sequencing data based on the edge-based LLM 155 stored locally at the apparatus 100. The edge-based LLM 155 is operatable at the apparatus 100 in an offline mode. As described above, apparatus 100 can efficiently runs localized large models on the end-side using the cost-effective hardware (e.g., FPGAs / GPUs) , providing a superior cost-performance ratio. In some examples, apparatus 100 utilizes knowledge enhancement RAG (Retrieval-Augmented Generation) to load libraries with minimal investment for DNA application analysis, including pathogens, rare diseases, dermatological conditions, and more. This capability enables real-time, cost-effective analysis of DNA sequencing data and converts this data into client services for medical applications. In some embodiments, apparatus 100 generates diagnostic outputs 156 including disease diagnosis results, therapeutic regimens, recommended medicines or treatments. In some examples, the generation of medical reports by LLMs 155 may require a set of prompt templates. Next, the apparatus 100 can providing, via a user interface 108, the diagnostic outputs 156 based on analysis results of the sequencing data 153. Each of the steps shown in FIG. 1C is described in more detail below. Sequencing-by-Synthesis

[0067] As described above, before performing basecalling and performing data analyses, sequencing the biological sample is performed to obtain sequencing data. FIG. 2 illustrates an exemplary sequencing-by-synthesis process 200 using an analytical system (e.g., system 110) of apparatus 100 in accordance with an embodiment of the present invention. In step 1 of process 200, the analytical system heats up a biological sample to break apart the two strands of a DNA molecule. One of the single strands will be used as the DNA template strand. FIG. 2 illustrates such a DNA template strand 202, which can be a genomic DNA. A genomic DNA is the complete DNA sequence of an organism’s genome, which is the total genetic information of an organism. Template strand 202 may be a strand that includes a sequence of nucleotide bases (e.g., a long sequence having few hundreds or thousands of bases) . It is understood that there may be many such templated strands generated from using the polymerase chain reaction (PCR) techniques. It is further understood that there may also be other isolation and purification processes applied to the biological sample to obtain the DNA template strands.

[0068] In step 2 of process 200, the analytical system generates many DNA fragments from the DNA template strand 202. These DNA fragments, such as fragments 204A-204D shown in FIG. 2, are smaller pieces containing fewer number of nucleotide bases. These DNA fragments can thus be sequenced in a massively parallel manner to increase the throughput of the sequencing process in NGS. Step 3 of process 200 performs adapter ligation. Adapters are oligonucleotides with sequences that are complementary to the priming oligos disposed on the flowcell (s) . The ends of the nucleic acid fragments are ligated with adapters to obtain ligated DNA fragments (e.g., 206A-D) to enable the subsequent sequencing process.

[0069] The DNA fragmentation and adapter ligation steps prepare the nucleic acids to be sequenced. These prepared, ready-to-sequence samples are referred to as “libraries” because they represent a collection of molecules that are sequenceable. After the DNA fragmentation and adapter ligation steps, the analytical system generates a sequencing library representing a collection of DNA fragments with adapters attached to their ends. In some embodiments, prepared libraries are also quantified (and normalized if needed) so that an optimal concentration of molecules to be sequenced is loaded to the system. In some embodiments, other processes may also be performed in the library preparation process. Such processes may include size selection, library amplification by PCR, and / or target enrichment.

[0070] After library preparation, process 200 proceeds to step 4 for clonal amplification to generate clusters of DNA fragment strands (also referred to as template strands) . In this step, each of the DNA fragments is amplified or cloned to generate thousands of identical copies. These copies form clusters so that fluorescent signals of the clusters in the subsequent sequencing reaction are strong enough to be detected by the analytical system. One such amplification process is known as bridge amplification. In a bridge amplification process, a portion 208 of flowcell 132 is used and priming oligos are disposed on the flowcell 132. As described in more detail below, flowcell 132 is an un-patterned flowcell and thus does not include any wells (e.g., nanoscale wells) like conventional flowcells. Thus, the surface of the flowcell 132 is planar. As a result, the surface properties of the flowcell 132 can be uniform or consistent without variation or disruption. Different portions of flowcell 132 can be substantially the same and the planar surface of flowcell 132 enables intelligent control of the cluster growing, as described in more detail below with FIGs. 3 and 4.

[0071] Continuing with the process 200 shown in FIG. 2, each DNA fragment in the library anneals to the primer oligo disposed on a portion 208 of flowcell 132 via the adapters attached to the DNA fragment. The complementary strand of a ligated DNA fragment is then synthesized. The complementary strand folds over and anneals with the other type of primer oligo disposed on the portion 208 of flowcell 132. A double-stranded bridge is thus formed after synthesis of the complementary strand.

[0072] The double-stranded bridge is denatured, forming two single strands attached to the portion 208 of flowcell 132. This process of bridge amplification repeats many times. The double-stranded clonal bridges are denatured, the reverse strands are removed, and the forward strands remain as clusters for subsequent sequencing. Two such clusters of strands are shown as clusters 214 and 216 in FIG. 2. Many clusters having different DNA fragments can be attached to the flowcell 132 at a same portion 208 or different portions. For example, cluster 214 may be a cluster of ligated fragmented DNA 206A disposed on portion 208; and cluster 216 may be a cluster of ligated fragmented DNA 206B also disposed on portion 208. The subsequent sequencing can be performed in parallel to some or all of these different clusters disposed on a portion and in turn, some or all the clusters disposed on many portions of the flowcell (s) . The sequencing process can thus be massively parallel.

[0073] Referencing FIG. 2, after the clonal amplification in step 4, process 200 proceeds to step 5, where the clusters are sequenced by synthesis (SBS) . In this SBS step, nucleotides are incorporated by a DNA polymerase into the complementary DNA strands of the clonal clusters of the DNA fragments one base at a time in each synthesis cycle. For example, as shown in FIG. 2, if cycle 1 is a beginning cycle, a first complementary nucleotide base is incorporated to the complementary DNA strand of each strand in cluster 214. FIG. 2 only shows one strand in cluster 214 for simplicity. But it is understood that similar processes can occur to some or all other strands of cluster 214, some or all other clusters on portion 208 of flowcell 132, some or all other portions, and some or all other flowcells. This synthesis process repeats in cycle 2, where a second complementary nucleotide base is incorporated to the complementary DNA strand. This synthesis process then repeats in cycles 3, 4, and so on, until complementary nucleotide bases are incorporated for all bases in the template strand 206A or until a predetermined number of cycles is reached. Thus, if the template strand 206A has “n” nucleotide bases, there may be “n” cycles or a predetermined number of cycles (less than “n” ) for the entire sequencing-by-synthesis process. The complementary strand 207A is at least partially completed after all the synthesis cycles. In some embodiments, this synthesis process can be performed for some or all strands, clusters, portions, and flowcells in parallel.

[0074] Step 6 of process 200 is an imaging step that can be performed after step 5 or in parallel with step 5. As one example, a flowcell can be imaged after the sequencing-by-synthesis process is completed for the flowcell. As another example, a flowcell can be imaged while the sequencing-by-synthesis process is being performed on another flowcell, thereby increasing the throughput. Referencing FIG. 2, in each cycle, the analytical system captures one or more images of the portion 208 (e.g., images 228A-D) of flowcell 132. The images represent the fluorescent signals detected in the particular cycle for all the clusters disposed on the portion 208 of flowcell 132. In some embodiments, the analytical system can have a four-channel configuration, where four different fluorescent dyes are used for identifying the four nucleotide bases. For example, the four fluorescence channels use different types of dyes for generating fluorescent signals having different spectral wavelengths. Different dyes may each bind with a different target and produce signals with a different fluorescence color or spectrum. Examples of the different dyes may include a Carboxyfluorescein (FAM) based dye that produces signals having a blue fluorescence color, a Hexachloro-fluorescein (HEX) based dye that produces signals having a green fluorescence color, a 6-carboxy-X-rhodamine (ROX) based dye that produces signals having a red fluorescence color, a Tetramethylrhodamine (TAMRA) based dye that produces signals having a yellow fluorescence color.

[0075] In a four-channel configuration, the analytical system captures an image of the same portion of the flowcell for each channel. Therefore, for each portion, the analytical system produces four images in each cycle. This imaging process can be performed with respect to some or all the portions and flowcells, producing a massive number of images in each cycle. These images represent the fluorescent signals detected in that particular cycle for all the clusters disposed on the tile. The images captured for all cycles can be used for basecalling to determine the sequences of the DNA fragments. A sequence of an DNA fragment includes an ordered combination of nucleotide bases having four different types, i.e., Adenine (A) , Thymine (T) , Cytosine (C) , and Guanine (G) . The sequences of multiple DNA fragments can be integrated or combined to generate the sequence of the original genomic DNA strand. Embodiments of this invention described below can process the massive numbers of images in an efficient way using high-throughput semiconductor-based image sensors and perform basecalling using improved architectures of deep learning neural networks. The basecalling process according to the embodiments of this invention thus has a faster speed and a lower error rate. While the above descriptions use DNA as an example, it is understood that the same or similar processes can be used for other nucleic acid, such as RNA and artificial nucleic acid. Intelligent Clustering Technologies

[0076] As described above, in step 4 of process 200, apparatus 100 performs clonal amplification to generate clusters of DNA fragment strands, or clusters in short, on flowcells 132. Conventionally, a flowcell used for the sequencing process is a patterned flowcell 310 as shown in FIG. 3. Flowcell 310 has many (e.g., billions to tens of billions) nanoscale wells precisely situated across the entire expanse of the flowcell’s surface. FIG. 3 shows an enlarged view 312 of some of these nanoscale wells 311. The clusters 316 are formed at the location of these nanoscale wells 311 but are not in other portions of the flowcell 310 (e.g., they do not form between nanoscale wells) . Patterned flowcell 310, as shown, is a structured arrangement to facilitate a consistent and even dispersion of sequencing clusters 316. But the clusters 316 are confined to be formed within the nanoscale wells 311 only. Therefore, the patterned flowcell 310 has a fixed density of clusters, depending on the arrangement of the nanoscale wells. This limits the density of the clusters that can be formed, and thereby limiting the data throughput capacity. In other words, the data throughput capacity of a system using patterned flowcells are predetermined and fixed. Such a system has no flexibility to adjust aspects of the sequencing process on demand.

[0077] Moreover, the intricate and complex fabrication process of patterned flowcells and to create these patterned surfaces contributes to a higher manufacturing cost for each flowcell, which is a factor that can significantly impact the overall expense of DNA sequencing process. In addition, the nanoscale wells of patterned flowcells can have adverse effects on the biochemical reactions, because the nanoscale wells disrupts the surface properties of the flowcells. In a patterned flowcell, surface properties may vary. For example, the surface chemical distribution, concentration, surface roughness, surface tension, etc. may vary from the edge of the nanoscale well to the center of the nanoscale well, making biochemical reactions uneven and making the cluster formation process difficult to control.

[0078] With reference to FIG. 4A, in some embodiments of the present invention, the apparatus (e.g., apparatus 100) for performing an integrated process of sequencing of a biological sample and providing diagnostic outputs can be configured to perform the sequencing process using un-patterned flowcells 432.  The apparatus can further be configured to control the cluster formation process. Unlike flowcell 310, an un-patterned flowcell 432 has no wells. For example, as shown in FIG. 4A, an un-patterned flowcell 432 has a planar surface and has no nanoscale wells (or any wells) for accommodating droplets of the biological samples. The droplets of the biological samples are distributed to different portions of the planar surface, on which the clusters of strands are formed.

[0079] Specifically, in one embodiment shown in FIGs. 4A and 4C, the fluidic sub-system 112 of apparatus 100 receives (step 452) the biological sample 101. The fluidic sub-system 112 can prepare (step 454) the biological sample to form clusters of strands on a plurality of un-patterned flowcells having planar surfaces. The preparation of the biological sample is described above using FIG. 2, steps 1-3. Once the biological sample is prepared, the fluidic sub-system 112 can perform clonal amplification to generate clusters 416 of DNA fragment strands on the planar surface of the un-patterned flowcells 432.

[0080] Because the un-patterned flowcells 432 has planar surfaces and no wells, the cluster formation can be much more flexible and controllable. For example, because the cluster formation is not confined in wells, the density of the clusters 416 can be increased; the size of the cluster 416 can be varied; and the spatial arrangements of the clusters 416 can be controlled. As shown in FIG. 4A and FIG. 4C, the control sub-system 114 can be configured to control (step 456) , for each of the plurality of flowcells 432, one or more aspects of the cluster formation. Some of these aspects include spatial arrangement (458A) , density of the clusters of strands (458B) , and sizes of the clusters of the strands (458C) .

[0081] In some examples, the control sub-system 114 controls these aspects of the cluster formation process based on user input 459. The user input 459 may be received at the user interface 108 of a user device 107. In some examples, the control sub-system 114 may be configured to control the aspects of the cluster formation process automatically using a trained neural network model. The neural network model can adjust the control of the cluster formation process based on the previous one or more sequencing results or diagnostic outputs. For instance, if more data points or higher throughput are needed based on the previous sequencing results, the neural network model of control sub-system 114 can increase the density of the clusters 416. In another example, based on the user input 459, the sequencing data can be filtered to remove the low-quality data that does not meet the user requirements or sequencing applications’ requirements.  The user input 459 may also be used to fine tune the sequencing process described above in connection with FIG. 2.

[0082] FIG. 4B is a block diagram illustrating various aspects of clusters controllably formed on un-patterned flowcells, in accordance with an embodiment of the present invention. In FIG. 4B, on un-patterned flowcell 432A, the spatial arrangement of clusters 416A are controlled such that spatial utilization of the flowcell can be optimized (e.g., maximized) . As described above, when clusters form on a patterned flowcell, their spatial locations are fixed such that they can only be formed at the location of the nanoscale wells. The un-patterned flowcell removes such a limitation, and as shown in FIG. 4B, clusters 416A can be formed at any portion of the un-patterned flowcell 432A, making the spatial utilization much more improved. FIG. 4B illustrates, for example, several clusters 416A are formed close to each other and do not need to be confined in certain locations. This adaptability enables the flowcell to operate at peak spatial utilization, judiciously utilizing its capacity to the fullest extent possible. In addition, because the un-patterned flowcell has a planar surface, the surface properties can be uniform and constant across the surface with no variations or with minimum variations. As a result, the biochemical reactions for forming the clusters are not affected.

[0083] FIG. 4B also illustrates another un-patterned flowcell 432B, on which clusters 416B are formed. The density of the clusters 416B is controlled to be high, thereby increasing the number of clusters on each flowcell. For instance, density of the cluster may be increased from about 22,000 clusters per mm2 when using a patterned flowcell to 300,000 clusters per mm2. This in turn may also increase the throughput of the sequencing process and the overall sequencing performance. FIG. 4B further illustrates another un-patterned flowcell 432C, on which clusters 416C are formed. Some of clusters 416C may be controlled to have a smaller size, compared to those clusters that are formed on a patterned flowcell. The smaller size of the clusters may also improve the effective utilization of the flowcell to realize high-throughput sequencing process. It is understood that FIGs. 4A-4B provide several examples of controlling one or more aspects of the clusters during a sequencing process, and clusters can be controlled in other ways (e.g., clusters that have larger sizes) not limited to those described above.

[0084] As described above, by using un-patterned flowcells and controlling the cluster formation, the clusters formed on the flowcells can have a higher density and tighter spatial arrangement, and / or smaller sizes. Controlling the cluster formation based on demand or sequencing requirements is referred to as intelligent clustering technology. One challenge associated with the intelligent clustering technology is that it imposes higher requirements on the downstream imaging process and basecalling process. As described above, images of the portions of a flowcell are captured during an imaging process. The images represent the fluorescent signals detected in a particular cycle for all the clusters disposed on the portions of the flowcell. Traditionally, images are captured by optical systems (e.g., microscopes) that have a limited resolution. Therefore, increasing the density of the clusters and / or reducing the sizes of the clusters may render the images captured difficult to analyze for basecalling. For example, certain fluorescent signals may not be clearly distinguishable from the adjacent fluorescent signals. In turn, the basecalling results may be negatively affected and the sequencing results may not be accurate. High-throughput semiconductor-based imaging system

[0085] To resolve the aforementioned difficulties associated with using the un-patterned flowcells, the present disclosure provides the image sub-system 118 of apparatus 100, which comprises a plurality of semiconductor-based image sensors (e.g., CMOS sensors) configured to perform high-throughput imaging sensing with high resolution. FIG. 5 is a block diagram illustrating an example image sub-system 118 used to sense the fluorescent signals and obtain images of the fluorescent signals in any particular cycles for one or more clusters. As shown in FIG. 5, imaging sub-system 118 includes many image sensors 522 and is scalable (e.g., can be scaled to include as many image sensors 522 as desired) . An image sensor is a sensor that detects photons, generate electrical signals (also referred to as photoelectrons) based on the detected photons, and transmit the electrical signals for further signal processing. In an image sensing system, photons can be generated as a result of fluorescence or chemiluminescence emissions from biological or chemical samples being analyzed. The photons are then collected and detected by photosensitive elements (e.g., pixels) included in an image sensor. Photosensitive elements can include, for example, photodiodes (e.g., silicon-based photodiodes) for detecting photons and generating electrical signals based on detected photons. In some embodiments, photosensitive elements may also include amplifiers (e.g., avalanche amplification) . The electrical signals generated by an image sensor can represent various photon information including the number of photons collected, the position of photons, and / or the intensity of photons. An image sensor described in this disclosure is not limited to a sensor that transmits electrical signals or information for generating an image. An image sensor used for analyzing biological or chemical samples (e.g., nucleotide acid sequencing applications, polymerase chain reaction applications) can include sensors that detect photons and transmit electrical signals for any type of signal processing with or without generating an image.

[0086] With reference to FIG. 5, imaging sub-system 118 includes wafer-level packaged semiconductor dies and their arrangements on a single semiconductor wafer. In some embodiments, wafer-level packaging of the semiconductor dies can include one or more of: forming through-silicon vias (TSV) , depositing redistribution layers, depositing passivation layers, forming electrically-conductive spheres, and disposing solder mask layers. In this disclosure, wafer-level packaged semiconductor dies are sometimes also referred to as packaged semiconductor dies or TSV-packaged semiconductor dies. In FIG. 5, each individual block (e.g., block representing image sensors 522A-522F) can represent a packaged semiconductor die of a semiconductor wafer. A semiconductor die is a unit or a single block of semiconductor material on which integrated circuits or other devices (e.g., sensors) are fabricated. For example, an image sensor having a plurality of photosensitive elements (e.g., pixels) can be fabricated on each semiconductor die represented by an individual block associated with image sensors 522A-522F as shown in FIG. 5.

[0087] In some embodiments, each image sensor can be fabricated on an individual semiconductor die. Thus, image sensor 522A is fabricated on a semiconductor die; image sensor 522B is fabricated on another neighboring die, and so forth. Each block associated with an image sensor 522 in FIG. 5 represents a separate semiconductor die. An image sensor may have a pre-configured or a pre-determined throughput capacity represented by a pixel array size. For example, an image senor may have a pixel array size of 8 megapixels, 16 megapixels, 32 megapixels, etc. Typically, for a given semiconductor process (e.g., a 45nm CMOS image sensor process) , a larger pixel array size requires more photosensitive elements such as more photodiodes. As a result, an image sensor with higher throughput capacity may require a larger physical area of a semiconductor die. In some embodiments, rather than increasing the area of a single semiconductor die, a high throughput sensing system or throughput-scalable sensing system can also be obtained based on group dicing of multiple packaged semiconductor dies.

[0088] FIG. 5 illustrates an exemplary a high throughput imaging sub-system 118 based on group dicing of packaged semiconductor dies from a semiconductor wafer as groups. As shown in FIG. 5, multiple packaged semiconductor dies can be diced as a group 520, instead of individually, from a semiconductor wafer. Dicing, sometimes also referred to as wafer dicing, is a process by which packaged semiconductor dies are physically separated from a semiconductor wafer or a wafer-level packaged semiconductor wafer. A dicing process may include scribing, breaking, mechanical sawing, and / or laser cutting. Dicing is typically performed at or near dicing streets (e.g., dicing streets 535) between the packaged semiconductor dies. The dicing streets can be, for example, 80 micrometers (um) wide.

[0089] In FIG. 5, an exemplary dicing group 520 is illustrated. The exemplary dicing group 520 may include a plurality of individual packaged semiconductor dies (e.g., 8, 16, 32, 64, etc. ) , on which image sensors 522 are fabricated. An image sensor with a particular throughput capacity may be fabricated on each packaged semiconductor die in the group 520. Therefore, the blocks shown in FIG. 5 can also represent a group of image sensors disposed on the corresponding packaged semiconductor dies.

[0090] FIG. 5 illustrates that imaging sub-system 118 is a throughput-scalable image sensing system obtained based on dicing multiple packaged semiconductor dies as a group from a semiconductor wafer or a wafer-level packaged semiconductor wafer. In FIG. 5, an image sensor is pre-fabricated or disposed on each packaged semiconductor die. For example, one image sensor 522A can be fabricated or disposed on a packaged semiconductor die. The image sensor 522A may include, for example, a plurality of photosensitive elements 524, a plurality of electrically-conductive layers (not shown in FIG. 5) , a plurality of electrically-conductive pads 526, and a semiconductor (e.g., silicon) substrate 528. The substrate 528 is a common substrate that is shared among all semiconductor dies of the image sensors 522.

[0091] As shown in FIG. 5, based on a group dicing plan, the packaged semiconductor dies of the semiconductor wafer can be diced in groups. The example illustrated in FIG. 5 shows that a group of six packaged semiconductor dies of image sensors 522A-522F are separated from the semiconductor wafer by, for example, laser cutting along the dicing streets at dicing streets 530, while keep the semiconductor dies of image sensors 522A-522F together. In other words, the laser cutting is only to physically separate the group 520 as a whole from a wafer, without physically separating the individual packaged semiconductor dies of image sensors 522A-522F from one another. Dicing streets 530 represent the perimeter of the group of dies of image sensors 522A-522F. And therefore, in group dicing, the laser cutting is performed along the perimeter of the group of dies of image sensors 522A-522F, but not between the dies. Each of packaged semiconductor dies of image sensors 522A-522F can be pre-fabricated or disposed with an image sensor having a particular throughout capacity (e.g., pixel array size) . The six image sensors pre-fabricated or disposed on packaged semiconductor dies of image sensors 522A-522F can thus form an image sensing system 220 that has six-times throughput capacity than each individual image sensor. In general, if each image sensor in the image sub-system 118 has a pixel array size of M megapixels and there are N number of image sensors in the image sensing system, the total pixel array size of the image sensing system is then M x N. In the example illustrated in FIG. 5, if each image sensor 522 has a pixel array size of 64 megapixels, and a group of six packaged semiconductor dies of image sensors 522A-522F form image sub-system 118, image sub-system 118 can be scaled to have a pixel array size of 384 megapixels.

[0092] While FIG. 5 illustrates that image sub-system 118 include six image sensors 522A-522F disposed on six packaged semiconductor dies, it is appreciated that the number of image sensors in a particular image sensing system can be determined or preconfigured to any desired number satisfying a throughput scaling requirement. For example, if a particular image sensing system used for a nucleotide acid sequencing application requires a total pixel array size of 1000 megapixels (or 1 Gigapixels) , and if each image sensor has a pixel array size of 64 megapixels, the number of image sensors required for such an image sensing system would be about 16 (e.g., 1000 / 64) . Correspondingly, 16 packaged semiconductor dies can be diced from the semiconductor wafer as a group (i.e., without separating the 16 dies from one another) . Accordingly, based on the throughput scaling requirement for the image sensing system and based on a throughput capacity of each image sensor, the number of image sensors in the image sensing system can be readily determined. As a result, the throughput capacity of the image sensing system can easily be scalable based on requirements of the specific applications (e.g., a DNA sequencing application, a PCR application) of the image sensing system. Such a throughput-scalable system does not require complex and costly redesign of the image sensor itself (e.g., redesign to add more photosensitive elements in a single semiconductor die) . The throughput-scalable imaging sub-system can be used to detect fluorescent signals from high-density clusters as described above. Furthermore, as the size of a pixel of the image sensor 522 becomes smaller, the resolution of the imaging sub-system 118 can be improved too.

[0093] As shown in FIG. 5, in one example of an imaging sub-system 118, multiple image sensors 522 are disposed on packaged semiconductor dies. As illustrated in FIG. 5, photosensitive elements of the multiple image sensors are not physically continuous or connected with one another. For example, there is a gap between photosensitive elements 524 of the image sensor 522A disposed on a corresponding die and photosensitive elements 534 of the image sensor 522B disposed on another die. While the two dies physically share a common substrate 528, the image sensors 522A and 522B are not continuously placed. Between the photosensitive elements of different image sensors, other device structures or components (e.g., pads 526) and dicing streets (e.g., dicing street 535 between dies of image sensors 522A and 522B) may exist. As a result, an image generated by multiple image sensors 522A-522F disposed on separate packaged semiconductor dies may not be continues or may have one or more image gaps between different portions of the image. Image gaps can be blank or dark areas between different portions of an image due to the lack of photon sensing between the photosensitive elements. Such image gaps may not be acceptable for certain imaging applications that require a continuous image to be provided. Such applications may include, for example, traditional photo-capturing applications (e.g., taking portrait photos, picturing a real-world object, etc. ) , surveillance camera applications, or security monitoring applications.

[0094] Further, for those applications in which image gaps are not acceptable or tolerable, if a raw image generated by multiple image sensors is not continues or has image gaps, significant image processing efforts may be required to remove or mitigate the image gaps. For example, post-capturing image processing may be applied to stitch portions of the images together to provide an acceptable image without image gaps. Thus, an image sensing system with multiple image sensors that have discretely-positioned photosensitive elements (e.g., elements that are not physically continuous or connected with one another) may not be easily designed or implemented for certain imaging applications. In contrast, such an imaging system may not have or may have minimum impact on performance of a biological or chemical sample analysis application such as a nucleotide acid sequencing application. For many biological or chemical sample analysis applications, the image sensors are used to count photons emitted from the samples and generate an image base on the photons. The image can be allowed to have image gaps, because the analysis results can be derived based on the information related to photon detections (e.g., the intensity of photons, position of photons, pattern of photons, etc. ) . The derivation of the analysis results does not require the image to be continuous or without image gaps. Therefore, a high-throughput image sensing system (e.g., imaging sub-system 118) including multiple image sensors obtained based on group dicing technologies can be readily used for many biological or chemical sample analysis application or any other photon counting based applications, without requiring any mitigation effort to remove the image gaps caused by the discretely-positioned photosensitive elements. Further details of such a high-throughput image sensing system can be found in U.S. non-provisional application No. 17 / 246,487 (now patent No. US 11,175,219) , entitled “THROUGHPUT-SCALABLE ANALYTICAL SYSTEM USING SINGLE MOLECULE ANALYSIS SENSORS, ” the content of which is incorporated herein by reference in its entirety for all purposes. Deep-learning neural network based Basecalling

[0095] As described above, in the sequencing process shown in FIG. 2, after the imaging sub-system (e.g., system 118) obtains images of the clusters, the basecalling process can be performed using deep-learning neural network models. FIGs. 6A-6C illustrate such a deep-learning based basecalling process. In particular, FIG. 6A is a flowchart illustrating a method 600 of performing basecalling based on a self-attention transformer-based neural network. FIG. 6B is a block diagram illustrating an encoder-decoder based convolutional neural network (CNN) , in accordance with an embodiment of the present invention. FIG. 6C is a block diagram illustrating a transformer-based neural network, in accordance with an embodiment of the present invention.

[0096] With reference to FIGs. 1A and 6A, the computing device (s) 103 obtains the images 140 of fluorescent signals from imaging sub-system 118 and performs deep-learning neural network based basecalling. In some examples, the images 140 may be preprocessed including light correction, image registration, image normalization, image enhancement, etc. In some embodiments, the images 140 can be partitioned into a grid of patches (step 602 in FIG. 6A) . As shown in FIG. 6C, the images 140 can be partitioned to a grid of patches 644. In this process, the input multi-channel images 140 may be segmented into a grid of patches, each of a predetermined size. The grid of patches 644 are provided to a transformer-based neural network 640 as shown in FIG. 6C.

[0097] In step 604 of method 600 shown in FIG. 6A, the computing device (s) 103 can perform linear projection of the grid of patches to obtain position embeddings. FIG. 6C illustrates such linear projections 645, in which each patch is firstly transformed into a one-dimensional vector (e.g., a sequence) . From these flattened vectors, lower-dimensional linear embeddings are constructed, capturing the essence of each patch while reducing dimensionality. The results of the linear projections 645 are position embeddings 646 (also referred to as feature embedding vectors) .

[0098] In step 606 of method 600 shown in FIG. 6A, the computing device (s) 103 can perform a position encoding addition process to add class tokens to the position embeddings. Positional embeddings are thus integrated to impart information about the relative positions of the patches within the original image. The asterisk (*) shown in position embedding 646 of FIG. 6C represents the class token that starts as a randomly initiated vector that is added to the sequence of the image patches. During the forward pass, the self-attention mechanism allows the class token to interact with all other tokens in the sequence.

[0099] In step 608 of method 600 shown in FIG. 6A, the computing device (s) 103 can encode, by an encoder, outputs of the position encoding addition process using a self-attention transformer based neural network to obtain encoder outputs. FIG. 6C illustrates such a transformer-based encoder 647. The transformer-based encoder 647 mainly relies on the self-attention mechanism. This mechanism computes weightings for every pixel within the image, quantifying its relevance in relation to all other pixels. Moreover, the multi-head attention (MLP) component augments this process by enabling the model to concurrently focus on various segments of the input sequence. This diversification of attention facilitates a more holistic understanding of the input data, as the model can consider multiple aspects or features simultaneously, enhancing its ability to capture nuanced patterns and dependencies across different parts of the input. Concurrently, the output from the self-attention layer is subjected by the multi-layer perceptron (MLP) to add nonlinearity and enrich the representation with more complex features.

[0100] FIG. 6C illustrates an example of the transformer encoder 647 in the self-attention transformer-based neural network 640. In some examples, encoder 647 includes a multi-head attention layer 653. The multi-head attention layer 653 determines multiple attention vectors per element of the position encoded vectors provided by the position embedding 646 and takes a weighted average to compute a final attention vector for each element of the position encoded vectors. The final attention vectors capture the contextual relationship between elements of the position encoded vectors. The transformer encoder 647 also includes one or more normalization layers 652 and 655. The normalization layers control the gradient scales. In some embodiments, the normalization layer 652 is positioned before the multi-head attention layer 653, as illustrated in FIG. 6C. In some embodiments, the normalization layer 652 can be positioned after the multi-head attention layer 653. Similarly, normalization layer 655 can be positioned before or after multiplayer perceptron layer 656 as well. A normalization layer standardizes the inputs to the next layer, which has the effect of stabilizing the network’s learning process and reducing the number of training iterations required to train the deep learning network. Normalization layers 652 and 655 can perform batch normalization and / or layer normalization.

[0101] FIG. 6C also illustrates that transformer-based encoder 647 includes a multilayer perceptron (MLP) 656. MLP 656 can be a type of feedforward neural network. An MLP has layers of nodes including: an input layer, one or more hidden layers, and an output layer. Except for the input nodes, each node in an MLP is a neuron that uses nonlinear activation function. MLP 656 is applied to every normalized attention vector. MLP 656 can transform the normalized attention vectors to a form that is acceptable by the next encoder or decoder in the self-attention transformer-based neural network . In the example shown in FIG. 6C, one encoder is used. Thus, in FIG. 6C, the output of MLP 656 can be the basis of the encoder output vectors 657.

[0102] Encoder output vectors 657 are then provided to a decoder 648, which includes another MLP. In another example shown in FIG. 6B, a stacked encoder structure having two encoders 624 and 625 is used. Thus, the output vectors from encoder 624 are provided to the next encoder 625 as input vectors. And the encoder output vectors 626 in the self-attention transformer-based neural network 620 are provided by the second encoder 625.

[0103] With reference back to FIG. 6C, the encoder output vectors 657 are provided to the decoder 648, which include a MLP head too. The MLP head serves as the final component of the network 640 that performs the classification task. After image patches has passed through the transformer encoder layers in encoder 647, which have learned to represent the patches in a semantically meaningful way, the MLP head of decoder 648 takes the output of the transformer encoders and maps it to the desired number of base classes. In particular, the decoder 648 outputs probability distributions representing probabilities of bases (A, T, C, or G) . And a particular base is predicted using the highest probability for that base. The process can be repeated many times to produce basecalling results, which include the sequences of DNA fragments. FIGs. 6A and 6C shows that the computing device (s) 103 of apparatus 100 can classify (step 610) the encoder output to obtain the sequencing data 153. The sequencing data 153 includes base classes of the sequences of the biological sample under analysis. In one example, the sequencing data 153 can be in a certain file format (e.g., FastQ file) .

[0104] While FIG. 6C illustrates that network 640 has one encoder and one decoder, multiple stacked encoders and stacked decoders may be used. For example, FIG. 6B illustrates that two encoders 624 and 625 are stacked together to produce the encoder output vectors 626. The encoder output vectors 626 may be forwarded to two decoders 627 for decoding (e.g., classification) . The network 620 may also include a linear and softmax layer 628. The layer 628 helps to transforms its input to probability distributions representing probabilities of bases.

[0105] Self-attention transformer-based neural network 620 or 640 can be trained with training data. A trained transformer neural network starts by generating input embedding vectors representing features of clusters of fluorescent signals of all images captured in “n” synthesis cycles. Using self-attention mechanism implemented by multi-head attention, network 620 or 640 aggregates information from all of the features represented by input embedding vectors to generate encoder output vectors. Each encoder output vector is thus informed by the entire context. The generation of the multiple encoder output vectors are performed in parallel for all features represented by input embedding vectors. A decoder in network 620 or 640 attends not only to the encoder output vectors but also other previously generated embedding vectors. Network 620 or 640 can be trained with sufficient data to provide a high accuracy. The training of network 620 or 640 can also be repeated or updated over time to maintain or provide even higher accuracy.

[0106] It is understood that FIGs. 6B and 6C merely provide examples of the self-attention transformer network. Other variations of the network can also be configured. Such networks are described in more detail in U.S. non-provisional application No. 17 / 681,672, now U.S. Patent No. 11, 580,641, entitled “DEEP LEARNING BASED METHODS AND SYSTEMS FOR NUCLEIC ACID SEQUENCING, ” the content of which is hereby incorporated by reference in its entirety for all purposes.

[0107] Embodiments of the present invention improve and optimize the speed and accuracy of cluster detection and basecalling, particularly when non-patterned flowcells and high-throughput scalable imaging systems are used. In such situations, the clusters have a high density and the high-throughput scalable imaging system (e.g., imaging sub-system 118) generates a vast amount of image data for processing. Attention-based deep learning models are better suited to parallel processing. Specifically, attention-based deep learning models generate attention vectors that are independent from one another. Therefore, they can be processed in parallel, significantly improving the basecalling speed. The amount of image data generated during a synthesis process is directly proportional to the length of a sequence. Thus, attention-based deep learning models are particularly efficient in processing long sequences because of their capabilities of processing input data in parallel. Due to the attention mechanism, an attention-based deep learning model also better explains the relation between the input data elements. As a result, attention-based deep learning models also improve the basecalling accuracy and / or reduce the error rate. Various embodiments of the basecalling algorithms described herein can be applied for nucleic acid sequencing (e.g., DNA sequencing, RNA sequencing, artificial nucleic acid sequencing) and / or protein sequencing. Quality Filtering

[0108] In DNA sequencing, for example, a statistical model such as the AYB (All Your Base) model may be used to produce more accurate basecalling results. The neural network-based basecalling models such as RNN base models, CNN based models, and a combination of transformer and CNN based model may also be used. A self-attention transformer-based model is described above. Deep learning-based basecalling approaches can provide matching or improved performance compared to traditional basecalling approaches, while enable high-throughput basecalling processes. During the basecalling process, it is desired to achieve high sequencing quality such as high overall sequencing accuracy. Improving the DNA sequencing accuracy can be achieved from several aspects including, for example, improving the efficiency of biochemical reagents, improving the quality of the optical system, and / or improving the basecalling algorithms.

[0109] FIG. 7A is a flowchart illustrating a method 700 of performing quality filtering of the sequencing data and further sequencing data analysis, in accordance with an embodiment of the present invention. As described above, the basecalling process generates the sequencing data, sometimes also referred to as the raw sequencing data. As shown in FIGs. 7A and 7B, the sequencing data 701 may be stored as a FastQ file. In some examples, the sequencing data 153 may be further processed to improve the quality. FIGs. 7A and 7B illustrate that a computing device (e.g., device 103) can perform quality filtering (step 710) of the raw sequencing data to obtain quality-improved sequences. The quality-improved sequences having higher quality scores than raw sequences included in the sequencing data.

[0110] In some embodiments, the basecalling algorithms may be improved by filtering an initial group of sequences to obtain a high-quality group of sequences. Existing techniques for filtering sequences to exclude low quality sequencing data use chastity filtering, which computes the intensity ratio between the DNA signal channel having the highest intensity and the channel having the second highest intensity. The intensity ratio represents the cluster quality. The chastity filtering method requires corrections of the data crosstalk and cluster phasing. As a result, the quality of the crosstalk and cluster phasing corrections limit the filtering performance. Furthermore, there is usually a limit to obtain a higher quality sequencing result with a single deep learning network.

[0111] In some embodiments, a hierarchical processing network structure for basecalling can be used to perform quality filtering. The network structure uses one or more network blocks for obtaining high quality sequences. Each network block comprises a base network and a sequence filter. The base network generates one or more sequencing quality indicators such as sequence quality filtering via network (SQFN) passing indices, a sequence dataset quality index, basecalling confidence level scores, etc. The sequencing quality indicators can represent the accuracy of basecalling of individual bases in a sequence, the qualities of one or more sequences individually, and the quality of a group of sequences as a whole. The sequence filter generates the filtered results based on preconfigured filtering strategies. Such filtering strategies may include basecalling quality filtering (BCQF) and passing indices synthetization (e.g., a logic AND operation) . In some embodiments, the filtered results obtained by one network block can be provided to another network block for further processing. For instance, multiple network blocks can be used in a hierarchical processing network structure to measure the DNA sequence signal quality to exclude the low-quality signals that are prone to be misidentified, thereby improving the final DNA sequencing quality. Details of the embodiments of the hierarchical processing network structure for quality filtering can be found in the U.S. non-provisional application No. 18 / 070,377, filed November 28, 2022, entitled “METHODS AND SYSTEMS FOR ENHANCING NUCLEIC ACID SEQUENCING QUALITY IN HIGH-THROUGHPUT SEQUENCING PROCESSES WITH MACHINE LEARNING, ” the content of which is incorporated by reference in its entirety for all purposes. Edge-based LLM data interpretation and diagnostic output generation

[0112] With reference still to FIGs. 7A and 7B, computing device (s) 103 of apparatus 100 can perform at least one of: aligning (step 722) the quality-improved sequences to one or more reference genomes, performing (step 724) a variant calling, or annotating (step 726) at least one of the quality-improved sequences to identify pathogens. These steps are sometimes also referred to as the secondary data analysis 720 and each of them is described below in greater detail.

[0113] In step 722, computing device (s) 103 of apparatus 100 can align a DNA sequence to a reference genome involves comparing the sequence of interest (e.g., a sequence in the sequencing data 153 stored in a FastQ file) to a known genome to identify similarities and differences. This process can be the basis for various applications in genomics, including variant calling, gene annotation, and understanding evolutionary relationships. In one example, the alignment process involves preparation, selection of reference genome, alignment, and output. In the preparation step, the DNA sequence (obtained through the sequencing technologies described above, and improved by quality filtering) is prepared for alignment, typically in the form of short reads or longer contiguous sequences. Next, a well-annotated reference genome is chosen from a database. This is a complete sequence of a species'DNA that serves as a standard for comparison. The actual alignment can then be performed using computational algorithms and software tools (e.g., deployed in the middleware layer 136 shown in FIG. 1B) installed on apparatus 100. These tools align the query sequence against the reference genome to find the best matching regions. The alignment process may generate a series of matches, mismatches, and gaps, resulting in an alignment file (often in formats like SAM or BAM) . This file contains information about where the query sequence aligns to the reference, including any variations such as insertions or deletions. Based on the alignment data, apparatus 100 can identify genetic variants, observing gene expression, or investigate genomic rearrangements, and perform pathogenic detections. Thus, the alignment process can be a fundamental step in sequencing data analysis.

[0114] In step 724, for example, the computing device (s) 103 of apparatus 100 can perform variant calling. Variant calling is the process of identifying variations between the DNA sequences of an individual (or a sample) and a reference genome. These variations can include single nucleotide polymorphisms (SNPs) , insertions, deletions, and structural variants. In a variant calling process, the alignment data from the previous step 722 may be preprocessed to, for example, remove duplicates and recalibrate base quality score. The preprocessing can identify and remove PCT duplicates to reduce bias. Further, the preprocessing can also adjust the quality scores of the bases based on known variant sites.

[0115] Next, in a variant calling process, algorithms and models may be used to analyze the aligned reads to identify positions where the sample differs from the reference. The algorithms and models can assess the depth of coverage (the number of reads at each position) and the quality of the observed variants.

[0116] After initial detection, variants can be filtered based on various criteria to reduce false positives. Factors used for such filtering include read depth, allele balance (the ratio of the variant allele to the reference allele) , and quality scores. The variants are also annotated to provide biological context, such as, whether they fall within coding regions or regulatory regions, potential effects on protein function, frequency in population databases (like gnomAD) .

[0117] The variants may be interpreted in a specific context, including assessing their clinical significance, potential impact on phenotype, or role in disease. Variant calling can be a required step in understanding genetic variation and its implications in health and disease.

[0118] In step 726, for example, the computing device (s) 103 of apparatus 100 can annotate at least one of the quality-improved sequences to identify pathogens. Pathogenic detection refers to the identification and characterization of pathogenic variants-genetic changes that are associated with diseases. This process is useful in clinical genetics, infectious disease diagnostics, and research into various health conditions. As described above, in step 722, the sequences in interest are aligned to a reference genome, and in step 724 variants are called. This step identifies differences from the reference sequence, which may include pathogenic variants.

[0119] Next, the variants are evaluated for pathogenicity, also referred to as pathogenicity assessment. The pathogenicity assessment may use one or more criteria including the clinical significance, the computational predictions, the population databases, and functional studies. Clinical significance criteria can be used to determine if the variant has been associated with a disease in clinical studies. In computational predictions, algorithms and models can be used to assess the potential impact of amino acid changes on protein function. Variants can be compared against population databases like gnomAD to see how common they are in healthy populations. In functional studies, experimental evidence can be gathered to confirm the effect of a variant on gene function.

[0120] As described above, variants can be annotated. Variants are annotated with information about their potential pathogenicity, location, and functional implications. This helps in understanding the biological context of the variants. Finally, reports can be generated including the variants classified as pathogenic, likely pathogenic, uncertain significance, likely benign, or benign. And recommendations for further testing or clinical management based on the detected variants can also be included in the report. An example report of diagnostic outputs 759 is shown in FIG. 7C.

[0121] In the embodiments described herein, some or all of the secondary data analysis 720 (e.g., alignment, variant calling, pathogenic detection) can be performed using an edge-based large language model (LLM) or models 155. LLMs 155 are typically based on transformer architecture, which allows them to process and generate text or other sequences efficiently. The transformer uses mechanisms like attention to focus on different parts of the input. An edge-based LLM is deployed locally to the computing device 103 of the apparatus 100 (e.g., installed in the apparatus) or at the location of the apparatus 100 (e.g., a server that is in the same room or same building as the apparatus) . As a result, the edge-based LLM is operatable in the offline mode without network access. Therefore, a cloud computing service may not be required for performing the secondary data analysis 720, reducing latency. As described, traditionally, some or all the secondary data analysis needs collaboration among molecular biology laboratories, bio-informatics teams, and clinical departments, this involves several laboratories, and these institutions are sometimes located in different places, so the processing time is very long, and the whole procedure often needs a day to several weeks. By deploying the edge-based LLM to the computing device of the apparatus 100 locally, the entire process of performing the secondary data analysis 720 can be significantly shortened. Therefore, efficiency is improved and the cost is reduced.

[0122] The edge-based LLM 155 can also be used for improving the tertiary data analysis. Traditionally, a human specialist, who has an intimate knowledge of interpretation of the data in the context of existing medical literature and databases, makes further data analysis manually (also referred to as the tertiary data analysis) and generates a comprehensive report referring to the secondary data analysis. This data interpretation and report generating process can also be performed by the computing device of apparatus 100 using the edge-based LLM 155. For example, the edge-based LLM is customizable to be trained with predetermined biomedical diagnostic knowledge in one or more particular biomedical knowledge areas. The LLM can thus be quite knowledgeable in certain biomedical knowledge areas and therefore can be used to interpret the secondary data analysis results and generate diagnosis report. For instance, if the user of apparatus 100 is a cardiologist, the edge-based LLM installed on apparatus 100 can be customized and trained with extensive knowledges of heart diseases. As such, the LLM is well equipped to perform the tertiary data analysis (e.g., data interpretation, report generation) based on the secondary data analysis results (e.g., pathogenic detection results) .

[0123] In some embodiments, one or more additional edge-based LLMs can be deployed to the computing device 103 of the apparatus 100, thereby forming a group of selectable LLMs with the edge-based LLM already deployed. In another embodiment, all desired edge-based LLMs 155 can be deployed to the apparatus 100 locally. Each of the groups of selectable LLMs corresponds to a different biomedical knowledge area. As such, the user of the apparatus 100 may be provided with options to select one or more edge-based LLMs for any particular task. For instance, one LLM may be customized and trained with knowledge of heart diseases; another LLM may be customized and trained with knowledge of diabetes; yet another LLM may be customized and trained with knowledge of cancer, and so forth. The user of the apparatus 100 can be provided with options on a user interface 108 to select one of the multiple LLMs available. The user interface 108 can be further configured to receive a user input indicating a selection from the group of the selectable LLMs.

[0124] In some embodiments, the edge-based LLM 155 is updatable with updated training data representing updated biomedical diagnostic knowledge in the one or more particular biomedical knowledge areas. The updating of the LLM can be performed periodically or as needed. Some example LLMs 155 include generative pre-trained transformer (GPT) , bidirectional encoder representations from transformers (BERT) , text-to-text transformer (T5) , Calude, etc. Any of these LLMs can be trained with particular knowledges in a certain biomedical knowledge area or areas. The trained LLM can thus be deployed to the apparatus 100. Training of the LLM can be performed remotely on a cloud or locally at apparatus 100.

[0125] FIG. 7A further illustrates that the biomedical knowledge database used for a LLM 155 can be a vectorized database 730. A vectorized database 730 includes many vectors representing the specific biomedical knowledge, such as nucleotide sequencing information, disease knowledge, and corresponding treatment protocols. The vectors from database 730 are one of the inputs to LLM 155. The vectors are used to enhance the background biomedical knowledge of LLM 155. The vectorized biomedical knowledge database 730 is stored locally at apparatus 100, and is updatable with knowledges stored external to the apparatus 100. For example, database 730 can be updated using documents, websites, references, etc. available on the Internet or other locations. With reference to FIG. 1A and FIG. 7A, the computing device103 of apparatus 100 can provide a diagnostic output 156 (e.g., disease diagnosis results, therapeutic regimes, recommended medicines or treatments) to the user.

[0126] FIG. 7C is a diagram illustrating data flow deployment based on an edge-based large-language model (LLM) 758 stored locally at an apparatus 100. As described above, the edge-based LLM can be used for any tasks described above for the secondary data analysis (e.g., alignment, variant calling, and pathogenic detection) and the tertiary data analysis (e.g., data interpretation and diagnosis report generation) . As an example, the description of FIG. 7C illustrates a data flow deployment process for data interpretation and diagnosis report generation using an edge-based LLM 758. In FIG. 7C, for example, a query 752 may be received, the query 752 may include the data to be interpreted (e.g., the variant calling results, the pathogenic detection results, etc. ) Computing device 103 of apparatus 100 can convert the query into an query vector and store the query vector in a vector database 754. The vector database 754 may further include other vectors representing known nucleotide sequencing related data (e.g., DNA related data) . The vector database 754 may include dense vectors.

[0127] The query vectors and similar vectors are converted to text. Similar vectors are those representing related data (e.g., DNA related data) . The original query represented by the query vector and the related data represented by the similar vectors can form an input 756 to the edge-based LLM 758. The input 756 can have certain format by fitting the query and related data into a prompt template. The edge-based LLM 758 can be a trained and customized LLM as described above. The LLM 758 is deployed locally to apparatus 100, such that it can operate in an offline mode. LLM 758 performs data interpretation based on the input 756, and generates answers 757. The answers 757 can be rearranged to produce diagnostic outputs 759, which in one example, comprises at least one of: disease diagnosis results, therapeutic regimens, or recommended medicines or treatments. The diagnostic outputs 759 can be in a report format as shown in FIG. 7C. For example, it can have a section for pathogenic detection results, another section for expected resistance information, another section for pathogenic bacteria information, another section for suggested treatments, etc.

[0128] As described above, traditionally, the integration and organization of data, along with generating reports annotated with disease information, require expert personnel to review tremendous literature and databases, which can be very time-consuming. In addition, lower-level facilities, such as community clinics or rural hospitals, often lack experience in interpreting reports and making clinical decisions, necessitating intervention from higher-level hospitals to reach diagnostic conclusions and formulate treatment plans.

[0129] The apparatus 100 described herein is an all-in-one intelligent edge-based LLM diagnostic machine that can provide an approach to utilize large language models to enable real-time generation of DNA sequencing reports including advice on disease diagnosis results and therapeutic regimens. The application of edge LLMs in the sequencing technology area refers to deploying complex artificial intelligence models directly at the data collection points or on user devices, instead of on remote cloud servers. This approach brings significant advantages to sequencing technology, particularly in terms of real-time processing, data privacy protection, latency reduction, and decreased dependence on network connectivity.

[0130] The technology described herein efficiently runs localized large models on the end-side using the cost-effective FPGA / GPU / TPU / NPU / CPU, providing a superior cost-performance ratio. It utilizes knowledge enhancement RAG (Retrieval-Augmented Generation) to load libraries with minimal investment for DNA application analysis, including pathogens, rare diseases, dermatological conditions, and more. This capability enables real-time, cost-effective analysis of DNA sequencing data and converts this data into client services for medical applications. In some examples, the generation of medical reports by LLMs needs a set of prompt templates carefully designed.

[0131] While the above examples use DNA sequencing and data analysis as an example, it is understood that the same or similar techniques can be applied to protein or DNA / protein or other multi-omics data analysis. System adjustments based on user feedback

[0132] After an apparatus (e.g., apparatus 100) performs the sequencing process, obtains sequencing data, performs secondary data analysis, interprets data, and generates diagnostic report, the diagnostic report can be reviewed by a user (e.g., a doctor or a clinician) . The user may provide some user feedback. The user may confirm that the diagnostic report is accurate or that the diagnostic report may need to be revised, corrected, updated, or supplemented. For example, if new medical knowledge or development has been available, the user feedback may indicate that the diagnostic report is not accurate because it did not take the latest medical development and knowledge into account. As a result, the diagnosis results, therapeutic regimens, and / or recommended medicines or treatments may need to be adjusted.

[0133] Such user feedback may be provided back to the apparatus 100 for improving its analysis accuracy. FIG. 8 is a flowchart illustrating a method 800 performed by the apparatus (e.g., apparatus 100) for receiving user feedback and making adjustment of the integrated process, in accordance with an embodiment of the present invention. With reference to FIG. 8, in step 802, the apparatus 100, via the user interface (e.g., user interface 108) , receives a user input representing user feedback of the diagnostic outputs. The user interface may be configured to provide the user options to confirm the diagnostic outputs or correct at least some of the diagnostic outputs. For example, with reference briefly back to FIG. 7C, the diagnostic outputs 759 may include a report having multiple categories like pathogenic detection results, expected resistance information, pathogenic bacteria information, suggested treatments, etc. For each category, the user interface 108 of the apparatus 100 can provide an area for comments or for confirmation. The area may include, for example, buttons, input fields, icons, text fields, checkboxes, etc. Using the area for comments or for confirmation, the user can provide feedback for each category. For instance, the user may add or remove certain items in the suggested treatment plan based on the patient’s personal situation, new medicine development and availability, and / or other factors. The user may also question the pathogenic detection results and provide some observations. In some embodiments, the user may request some of the processes (e.g., sequencing, basecalling, secondary data analysis, tertiary data analysis, report generation) or the entire process be re-run, with the existing sample or with a new sample.

[0134] With reference back to FIG. 8, in step 804, based on the user input representing the user feedback of the diagnostic outputs, the apparatus 100 can adjust at least one aspect of the integrated process described above. For example, in step 806, the control sub-system 114 of apparatus 100 can control the fluidic sub-system 112 to adjust at least one aspect of clusters of strands formed on un-patterned flowcells. In some examples, the control sub-system 114 automatically maximizes the utilization of the un-patterned flowcells. As described above, using un-patterned flowcells, the control sub-system 114 can vary one or more aspects of the clusters like spatial arrangements, density, and sizes. Based on the user input, the control sub-system 114 can send control signals to the fluidic sub-system 112 for making more optimized use of the un-patterned flowcells (e.g., increasing the density, maximizing the spatial arrangements, reducing the cluster sizes, etc. ) .

[0135] In some embodiments, based on the user input, the control sub-system 114 of apparatus 100 can control the imaging sub-system 118 to adjust (step 808) at least one aspect of image sensing. For example, the control sub-system 114 can send control signals to the imaging sub-system 118 to focus on certain portions of the flowcells, to dynamically control the imaging throughput (e.g., by activating additional image sensors when higher imaging throughput is needed and by deactivating image sensors when lower imaging throughput is needed) , by applying or removing filters associated with the image sensors, etc. Various aspects of the imaging sub-systems 118 that can be controlled by the control sub-system 114 are described in more detail in U.S. non-provisional application No. 17 / 246,487 (now patent No. US 11,175,219) , entitled “THROUGHPUT-SCALABLE ANALYTICAL SYSTEM USING SINGLE MOLECULE ANALYSIS SENSORS, ” the content of which is incorporated herein by reference in its entirety for all purposes.

[0136] In some embodiments, based on the user input, the control sub-system 114 can control (step 810) the basecalling for obtaining the sequencing data. As described above, the basecalling process involves partitioning the images into grids of patches, performing linear projection to obtain position embeddings, performing a position encoding addition process to add class tokens, encoding using self-attention neural network to obtain encoder outputs, and classifying the encoder outputs to obtain the sequencing data. One or more parameters and models in the basecalling process can be adjusted based on the user input. For instance, the neural network parameters of the self-attention transformer-based models can be adjusted / retrained / updated by using the user input as a part of new training data. The model can thus be fine tuned to generate more accurate sequencing data. As described above, the images of fluorescent signals may be pre-processed to improve the image quality. The image pre-processing can also be adjusted using the user input to improve the image quality. Various aspects of the imaging sub-systems that can be controlled by the control sub-system are described in more detail in U.S. non-provisional application No. 17 / 681,672, now U.S. Patent No. 11,580,641, entitled “DEEP LEARNING BASED METHODS AND SYSTEMS FOR NUCLEIC ACID SEQUENCING, ” the content of which is hereby incorporated by reference in its entirety for all purposes.

[0137] In some embodiments, with reference to FIG. 8, the control sub-system 114 can adjust (step 812) at least one aspect of the sequencing data analysis process. As described above, the sequencing data analysis process may include performing quality filtering of the sequencing data, aligning the quality-improved sequences to one or more reference genomes, performing a variant calling, and annotating at least one of the quality-improved sequences to identify pathogens. One or more of these processing steps can be adjusted based on user input. For instance, based on the user input, the control sub-system 114 can enhance the filtering of the sequencing data to obtain higher-quality data to meet the quality requirement of the users or the NGS applications. To do that, the parameters of the quality filtering network structure may be adjusted automatically. Various aspects of the quality filtering network structure that can be controlled by the control sub-system are described in more detail in the U.S. non-provisional application No. 18 / 070,377, filed November 28, 2022, entitled “METHODS AND SYSTEMS FOR ENHANCING NUCLEIC ACID SEQUENCING QUALITY IN HIGH-THROUGHPUT SEQUENCING PROCESSES WITH MACHINE LEARNING, ” the content of which is incorporated by reference in its entirety for all purposes.

[0138] Furthermore, as described above, an edge-based LLM model may be deployed to the computing device of the apparatus for performing some analyses of the sequencing data to obtain the diagnostic outputs. Therefore, the LLM model parameters can also be adjusted / updated / retrained based on the user input. In some examples, the user may want to choose another edge-based LLM that has been trained using a vectorized database of a different biomedical knowledge area.  The user interface 108 of the apparatus 100 can be further configured to receive a user input indicating a selection of a different LLM model from the group of the selectable LLMs. As described above, the vectorized biomedical knowledge database, as one of the inputs of LLM, can help to enhance the background biomedical knowledge of LLM. Thus, one LLM model may result in better and more accurate diagnostic report compared to another LLM model. The user interface may also provide the user with options to select multiple LLM models to compare their results. It is understood that the apparatus 100 can be configured to make any adjustments desired based on user input, including adjustments to the sequencing process, the clustering, the basecalling process, the quality filtering process, the secondary and tertiary data analysis process, and the report generating process. The adjustments the apparatus can make are not limited to those described above. Exemplary computing device embodiment

[0139] FIG. 9 is an example block diagram of a computing device 900 that may incorporate embodiments of the present invention. FIG. 9 is merely illustrative of a machine system to carry out aspects of the technical processes described herein, and does not limit the scope of the claims. One of ordinary skill in the art would recognize other variations, modifications, and alternatives. In one embodiment, the computing device 900 typically includes a monitor or graphical user interface 902, a data processing system 920, a communication network interface 912, input device (s) 908, output device (s) 906, and the like.

[0140] As depicted in FIG. 9, the data processing system 920 may include one or more processor (s) 904 that communicate with a number of peripheral devices via a bus subsystem 918. These peripheral devices may include input device (s) 908, output device (s) 906, communication network interface 912, and a storage subsystem, such as a volatile memory 910 and a nonvolatile memory 914. The volatile memory 910 and / or the nonvolatile memory 914 may store computer-executable instructions and thus forming logic 922 that when applied to and executed by the processor (s) 904 implement embodiments of the processes disclosed herein.

[0141] The input device (s) 908 include devices and mechanisms for inputting information to the data processing system 920. These may include a keyboard, a keypad, a touch screen incorporated into the monitor or graphical user interface 902, audio input devices such as voice recognition systems, microphones, and other types of input devices. In various embodiments, the input device (s) 908 may be embodied as a computer mouse, a trackball, a track pad, a joystick, wireless remote, drawing tablet, voice command system, eye tracking system, and the like. The input device (s) 908 typically allow a user to select objects, icons, control areas, text and the like that appear on the monitor or graphical user interface 902 via a command such as a click of a button or the like. Graphical user interface 902 can be used in various steps of the methods described above (e.g., method 800) to receive user inputs for making the adjustments of the various processes performed by the apparatus.

[0142] The output device (s) 906 include devices and mechanisms for outputting information from the data processing system 920. These may include the monitor or graphical user interface 902, speakers, printers, infrared LEDs, and so on as well understood in the art.

[0143] The communication network interface 912 provides an interface to communication networks (e.g., communication network 916) and devices external to the data processing system 920. The communication network interface 912 may serve as an interface for receiving data from and transmitting data to other systems. Embodiments of the communication network interface 912 may include an Ethernet interface, a modem (telephone, satellite, cable, ISDN) , (asynchronous) digital subscriber line (DSL) , FireWire, USB, a wireless communication interface such as Bluetooth or WiFi, a near field communication wireless interface, a cellular interface, and the like. The communication network interface 912 may be coupled to the communication network 916 via an antenna, a cable, or the like. In some embodiments, the communication network interface 912 may be physically integrated on a circuit board of the data processing system 920, or in some cases may be implemented in software or firmware, such as "soft modems" , or the like. The computing device 900 may include logic that enables communications over a network using protocols such as HTTP, TCP / IP, RTP / RTSP, IPX, UDP and the like.

[0144] The volatile memory 910 and the nonvolatile memory 914 are examples of tangible media configured to store computer readable data and instructions forming logic to implement aspects of the processes described herein. Other types of tangible media include removable memory (e.g., pluggable USB memory devices, mobile device SIM cards) , optical storage media such as CD-ROMS, DVDs, semiconductor memories such as flash memories, non-transitory read-only-memories (ROMS) , battery-backed volatile memories, networked storage devices, and the like. The volatile memory 910 and the nonvolatile memory 914 may be configured to store the basic programming and data constructs that provide the functionality of the disclosed processes and other embodiments thereof that fall within the scope of the present invention. Logic 922 that implements embodiments of the present invention may be formed by the volatile memory 910 and / or the nonvolatile memory 914 storing computer readable instructions. Said instructions may be read from the volatile memory 910 and / or nonvolatile memory 914 and executed by the processor (s) 904. The volatile memory 910 and the nonvolatile memory 914 may also provide a repository for storing data used by the logic 922. The volatile memory 910 and the nonvolatile memory 914 may include a number of memories including a main random access memory (RAM) for storage of instructions and data during program execution and a read only memory (ROM) in which read-only non-transitory instructions are stored. The volatile memory 910 and the nonvolatile memory 914 may include a file storage subsystem providing persistent (non-volatile) storage for program and data files. The volatile memory 910 and the nonvolatile memory 914 may include removable storage systems, such as removable flash memory.

[0145] The bus subsystem 918 provides a mechanism for enabling the various components and subsystems of data processing system 920 communicate with each other as intended. Although the communication network interface 912 is depicted schematically as a single bus, some embodiments of the bus subsystem 918 may utilize multiple distinct buses.

[0146] It will be readily apparent to one of ordinary skill in the art that the computing device 900 may be a device such as a smartphone, a desktop computer, a laptop computer, a rack-mounted computer system, a computer server, or a tablet computer device. As commonly known in the art, the computing device 900 may be implemented as a collection of multiple networked computing devices. Further, the computing device 900 will typically include operating system logic (not illustrated) the types and nature of which are well known in the art.

[0147] One embodiment of the present invention includes systems, methods, and a non-transitory computer readable storage medium or media tangibly storing computer program logic capable of being executed by a computer processor. The computer program logic can be used to implement embodiments of processes and methods described herein, including method 300 for basecalling, method 400 for image preprocessing, method 800 for cluster detection, method 1000 for feature extraction, and various deep learning algorithms and processes.

[0148] Those skilled in the art will appreciate that computer system 900 illustrates just one example of a system in which a computer program product in accordance with an embodiment of the present invention may be implemented. To cite but one example of an alternative embodiment, execution of instructions contained in a computer program product in accordance with an embodiment of the present invention may be distributed over multiple computers, such as, for example, over the computers of a distributed computing network.

[0149] While the present invention has been particularly described with respect to the illustrated embodiments, it will be appreciated that various alterations, modifications and adaptations may be made based on the present disclosure and are intended to be within the scope of the present invention. While the invention has been described in connection with what are presently considered to be the most practical and preferred embodiments, it is to be understood that the present invention is not limited to the disclosed embodiments but, on the contrary, is intended to cover various modifications and equivalent arrangements included within the scope of the underlying principles of the invention as described by the various embodiments referenced above and below.

Claims

An apparatus for performing an integrated process of sequencing a biological sample and providing diagnostic outputs based on an edge-based large-language model (LLM) , the apparatus comprising:a fluidic sub-system configured to generate fluorescent signals based on analysis of the biological sample;an imaging sub-system coupled to the fluidic sub-system to sense the fluorescent signals and obtain images of the fluorescent signals;one or more processors of a computing device; andmemory storing one or more instructions, when executed by the one or more processors, cause the computing device to perform:basecalling using the images of the fluorescent signals to obtain sequencing data of the biological sample;analyzing the sequencing data based on the edge-based LLM stored at the apparatus, the edge-based LLM being operatable at the apparatus in an offline mode;providing, via a user interface, the diagnostic outputs based on analysis results of the sequencing data.The apparatus of claim 1, wherein:the one or more processors include at least one of: one or more graphical processing units (GPUs) , one or more central processing units (CPUs) , one or more tensor processing units (TPUs) , or one or more programmable logic devices (PLDs) ; andthe one or more processors enable the apparatus to operate in the offline mode to perform the integrated process of performing the basecalling, analyzing the sequencing data, and providing the diagnostic outputs based on the biological sample.The apparatus of claim 1, wherein the fluidic sub-system, the imaging sub-system, the one or more processors, the user interface, and optionally one or more other sub-systems are mounted within a housing.The apparatus of claim 1, wherein the imaging sub-system comprises a plurality of semiconductor-based image sensors.The apparatus of claim 3, wherein each of the plurality of semiconductor-based image sensors is disposed in a different semiconductor die sharing a common substrate.The apparatus of claim 1, wherein the fluidic sub-system is further configured to:receive the biological sample, andprepare the biological sample to form clusters of strands on a plurality of un-patterned flowcells having planar surfaces, wherein the clusters of strands are formed on planar surfaces of the plurality of un-patterned flowcells.The apparatus of claim 6, wherein the un-patterned flowcells have no wells for accommodating droplets of the biological samples.The apparatus of claim 6, further comprising a control sub-system, the control sub-system is configured to control, for each flowcell of the plurality of flowcells, at least one aspect of the clusters of strands formed on the planar surface of the flowcell, the at least one aspect including one or more of the following:spatial arrangement of the clusters of strands;density of the clusters of strands; orsizes of the clusters of strands.The apparatus of claim 8, wherein the control sub-system is further configured to control the at least one aspect of the clusters of strands based on a user input representing feedback of the diagnostic outputs, the user input being received via the user interface.The apparatus of claim 1, wherein performing the basecalling using the images of the fluorescent signals to obtain the sequencing data of the biological sample comprises:performing the basecalling based on a neural network.The apparatus of claim 10, wherein performing the basecalling based on the self-attention transformer-based neural network comprises:partitioning the images of the fluorescent signals into a gird of patches;performing linear projection of the gird of patches to obtain position embeddings;performing a position encoding addition process to add class tokens to the position embeddings; andencoding, by an encoder, outputs of the position encoding addition process using the self-attention transformer-based neural network to obtain encoder outputs.The apparatus of claim 11, wherein performing the basecalling based on the self-attention transformer-based neural network further comprises:classifying encoder outputs to obtain the sequencing data, wherein the sequencing data include base classes of the sequences of the biological sample.The apparatus of claim 1, wherein the edge-based LLM is deployed locally to the computing device of the apparatus such that the edge-based LLM is operatable in the offline mode without network access.The apparatus of claim 1, wherein the edge-based LLM is customizable to be trained with predetermined biomedical diagnostic knowledge in one or more particular biomedical knowledge areas.The apparatus of claim 14, wherein the edge-based LLM is updatable with updated training data representing updated biomedical diagnostic knowledge in the one or more particular biomedical knowledge areas.The apparatus of claim 1, wherein one or more additional edge-based LLMs are deployed to the computing device of the apparatus, the edge-based LLM and the additional edge-based LLMs form a group of selectable LLMs, each of the group of selectable LLMs corresponds to a different biomedical knowledge area.The apparatus of claim 16, wherein the user interface is further configured to receive a user input indicating a selection from the group of the selectable LLMs.The apparatus of claim 1, wherein analyzing the sequencing data based on the edge-based LLM stored locally at the apparatus comprises:performing quality filtering of the sequencing data to obtain quality-improved sequences, the quality-improved sequences having higher quality scores than raw sequences included in the sequencing data; andperforming at least one of:aligning the quality-improved sequences to one or more reference genomes,performing a variant calling, andannotating at least one of the quality-improved sequences to identify pathogens.The apparatus of claim 1, wherein the diagnostic outputs comprise at least one of: disease diagnosis results, therapeutic regimens, or recommended medicines or treatments.The apparatus of claim 1, wherein analyzing the sequencing data is further based on a vectorized biomedical knowledge database, the vectorized biomedical knowledge database comprising vectors representing nucleotide sequencing information, disease knowledge, and corresponding treatment protocols.The apparatus of claim 20, wherein the vectorized biomedical knowledge database is stored locally at the apparatus, and is updatable with knowledges stored external to the apparatus.The apparatus of claim 1, wherein the computing device is further caused to perform:receiving, via the user interface, a user input representing user feedback of the diagnostic outputs; andadjusting, based on the user input representing the user feedback, at least one of:controlling the fluidic sub-system to adjust at least one aspect of clusters of strands formed on un-patterned flowcells,controlling the imaging sub-system to adjust at least one aspect of image sensing,the basecalling to obtain the sequencing data, orthe analyzing of the sequencing data.A method performed by an apparatus for controlling at least a part of a sequencing process for a biological sample, the method comprises:receiving the biological sample, anddistributing, by a fluidic sub-system of the apparatus, the biological sample to a plurality of un-patterned flowcells having planar surfaces, wherein clusters of strands are formed on the planar surfaces of the plurality of un-patterned flowcells; andcontrolling, by a control sub-system of the apparatus, at least one aspect of the clusters of strands formed on the planar surfaces of the plurality of flowcells.The method of claim 23, wherein controlling the at least one aspect of the clusters of strands comprises, for each flowcell of the plurality of flowcells, controlling one or more of the following:spatial arrangement of the clusters of the strands;density of the clusters of strands; orsizes of the clusters of the strands.The method of claim 23, wherein the controlling the at least one aspect of the clusters of strands is further based on a user input representing user feedback of diagnostic outputs provided by the apparatus, the user input being received via a user interface of the apparatus.A non-transitory computer readable medium comprising a memory storing one or more instructions which, when executed by one or more processors of at least one computing device, cause the at least one computing device to perform:basecalling using images of the fluorescent signals to obtain sequencing data of a biological sample;analyzing the sequencing data based on an edge-based LLM stored at an apparatus comprising the computing device, the edge-based LLM being operatable at the apparatus in an offline mode; andproviding, via a user interface, diagnostic outputs based on analysis results of the sequencing data.