Nucleic acid analysis device and nucleic acid analysis method

By using the base prediction part and the alignment part in nucleic acid analysis for image alignment, combined with supervised learning and machine learning methods, the problem of reduced image alignment accuracy is solved, and higher base alignment accuracy and reliability are achieved.

CN115516075BActive Publication Date: 2025-09-02HITACHI HIGH TECH CORP
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
CN202080100282.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2020-05-12
Publication Date
2025-09-02
Estimated Expiration
2040-05-12

AI Technical Summary

Technical Problem

In the existing nucleic acid analysis technology, image alignment accuracy is susceptible to fluorescence intensity attenuation and focus differences, resulting in an increase in the possibility of base call errors, especially during multiple cycles.

Method used

The base prediction unit and the alignment unit are used to perform image alignment, and the luminescent images of multiple phosphors are detected through sensors. The image feature quantity extraction and base prediction are used to perform image quality reduction and base prediction, and the training data is updated in combination with machine learning methods to improve the alignment accuracy.

Benefits of technology

The image alignment accuracy of nucleic acid analysis is improved, the possibility of base call errors is reduced, and the accuracy and reliability of base arrangement are ensured.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115516075B_ABST
    Figure CN115516075B_ABST
Patent Text Reader

Abstract

The object is to provide a nucleic acid analysis technology that is robust to image alignment accuracy. A preferred aspect of the present invention provides a nucleic acid analysis device characterized by comprising: a base prediction unit configured to perform base prediction using as input a plurality of images obtained by detecting luminescence from a bio-related substance disposed on a substrate; an alignment unit configured to align the plurality of images with a reference image; and an extraction unit configured to extract bright spots from the plurality of images. The base prediction unit, configured to extract features of the images, including the surrounding pixels of the extracted bright spots within the plurality of images, and to predict bases based on the features, takes as input images of the plurality of images, including the location of the extracted bright spots.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a nucleic acid analysis technology for measuring organism-related substances. Background Art

[0002] In recent years, a method has been proposed for nucleic acid analysis devices that simultaneously determines the base sequences of a large number of DNA fragments to be analyzed by loading them into a flow cell constructed from a glass or silicon substrate. In this method, a substrate containing a fluorescent dye corresponding to the bases is introduced into the analysis region of the flow cell containing the large number of DNA fragments. Excitation light is then irradiated onto the flow cell to detect fluorescence emitted from each DNA fragment, thereby identifying (calling) the bases.

[0003] To analyze a large number of DNA fragments, the analysis area is typically divided into multiple detection fields. Each time the irradiation is performed, the detection field is changed and analysis is performed in all detection fields. A polymerase extension reaction is then used to introduce a new fluorescent dye-containing substrate, and each detection field is analyzed using the same procedure as above. Repeating this cycle allows for efficient base sequence determination (see Patent Document 1).

[0004] In this analysis, fluorescence emitted from an amplified DNA sample (hereinafter referred to as a colony) secured to a substrate is captured and image processing is used to identify bases. Specifically, each colony within the fluorescence image is identified, and the fluorescence intensity corresponding to each base at the position of each colony is obtained. The base is then identified based on this fluorescence intensity (see Patent Document 2).

[0005] Generally speaking, even when capturing the same detection field of view, fluorescence images captured by different cameras will be captured at different locations on the flow cell due to the control accuracy limitations of the drive device used to change the field of view. Consequently, a given colony is captured at different coordinate positions within each fluorescence image. Therefore, to accurately identify each colony, it is necessary to accurately determine the coordinate position of each colony on the flow chip.

[0006] To achieve this goal, there are methods that place reference marks on the substrate to determine positions on the substrate, and methods that detect the positions of individual clusters in a captured image through image correlation matching with a reference image. Here, a reference image is an image whose positional information, representing the positional coordinates of the clusters on the image, is known and generated based on the design information of the flow path chip. Alternatively, any one of multiple images captured in each detection field of view can be used as the reference image, and clusters in other images can be aligned with the clusters on the reference image. Hereinafter, this process of aligning the positional coordinates between the reference image and the target image will be referred to as alignment.

[0007] Prior art literature

[0008] Patent Document 1: Japanese Patent Application Laid-Open No. 2020-60

[0009] Patent Document 2: WO2017-203679 Summary of the Invention

[0010] However, alignment depends on the image pattern of the colony and the degree of focus of the image. Generally speaking, as the extension reaction and imaging cycles are repeated, the DNA degrades and the fluorescence intensity decays, so the accuracy of alignment decreases as the number of cycles increases. In addition, in the case of poor focus during imaging, alignment accuracy also decreases. Moreover, as described later, when alignment is performed between fluorescence images of different cameras, there is also the possibility of reduced alignment accuracy due to large differences in lens skew. When the accuracy of alignment decreases, the reliability of the fluorescence intensity obtained at each colony position decreases, and the possibility of calling the wrong base increases.

[0011] The present invention has been made in view of such circumstances, and an object of the present invention is to provide a nucleic acid analysis technology that is robust with respect to image alignment accuracy.

[0012] A preferred aspect of the present invention is a nucleic acid analysis device, characterized in that it comprises: a base prediction unit for performing base prediction using as input a plurality of images obtained by detecting luminescence from a bio-related substance arranged on a substrate; a positioning unit for positioning the plurality of images to a reference image; and an extraction unit for extracting bright spots from the plurality of images. The base prediction unit takes as input an image of surrounding pixels including the position of the extracted bright spot within the plurality of images, extracts a feature value of the image, and predicts a base based on the feature value.

[0013] In an example of a more specific means, the multiple images are images obtained by a sensor detecting multiple types of luminescence from multiple types of fluorescent bodies introduced into a biologically related substance, and for each of the multiple types of luminescence, the sensor performing the detection and at least one of the optical paths to the sensor performing the detection are different.

[0014] In another more specific example of the method, the base prediction unit is composed of a predictor capable of supervised learning.

[0015] In another more specific example of the means, the base prediction unit receives as input an image of at least one cycle selected from the previous cycle and the next cycle in addition to the image of the cycle in which the prediction is performed.

[0016] Another preferred aspect of the present invention is a nucleic acid analysis method for performing base prediction using multiple images obtained by detecting luminescence from a bio-related substance as input to a base prediction device. The method comprises performing a colony position determination phase and a base sequence determination phase. The colony position determination phase comprises: an alignment process for aligning the multiple images; and a colony position determination process for extracting bright spots from the multiple images to determine the colony positions of the bio-related substance. In the base sequence determination phase, images containing pixels surrounding the extracted colony positions within the multiple images are input to the base prediction device, features of the images are extracted, and bases are predicted based on the features.

[0017] Another preferred aspect of the present invention is a machine learning method for a base predictor that performs base prediction using multiple images obtained by detecting luminescence from a bio-related substance as input. The method comprises: a first base prediction step of generating a first base prediction result based on the multiple images; a first training data generation step of generating first training data based on the alignment results of the first base prediction result and a reference array; a predictor update step of updating parameters of the base predictor using the first training data generated in the first training data generation step; a second base prediction step of generating a second base prediction result based on the multiple images using the base predictor updated in the predictor update step; a second training data generation step of generating second training data based on the alignment results of the second base prediction result and a reference array; and a training data update step of updating the first training data using the second training data.

[0018] According to the present invention, a nucleic acid analysis technology that is robust to image alignment accuracy can be provided. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] Figure 1 This is a block diagram showing a schematic configuration example of a nucleic acid analysis device according to each embodiment.

[0020] Figure 2 It is a line diagram showing the processing steps for decoding the base sequence of DNA in each embodiment.

[0021] Figure 3 It is a plan view for explaining the concept of the detection field on the flow cell of each embodiment.

[0022] Figure 4 These are explanatory diagrams showing the concepts of four types of bright spots in fluorescent images in each detection field of view according to each embodiment.

[0023] Figure 5 It is an explanatory diagram showing the concept of determining the base sequence in each example.

[0024] Figure 6This is a block diagram for explaining an example of the configuration of a computer of the nucleic acid analysis apparatus of Example 1.

[0025] Figure 7 This is a flowchart showing the flow of the base calling stage in Example 1.

[0026] Figure 8 This is a flowchart showing the flow of the colony position determination phase in Example 1.

[0027] Figure 9 It is an explanatory diagram showing the concept of positional deviation between cycles in each embodiment.

[0028] Figure 10 This is an explanatory diagram showing the concept of measuring the positional deviation amounts of a plurality of locations within an image in each embodiment.

[0029] Figure 11 This is an explanatory diagram showing the concept of cluster extraction based on multiple cycles in Example 1.

[0030] Figure 12 This is an explanatory diagram showing the concept of determining the colony position based on multiple cycles in the first embodiment.

[0031] Figure 13 This is a flowchart showing the flow of the base sequence determination stage in Example 1.

[0032] Figure 14 This is an explanatory diagram showing the concept of the ROI image for each community in the first embodiment.

[0033] Figure 15 This is a block diagram showing the concept of the base predictor of Example 1.

[0034] Figure 16 This is a block diagram showing an example of the structure of the convolutional neural network of Example 1.

[0035] Figure 17 This is an explanatory diagram showing the concept of a base prediction device according to the second embodiment that takes ROI images of multiple cycles as input.

[0036] Figure 18 This is a block diagram illustrating the concept of base prediction using multiple base predictors in Example 3.

[0037] Figure 19 This is a flowchart showing the process of learning base prediction parameters in Example 4.

[0038] Figure 20 This is an explanatory diagram for explaining the concept of the alignment process of the fourth embodiment.

[0039] Figure 21This is an explanatory diagram for explaining the relationship between the bases of all the colonies and the bases of the aligned colonies in Example 4.

[0040] Figure 22 This is an explanatory diagram for explaining the concept of correct information constituting training data in the fourth embodiment.

[0041] Figure 23 This is an explanatory diagram for explaining the concept of updating training data in the fourth embodiment.

[0042] Figure 24 This is an explanatory diagram for explaining the concept of improving prediction performance through repeated learning in Example 4.

[0043] Figure 25 This is a graph for explaining the concept of improving the alignment rate and error rate through repeated learning in Example 4.

[0044] Figure 26 This is an explanatory diagram showing the concept of expanding training data by fuzzy processing in the sixth embodiment.

[0045] Figure 27 This is an explanatory diagram showing the concept of expanding training data by shift processing in the sixth embodiment.

[0046] Figure 28 This is an explanatory diagram showing the concept of selecting training data based on the signal strength of the colony position in the seventh embodiment.

[0047] Figure 29 This is an explanatory diagram showing the concept of selecting training data based on the reliability of community positions in Example 7.

[0048] Figure 30A This is an explanatory diagram used to illustrate the different system structures of the nucleic acid analysis device, base prediction unit, and learning unit of Example 8.

[0049] Figure 30B This is an explanatory diagram used to illustrate the different system structures of the nucleic acid analysis device, base prediction unit, and learning unit of Example 8.

[0050] Figure 30C This is an explanatory diagram used to illustrate the different system structures of the nucleic acid analysis device, base prediction unit, and learning unit of Example 8.

[0051] Figure 31 This is a video diagram showing an example of a screen for selecting a plurality of base predictors in Example 9.

[0052] Figure 32 This is a video diagram showing an example of a setting screen when using a base predictor separately for each cycle in Example 9.

[0053] Figure 33 This is a video diagram showing an example of a setting screen for adding and deleting a data group used for learning a base predictor in Example 9.

[0054] (Explanation of Symbols)

[0055] 100: Nucleic acid analysis device; 101, 122: Two-dimensional sensor; 102, 121: Imaging lens; 103: Bandpass filter; 104: Excitation filter; 105, 120: Dichroic mirror; 106: Filter module; 107: Light source; 108: Objective lens; 109: Circulation cell; 112, 115: Piping; 113: Reagent container; 114: Reagent storage unit; 116: Waste liquid container; 117: Loading table; 118: Temperature-controlled substrate; 119: Computer; 123: Analysis area; 124: Detection field of view; 800: Base calling unit; 801: Alignment unit; 802: Community extraction unit; 803: Base prediction unit; 804: Learning unit. DETAILED DESCRIPTION

[0056] The following describes embodiments of the present invention with reference to the accompanying drawings. In the drawings, functionally identical elements are sometimes denoted by the same reference numerals. Furthermore, the drawings illustrate specific embodiments and implementation examples based on the principles of the present invention. These are provided solely for understanding the present invention and are not intended to limit the present invention. In other words, it should be understood that the descriptions in this specification are merely exemplary and are not intended to limit the claims or application examples in any way.

[0057] While the various embodiments described below are described in sufficient detail to enable those skilled in the art to implement the present invention, it should be understood that other configurations and configurations are possible, and that structural and configurational changes and substitutions of various elements are possible without departing from the scope and spirit of the technical concept of the present invention. Therefore, the following description should not be construed as limiting these concepts. Furthermore, while the nucleic acid analysis devices of the various embodiments target DNA fragments for measurement and analysis, they can also target RNA, proteins, and other substances in addition to DNA. The present invention is applicable to all aspects of biologically relevant substances.

[0058] Furthermore, as described later, the embodiments of the present disclosure may be implemented as software that operates on a general-purpose computer, or as dedicated hardware, or as a combination of software and hardware.

[0059] Hereinafter, the various processes in the embodiments of the present disclosure will be described using the various processing units (e.g., the alignment unit, the community extraction unit, the base prediction unit, and the learning unit) as the subject (the main body of action) as "programs." However, since programs are executed by a processor and perform the specified processing while using memory and a communication port (communication control device), the processor can also be used as the subject for description. Part or all of the program can be implemented using dedicated hardware or modularized.

[0060] Various embodiments of the present invention are described below, sequentially with reference to the accompanying drawings. In a representative embodiment, a base calling method is provided, which uses as input an ROI (Region of Interest) image surrounding a colony in a fluorescence image after alignment between fluorescence images and colony detection within each fluorescence image. Furthermore, a method is provided for learning a base prediction engine by repeatedly updating training data and performing base calling based on comparisons of called base sequences with reference sequences.

[0061] Example 1

[0062] (1) Nucleic acid analysis device

[0063] Figure 1 The nucleic acid analysis device 100 includes a flow cell 109, a liquid delivery system, a transport system, a temperature control system, an optical system, and a computer 119. The flow cell 109 includes a nucleic acid analysis substrate of each embodiment described below.

[0064] The liquid delivery system provides components for supplying reagents to the flow cell 109. This system includes a reagent storage unit 114 that houses multiple reagent containers 113; a nozzle 111 that connects to the reagent container 113; a pipe 112 that introduces the reagents into the flow cell 109; a waste liquid container 116 that discards waste liquids such as reagents that have reacted with DNA fragments; and a pipe 115 that introduces waste liquids into the waste liquid container 116.

[0065] The conveying system is a system for moving the analysis area 123 of the flow cell 109 described later to a predetermined position. The conveying system is provided with a stage 117 on which the flow cell 109 is placed and a driving motor (not shown) for driving the stage. The stage 117 can move in each direction of the X-axis and the Y-axis that are orthogonal to each other in the same plane. In addition, the stage 117 can also move in the Z-axis direction orthogonal to the XY plane by a driving motor different from the stage driving motor.

[0066] The temperature control system is a system that adjusts the reaction temperature of the DNA fragments. The temperature control system is installed on the mounting table 117 and includes a temperature control substrate 118 that promotes the reaction between the DNA fragments to be analyzed and the reagents. The temperature control substrate 118 is implemented by, for example, a Peltier element.

[0067] The optical system provides components for irradiating the analysis area 123 of the flow cell 109 (described later) with excitation light and detecting fluorescence emitted from DNA fragments. The optical system consists of a light source 107, a condenser lens 110, an excitation filter 104, dichroic mirrors 105 and 120, a bandpass filter 103, an objective lens 108, imaging lenses 102 and 121, and two-dimensional sensors 101 and 122. The excitation filter 104, the dichroic mirror 105, and the bandpass filter 103, also known as an absorption filter, are included as a set within a filter module 106. The bandpass filter 103 and the excitation filter 104 determine the wavelength range that allows fluorescence of a specific wavelength to pass.

[0068] The following describes the process of excitation light irradiation in the optical system. Excitation light emitted from light source 107 is focused by condenser lens 110 and incident on filter module 106. The incident excitation light is transmitted only by the excitation filter 104 within a specific wavelength band. The transmitted light is reflected by dichroic mirror 105 and focused by objective lens 108 onto flow cell 109.

[0069] Next, the fluorescence detection process in the optical system will be described. Due to the focused excitation light, the fluorescent substance excited by the specific wavelength band mentioned above among the four fluorescent substances in the DNA fragment fixed on the flow cell 109 is excited. The fluorescence emitted by the excited fluorescent substance passes through the dichroic mirror 105, passes only the specific wavelength band through the bandpass filter 103, and then passes only the specific wavelength band through the dichroic mirror 120, while transmitting other wavelength regions. The light that has passed through the dichroic mirror 120 is imaged as a fluorescent spot on the two-dimensional sensor 101 through the imaging lens 102. In addition, the light reflected by the dichroic mirror 120 is imaged as a fluorescent spot on the two-dimensional sensor 122 through the imaging lens 121.

[0070] In this embodiment, only one type of phosphor is designed to be excited by a specific wavelength band. As will be described later, four base types can be distinguished based on the type of phosphor. Furthermore, to enable sequential detection of these four types of phosphors, two sets of filter modules 106 are prepared, each switching between the wavelength bands of the irradiation and detection light. The excitation filter 104, dichroic mirrors 105 and 120, and bandpass filter 103 within each filter module 106 are designed with transmission characteristics that enable detection of each phosphor with maximum sensitivity.

[0071] Computer 119, like a conventional computer, includes a processor (CPU), storage devices (various memories such as ROM and RAM), input devices (keyboard, mouse, etc.), and output devices (printer, display, etc.). This computer functions as a control processing unit. In addition to controlling the aforementioned liquid delivery system, transport system, temperature control system, and optical system, this control processing unit also analyzes the fluorescent images detected and generated by the two-dimensional sensors 101 and 122 of the optical system and performs base calling for each DNA fragment. However, the control of the aforementioned liquid delivery system, transport system, temperature control system, and optical system, image analysis, and base calling do not necessarily need to be performed by a single computer 119. To distribute the processing load and reduce processing time, these processes can be performed by multiple computers, each functioning as a control unit and a processing unit.

[0072] (2) Methods for interpreting DNA base sequences

[0073] refer to Figures 2 to 4 The method for deciphering the base sequence of DNA is described below. In addition, as described later, the same DNA fragment is pre-amplified and densely populated on the flow cell 109. In the amplification of DNA fragments, existing technologies such as emulsion PCR and bridge PCR can be used.

[0074] Figure 2 This diagram illustrates the processing steps for deciphering DNA base sequences. By repeating the loop process (S22) M times, a complete run (S21) for deciphering is performed. M represents the length of the desired base sequence and is predetermined. Each loop process is used to determine the kth (k = 1 to M) base and is divided into the chemical process (S23) and imaging process (S24) described below.

[0075] (A) Chemical treatment: Treatment for base extension

[0076] In the chemical treatment ( S23 ), the following processes (i) and (ii) are performed.

[0077] (i) If the cycle is not the first cycle, the fluorescently labeled nucleotides (described below) from the previous cycle are removed from the DNA fragments and washed. The reagent for this purpose is introduced into the flow cell 109 via the pipe 112. The waste liquid after washing is discharged into the waste liquid container 116 via the pipe 115.

[0078] (ii) Reagents containing fluorescently labeled nucleotides flow through tubing 112 to analysis area 123 on flow cell 109. The flow cell temperature is adjusted using temperature control substrate 118, causing DNA polymerase to initiate an extension reaction, and complementary fluorescently labeled nucleotides are incorporated into the DNA fragments on the colony.

[0079] Here, fluorescently labeled nucleotides refer to the use of 4 types of fluorescent bodies (FAM, Cy3, Texas Red (TxR), Cy5) to label 4 types of nucleotides (dCTP, dATP, dGTP, dTsTP). Each fluorescently labeled nucleotide is recorded as FAM-dCTP, Cy3-dATP, TxR-dGTP, and Cy5-dTsTP. These nucleotides are complementarily incorporated into the DNA fragment, so if the base of the actual DNA fragment is A, dTsTP is incorporated, if it is base C, dGTP is incorporated, dCTP is incorporated in base G, and dATP is incorporated if it is base T. That is, the fluorescent body FAM corresponds to base G, Cy3 corresponds to base T, TxR corresponds to base C, and Cy5 corresponds to base A. In addition, with respect to each fluorescently labeled nucleotide, in order to prevent extension to the next base, the 3' end is blocked.

[0080] (B) Imaging processing: Processing to generate fluorescence images

[0081] The imaging process (S24) is performed by repeating the imaging process (S25) for each detection field described below N times. Here, N is the number of detection fields.

[0082] Figure 3 This diagram illustrates the concept of the detection field of view. The detection field of view (FOV) 124 corresponds to each of the N regions when the entire analysis region 123 is divided. The size of the detection field of view 124 is the size of the area detectable by the two-dimensional sensors 101 and 122 during a single fluorescence detection, and is determined by the design of the optical system. As described later, a fluorescence image corresponding to the four types of fluorescent substances is generated for each detection field of view 124.

[0083] (B-1) Imaging Processing for Each Detection Field

[0084] In the detection field imaging process ( S25 ), the following processes (i) to (iv) are performed.

[0085] (i) The stage 117 is moved (S26) so that the detection field 124 for fluorescence detection is brought to a position irradiated with the excitation light from the objective lens 108. At this time, the objective lens 108 may be driven to adjust the focus position in order to correct the vertical deviation caused by the movement of the stage 117.

[0086] (ii) The filter module 106 is switched to a setting corresponding to the fluorescent substance (FAM / Cy3) ( S27 ).

[0087] (iii) By irradiating the two-dimensional sensors 101 and 122 with excitation light and exposing them simultaneously, a fluorescence image (FAM) is generated in the two-dimensional sensor 101 and a fluorescence image (Cy3) is generated in the two-dimensional sensor 122 ( S28 ).

[0088] (iv) The filter module 106 is switched to a setting corresponding to the phosphor (TxR / Cy5) ( S29 ).

[0089] (v) By irradiating the two-dimensional sensors 101 and 122 with excitation light and exposing them simultaneously, a fluorescence image (TxR) is generated in the two-dimensional sensor 101 and a fluorescence image (Cy5) is generated in the two-dimensional sensor 122 ( S30 ).

[0090] By executing the above processing, fluorescence images for four types of fluorophores (FAM, Cy3, TxR, and Cy5) are generated for each detection field. In these fluorescence images, the signals of the fluorophores corresponding to the base types of the DNA fragments immobilized in the flow cell 109 appear as clusters on the image. Specifically, the clusters detected in the FAM fluorescence image are determined to be base A, the clusters detected in the Cy3 fluorescence image are determined to be base C, the clusters detected in the TxR fluorescence image are determined to be base T, and the clusters detected in the Cy5 fluorescence image are determined to be base G.

[0091] Figure 4 : is a diagram showing the concept of the cluster of four types of fluorescent images in each detection field. Figure 4 As shown in (a), for example, in a certain detection field in a certain cycle, there are clusters at 8 positions P1 to P8, and the bases are A, G, C, T, A, C, T, and G.

[0092] At this time, the fluorescence images of the four types of fluorescent bodies (Cy5, Cy3, FAM, TxR) are shown in (b) to (e), and the communities are detected at positions P1 to P8 according to the corresponding base types. The positions of P1 to P8 are the same in the four fluorescent images. However, due to the design of the optical system, differences in the optical path are generated for each wavelength, so strictly speaking, there is a possibility of being different. Therefore, by performing the alignment processing described later as needed, the community positions of the four types of fluorescent images can be made the same. However, in the case of crosstalk where the wavelength characteristics of the filters of each fluorescent body overlap, a community of a certain base type is observed in more than two fluorescent images. The base type at this time can be identified by the ROI image of the community of the four types of fluorescent images as described later. Based on the above, the base category of each community detected in the detection field is determined.

[0093] (C) Repetition of loop processing

[0094] By repeating the above loop process for the number of times the desired base sequence length M is required, a base sequence of length M can be determined for each cluster.

[0095] Figure 5 This is a diagram showing the concept of determining the base sequence. Figure 5 As shown, in each colony (DNA fragment having the base sequence ACGTATACGT...), when one base is extended by chemical treatment (S23) in a certain cycle (#N), for example, Cy3-dATP is introduced. In the imaging process, this fluorescently labeled nucleotide is detected as a colony on the fluorescence image of Cy3. Similarly, in cycle (#N+1), it is detected as a colony on the fluorescence image of Cy5. In cycle (#N+2), it is detected as a colony on the fluorescence image of TxR. In cycle (#N+3), it is detected as a colony on the fluorescence image of FAM. Through the above cyclic processing from cycle #N to cycle #N+3, the base sequence in this colony is determined to be TACG.

[0096] (3) Details of base calling processing

[0097] As described above, the DNA fragment to be detected is observed as four bright spots on the fluorescent image, and the bases are called in each cycle.

[0098] Figure 6This is a block diagram illustrating the functional configuration of the computer 119 in the nucleic acid analysis device 100. The computer includes a control unit 806, which controls the aforementioned liquid delivery system, transport system, temperature control system, optical system, and base calling unit; a communication unit 807, which exchanges control commands and image data between the computer 119 and the device; a UI unit 808, which presents screens to the user and accepts user input; a storage unit 809, such as a memory or hard disk; and a base calling unit 800, which outputs base sequences. The base calling unit 800 includes an alignment unit 801, a community extraction unit 802, a base prediction unit 803, and a learning unit 804. The base calling process performed by the base calling unit 800 is described below.

[0099] In this embodiment, it is assumed that the control unit 806, communication unit 807, UI unit 808, and base calling unit 800 are implemented using software. Specifically, the computer 119 storage device stores programs for performing calculations or processing in each unit, and the processing unit 870 executes these programs, performing processing in cooperation with hardware such as the input device 880, output device 890, and storage unit 809. As described above, the control unit 806, communication unit 807, UI unit 808, and base calling unit 800 may also be implemented using hardware instead of software.

[0100] Figure 7 : is a diagram showing the flow of base calling processing. The base calling processing is performed in two stages: the colony position determination stage (S90) and the base sequence determination stage (S91).

[0101] (A) Community location determination stage

[0102] The base calling unit 800 executes the colony position determination stage (S90). In this embodiment, in the colony position determination stage (S90), the images from the first cycle to the Nth cycle are looped to determine the colony to be called.

[0103] refer to Figure 8 The process of the colony position determination stage (S90) is described below. The alignment unit 801 obtains the four-color fluorescence image in the first FOV of the first cycle (Cycle) through steps S101, S102, and S103, and performs alignment processing with the reference image (S104). The alignment processing is described below.

[0104] (A-1) Image Alignment

[0105] As described above, the nucleic acid analyzer 100 acquires four fluorescence images using the two sensors 101 and 122, resulting in positional shifts between the fluorescence images. Furthermore, when repeatedly capturing the same detection field of view in each cycle, the stage 117 is moved to change the detection field of view in each cycle. Consequently, positional shifts occur between cycles for the same detection field of view due to control errors during stage movement.

[0106] Figure 9 is a diagram showing the concept of positional offset between cycles. Figure 9 As shown, for a given detection field of view (FOV), there is a possibility that the imaging position may shift due to stage control errors between the Nth cycle (a) and the (N+1)th cycle (b). Consequently, the DNA fragment positions (P1-P8) in the fluorescence image of cycle N are detected as different positions (P1'-P8', respectively) in the fluorescence image of cycle (N+1). However, these bright spots are all caused by the same DNA fragment. Therefore, in order to determine the base sequence of each colony, it is necessary to correct the positional shift between the colonies detected in each fluorescence image.

[0107] In order to correct this positional offset, it is necessary to align each fluorescent image with respect to a common reference image. Here, the reference image refers to a common image used in the position coordinate system of the colony. For example, if the position of the colony is known as design information, a reference image can also be made based on the known colony position. As an example, an image of brightness of a two-dimensional Gaussian distribution with a predefined variance corresponding to the colony size can be made with the colony position (x, y) as the center. Alternatively, a reference image can be made based on any actual image in the actual image captured. As an example, the images of each detection field of view in the initial cycle can be used as the reference image, and the images of each detection field of view after the second cycle can be aligned to the reference image.

[0108] Known matching techniques can be applied to image alignment. As an example, a cross-correlation function m(u, v) is calculated between a template image t(x, y) obtained by cutting out a portion of a reference image and a target image f(x, y) obtained by cutting out a portion of the input image. The positional offset is determined as S_1 = (u, v), which gives the maximum value. Here, an example of t(x, y) is a 256×256 pixel image at the center of the reference image. Similarly, an example of f(x, y) is a 256×256 pixel image at the center of the input image. To calculate the positional offset, instead of using a cross-correlation function, a normalized cross-correlation that takes into account brightness differences or a phase-limited correlation can be used. Furthermore, when detecting the angular offset between images, polar coordinate transformation of the images can be performed, allowing the cross-correlation or phase-limited correlation described above to be applied to the images after the angular orientation has been transformed to the horizontal.

[0109] Furthermore, the positional deviation amount may be obtained at a plurality of points depending on the degree of image distortion.

[0110] Figure 10 For example, if it can be assumed that there is no distortion in the image and that all pixels are offset at the same position, that is, the same offset caused by the stage, it can be applied. Figure 10 The left side of (a) shows the position offset S_1(u, v).

[0111] On the other hand, for example, in the case where there is distortion in the image and the amount of positional shift varies depending on the position in the image (the case where the flow cell 109 is deformed due to heating and the positional shift varies), Figure 10 As shown on the right side of (a), the position offset is calculated at n points in the image, and the position offsets S_1, S_2, ... S_n at the multiple points are calculated. In calculating the position offset at each point, the image centered at the position of each point in the reference image and the input image is cut out and used as the template image and the object image respectively, and the position offset with the largest correlation is calculated as described above. And, based on the obtained n position offsets, the coefficients of the affine transformation and the polynomial transformation are calculated, for example, using the least squares method, so that the position offset of any pixel position can be formulated (refer to Figure 10 (b)). In addition, the transformation can also be calculated bidirectionally. That is, the mutual transformation between the coordinate system of the input image and the coordinate system of the reference image can be defined.

[0112] (A-2) Bright spot extraction processing

[0113] exist Figure 8In the aligned fluorescence image, the cluster extraction unit 802 extracts the bright spot positions representing the clusters (S105). One example of a method for obtaining the bright spot positions is to use a predetermined threshold value to divide the input image into bright spot areas and non-bright spot areas, and then search for a maximum value in the bright spot areas.

[0114] Before the bright spot extraction process, the input image can also be subjected to noise removal using a low-pass filter, median filter, or the like. Furthermore, background correction can be performed to address situations where uneven brightness occurs within the image. As an example of background correction, the following method can be used: an image obtained by pre-photographing an area where no DNA fragments are present is used as a background image and subtracted from the input image. Alternatively, a high-pass filter can be applied to the input image to remove low-frequency background components.

[0115] Furthermore, even if a cluster is included in any of the four types of fluorescence images, there is a possibility that a bright spot from a single cluster may be included in multiple fluorescence images due to the influence of crosstalk as described above. Bright spots from different fluorescence images that are determined to be close to each other through alignment can also be merged as described later.

[0116] As described above, in this embodiment, the cluster position is determined using images from the beginning to the Nth cycle. Here, N is expressed as the number of cluster determination cycles. N can be about 1 to 8.

[0117] Depend on Figure 11 This figure illustrates the advantages of using multiple cycles to extract colonies. The figure schematically shows three cycles' worth of bases from adjacent colonies #1 to #5. If adjacent colonies in the same cycle contain the same base, they appear adjacent in a single fluorescence image, making it difficult to distinguish between them. Therefore, referring to multiple cycles makes it easier to distinguish colonies at sites with different bases.

[0118] exist Figure 11 In the example shown in Figure 2, in cycle 1, colonies #2 and #3 (114), and colonies #4 and #5 (115) with the same base are adjacent, making it difficult to distinguish these colonies in the fluorescence image of cycle 1 alone. In cycle 2, colonies #4 and #5 (113) have different bases, making them easier to distinguish. Similarly, in cycle 3, colonies #2 and #3 (111) have different bases, making them easier to distinguish. In this way, multiple cycles can be used to improve the distinguishability of colonies.

[0119] exist Figure 8In the process, the cluster extraction unit 802 repeats this bright spot extraction process (S105) for each detection field within each loop (S106). For the last detection field in the loop, the process shifts to the next loop (S102), and then to the first detection field (S103), repeating the alignment process (S104) and bright spot extraction (S105). When the number of loops from the beginning reaches the number of loops N (N for cluster determination) (S107), the clusters obtained up to N loops are merged (S108).

[0120] (A-3) Community merging process

[0121] In the cluster merging process ( S108 ), the cluster extracting unit 802 merges the bright spots extracted from the fluorescence images corresponding to N cycles and converted into the coordinate system of the reference image by position alignment.

[0122] Figure 12 The concept of community merging is shown. As shown in the figure, even if it is the same community, it will not be accurately configured at the same coordinates in the coordinate system of the reference image due to errors in the alignment calculation. Therefore, in the community merging process, communities that are adjacent to each other within a certain distance can be regarded as one (Figure a), or they can be regarded as a valid community when there is only one community (Figure b). In addition, in the case where the size of adjacent communities exceeds a certain threshold, they can be divided into two communities. The center of gravity of the merged new communities can also be recalculated. In these merging algorithms, existing grouping technologies such as the k-means method can also be applied.

[0123] Through the above-described processing, positional shifts between the four types of fluorescence images and positional shifts between images of a plurality of cycles can be corrected, and the colony position in each image can be determined.

[0124] (B) Base sequence determination stage

[0125] Next, the narrative Figure 7 Details of the base sequence determination stage (S91) in the base calling process of In this stage, bases of all cycles are called to determine the base sequence for the cluster determined in the cluster position determination stage (S90).

[0126] Figure 13 The flow of the base sequence determination stage (S91) is shown below: The process moves to the first FOV of the first loop through steps S131, S132, and S133, and then the following processing is performed for each FOV.

[0127] (B-1) Positioning Process (S134)

[0128] The alignment unit 801 aligns the four fluorescence images of the target FOV with respect to the reference image. This method is the same as that described in (A-1). However, since alignment has already been performed in the previous stage for the images up to the colony position determination cycle number, the alignment results from that stage can also be used.

[0129] (B-2) Community Position Coordinate Transformation (S135)

[0130] The colony extraction unit 802 transforms the coordinates of all colonies on the reference coordinate system determined in the previous stage into the coordinate system of the four fluorescence images being processed. The alignment results of step S134 are used in this transformation. Thus, the colony positions on each fluorescence image are obtained.

[0131] (B-3) ROI Image Extraction (S136)

[0132] The colony extraction unit 802 extracts a ROI (Region of Interest) image centered on the colony position on each fluorescence image.

[0133] Figure 14 This section illustrates the concept of ROI image extraction. The "+" represents the center of the colony in the fluorescence image, and a W-pixel x H-pixel region centered therein is extracted. W and H are appropriately determined in advance based on the size of the colony and image resolution. It is preferable to minimize the reflection of adjacent colonies. Furthermore, the pixel values ​​of the fluorescence image can be normalized based on base prediction, described below, before extracting the ROI image.

[0134] By extracting not only the bright spots in the fluorescence image but also the surrounding pixels, we can obtain incidental information about the fluorescence image, such as image positional shift, defocus, and crosstalk. This increased information content in the fluorescence image can improve the prediction accuracy of base prediction tools using machine learning, as described later.

[0135] (B-4) Base prediction (S137)

[0136] The base prediction unit 803 performs base prediction using the sets of ROIs of the four-color fluorescence images as input.

[0137] Figure 15 An example of a base predictor within the base prediction unit 803 is shown. The base predictor is composed of a feature quantity calculator and a multinomial classifier. The feature quantity calculator calculates a feature quantity from the input image, and the multinomial classifier classifies the image into A, G, C, or T based on the feature quantity.

[0138] exist Figure 16As an example of such a base prediction machine, a structure using a CNN (Convolutional Neural Network) is shown. In the convolution layer (Conv in the figure), the following filtering operation is performed on the input image. CNN is an example of a neural network capable of supervised learning.

[0139] [Formula 1]

[0140] u ijm =∑ k ∑ p ∑ q I i+p,j+q,k h pqkm +b ijm… (1)

[0141] Here, I is the input image, h is the filter coefficient, b is the addition term, k is the input image channel, m is the output channel, i and p are the horizontal positions, and j and q are the vertical positions.

[0142] In the ReLU layer, the following activation function is applied to the output of the above convolutional layer.

[0143] [Formula 2]

[0144] ReLU(x)=max(0,x)…(2)

[0145] In the activation function, nonlinear functions such as tanh function, logistic function, and normalized linear function (ReLU) can also be used.

[0146] The Pooling layer slightly reduces the positional sensitivity of the features extracted by the convolutional and ReLU layers, ensuring that the output remains constant even when the feature's position within the image changes slightly. Specifically, a representative value is calculated from a partial region of the feature at a fixed step size. The maximum value of these representative values ​​is calculated using the average value, for example. The Pooling layer does not have any parameters that change due to learning.

[0147] Affine layers, also known as fully coupled layers, define weighted coupling from all units in the input layer to all units in the output layer. Here, i is the index of the unit in the input layer, j is the index of the unit in the output layer, w is the weight coefficient between them, and b is the additive term.

[0148] [Formula 3]

[0149]

[0150] In CNN, the above-mentioned convolutional layer, ReLU layer, and Pooling layer are repeatedly executed, and the result of the Affine layer and ReLU layer is the image feature. Based on the image feature obtained in this way, multiple classification is performed, that is, the base discrimination of A, G, C, and T.

[0151] As an example of the multi-classification method, in this embodiment, Affine layer processing is further performed on the above-mentioned image feature amount, and logistic regression using the following softmax function is applied to the result.

[0152] [Formula 4]

[0153]

[0154] Here, y is a value representing the likelihood of the label (here, the base) corresponding to output unit k. In this embodiment, output unit k corresponds to the likelihood of base class k, and the base type with the highest likelihood is set as the final classification result.

[0155] The filter coefficients and additive terms of the convolutional layer and the weight coefficients and additive terms of the Affine layer are predetermined by the learning process of the learning unit 804, as described later. These coefficients are stored as predictor parameters in the storage unit 809. The base prediction unit 803 may also retrieve these coefficients from the storage unit 809 as appropriate during the base prediction process.

[0156] By performing the above-described base prediction process ( S137 ) on all FOVs of all loops ( S138 , S139 ), the base sequences in all FOVs of all loops are determined (the base sequence determination stage S91 ends).

[0157] As described above, in the nucleic acid analysis device of Example 1, the ROI images of each fluorescent color obtained by alignment and community extraction are used as input, their feature values ​​are calculated, and base prediction is performed based on the feature values, so that base prediction that is robust to image position offset and defocus can be achieved.

[0158] Example 2

[0159] according to Figure 17 To illustrate Example 2. In Example 2, the base prediction device is as follows Figure 17 As shown, as input to the base prediction device, in addition to the input of the ROI image of a certain cycle (Nth cycle), the ROI images of the previous and next cycles are also input. These ROI images of the previous and next cycles are pre-aligned and are set to be images from the same community. Figure 17In the figure, all ROI images are of the same size, and are input so that each fluorescence image corresponds to one channel. That is, in this figure, ROI images of 12 channels are input.

[0160] An advantage of using ROI images of previous and subsequent cycles as input is that base prediction can be performed taking into account the influence of attenuation between cycles.

[0161] Attenuation refers to the incomplete chemical reactions of DNA fragments in each cycle, which causes the extension reaction to progress unevenly. This leads to the contamination of not only the signals from bases in each cycle, but also signals from bases in previous and subsequent cycles. This attenuation is known to occur at a certain rate in each cycle, and as the cycles progress, this effect accumulates, contributing to reduced base calling accuracy.

[0162] Thus, during learning and prediction, the fluorescence signals of each cycle are mixed with fluorescence signals from the same community in the preceding and following cycles. Therefore, the base predictor uses a model that predicts bases based on input from ROI images that include both preceding and following cycles, enabling base prediction that takes attenuation into account. In this figure, only the preceding and following cycle ROI images are used as input, but it is also possible to use ROI images from two or more cycles, both preceding and following, as input. Alternatively, it is also possible to use only the preceding and following images as input.

[0163] As described above, in the nucleic acid analysis device of Example 2, by adding a plurality of ROI images before and after the cycle to be predicted to the input image to perform base prediction, it is possible to achieve highly accurate base prediction taking into account the influence of attenuation.

[0164] Example 3

[0165] according to Figure 18 To illustrate Example 3. In Example 3, Figure 18 As shown, the base prediction unit 803 combines multiple base prediction units described above to perform base prediction. In this figure, each base prediction unit uses an ROI image of the same size as input and is configured with different base prediction parameters. These different base prediction parameters are determined in advance under different conditions through learning, as described below. Different conditions here refer to different equipment, different room temperatures, different cycles, etc., and can be determined by taking into account variations in the camera images used during operation.

[0166] In the output layer of this diagram, the final base likelihood is output based on the outputs of multiple base predictors (the likelihoods of each base) determined under these different conditions. This structure enables more reliable base predictions that take into account various conditions. The output layer can output the maximum value of all base predictors or the average or weighted sum of the likelihoods of each base.

[0167] Furthermore, the CNN network structure, the number of loops before and after the input ROI image, and the ROI size may differ between the base predictors. Furthermore, the feature extraction method and the multinomial classification algorithm may also differ.

[0168] As described above, in the nucleic acid analysis device of Example 3, a plurality of base predictors determined under different conditions are used, and therefore base prediction that is robust to differences in various conditions and more accurate can be achieved.

[0169] Example 4

[0170] In Example 4, an example of a method for learning the base predictor in the base prediction unit 803 described in Example 1 is shown. In Example 4, the base predictor in Example 1 is trained by Figure 6 The following description will be given using as an example a structure in which the base calling unit 800 described in the preceding text is added with the learning unit 804. However, the learning of the base prediction unit may also be performed by a separate device. In this case, the learning unit 804 may be omitted. Figure 6 Learning Department 804.

[0171] Details of learning process

[0172] Figure 19 The flow of the base predictor learning process in the learning unit 804 is shown.

[0173] (A) Initial base prediction (S191)

[0174] The initial base prediction (S191) outputs the initial value of the base sequence for each community determined in the community position determination stage (S90). In the output of the base sequence, it can also be a prediction based on a simple rule such as selecting the base corresponding to the fluorescent color with the maximum brightness of each community. Alternatively, it can be achieved by setting the initial prediction parameters using the base predictor described in Example 1 (for example, the base prediction unit 803 in the initial setting state). However, since the correctness or incorrectness of the base sequence is determined by alignment with the reference sequence as described later, in the initial base prediction, it is preferred that the accuracy of the base sequence alignment be a certain amount.

[0175] (B) Alignment Processing (S192)

[0176] The base sequences obtained in the initial base prediction are aligned (S192). The alignment process is a process of matching the base sequences of all the clusters obtained with the reference sequence.

[0177] Figure 20 A conceptual diagram illustrating the alignment process. The reference alignment is a known, accurate alignment corresponding to the DNA sample measured by the nucleic acid analyzer. The reference alignment used here can be a widely available genomic alignment or an accurate alignment associated with a commercially available sample. Alignment algorithms can use retrieval methods based on known techniques such as the Burrows-Wheeler transform.

[0178] exist Figure 20 , shows a situation where base sequences 2003 and 2004 of a particular group within a set 2005 of base sequences of all groups are aligned to partial sequences 2001 and 2002 of a reference sequence 2000, respectively. For such aligned base sequences, bases that match the reference sequence are determined to be correct, while bases that do not match are determined to be incorrect. In this figure, base sequence 2003 and partial sequence 2001 match, so all bases in base sequence 2003 are determined to be correct. In base sequence 2004 and partial sequence 2002, the second base from the beginning is determined to be incorrect, while all other bases are determined to be correct. Furthermore, no correctness judgment is made for misaligned base sequences output by the base prediction unit 803.

[0179] Figure 21 The figure shows the relationship between the set of bases in a community and the set of bases in the aligned community. Set 2300 is the set of bases estimated based on the total cycle length by the base prediction unit 803 for all communities obtained by the community extraction unit 802 in Example 1. Set 2301 is the set of communities aligned as a result of the alignment process performed on set 2300. As shown in the figure, set 2301 can be further divided into correct bases and incorrect bases.

[0180] The following alignment evaluation index is calculated for the alignment result obtained in this way and stored in the storage unit 809 .

[0181] (1) Targeting rate: the ratio of the number of targeted clusters to the total number of extracted clusters

[0182] (2) Correct base rate (or incorrect base rate): The ratio of the number of correct bases (or incorrect bases) to the total number of bases in the aligned community

[0183] (C) Training Data Update (S193)

[0184] Information obtained by combining the fluorescence image corresponding to each base of the base sequence aligned in step S192 (or S196) and the correct base indicated by the reference sequence is considered as one correct information, and training data is created (S193).

[0185] Figure 22 shows the concept of correct information in the training data. As an example, Figure 20 , an example of correct information in a base sequence 2004 aligned with a partial sequence 2002 of a reference sequence is shown. The arrangement position of the base sequence 2004 corresponds to a cycle, and each base is estimated based on the ROI image of each cycle at its community position. The ROI image corresponding to each base and the correct base information represented by the partial sequence 2002 are grouped as follows. Figure 22 The correct information is represented by 2100 to 2104. Regarding the correct information 2100, 2102, 2103, and 2104, the predicted bases are correct.

[0186] In correct information 2101, the predicted base is T (in base sequence 2004), and while the prediction result is incorrect, correct information can be generated by pairing the correct base "A" indicated by the reference sequence with the ROI image. Furthermore, the ROI image for each colony can be stored as a pair of fluorescence image link information and colony position information on each fluorescence image. When input to the base prediction unit, the ROI image can be obtained from this information.

[0187] In this embodiment, the aligned sequence information can include both information indicating correctly predicted bases and information indicating incorrect bases. In particular, incorrect bases are inferred as ROI images, where base prediction is inherently difficult. Therefore, by including correct information about incorrect bases in the training data, it is expected that the performance of the base prediction device will be improved.

[0188] If training data already exists, correct information about bases that do not exist in the existing training data is added to the training data.

[0189] Figure 23 A conceptual diagram showing the updating of training data. An information table such as the one shown in the figure is stored in the storage unit 809 for all bases in all cycles of all clusters. As an example, the contents of the information table may include the following information.

[0190] (a) Link destination of each fluorescence image data (since it is common to all colonies, it can also be stored for each cycle)

[0191] (b) Location information of clusters in each image

[0192] (c) Predicted bases

[0193] (d) Whether it is aligned

[0194] (e) Correct base (if aligned)

[0195] (f) Likelihood of each base

[0196] (g) Whether the training data contains correct information

[0197] Referring to (g) above, if the correct information of the base is not included in the training data, the correct information of the base is added to the training data. At this time, the contents of (c), (d), and (f) can also be updated as needed based on the base prediction results described later.

[0198] (D) Base Prediction Module Update (S914)

[0199] Using this newly created or updated training data, learning is performed to update the parameters of the base predictor (S194). Known machine learning algorithms can be applied to learning. In the case of the convolutional neural network described in Example 1, the known error backpropagation method can be applied to determine the filter coefficients and additive terms of the convolutional layer and the weight coefficients and additive terms of the Affine layer. The cross-entropy error function can be used as the error function in this case.

[0200] The coefficients at the start of learning may be randomly initialized if it is the first time, or a known pre-learning method such as an autoencoder may be applied. If the base predictor update in step S194 is the second or later update, the predictor parameters determined previously may be applied.

[0201] The predictor parameters can be calculated by repeating the calculation for a predetermined number of iterations (passes) using a known method such as gradient descent to update the predictor parameters in a manner that minimizes the error function. The learning coefficients used to update the predictor parameters can also be appropriately modified using known methods such as AdaGrad and Adadelta.

[0202] Furthermore, in calculating the gradient of the error function used for parameter updating, the gradient can be calculated using the sum of the errors for all data using the gradient descent method, or the known probabilistic gradient descent method can be used to randomly divide the data into a set of M predetermined data, called mini-batches, and calculate the gradient for each mini-batch to update the predictor parameters. Furthermore, in the probabilistic gradient descent method, the data can be shuffled for each pass to reduce the impact of data bias.

[0203] In addition, during the above learning, a portion of the training data can be separated into verification data, and this verification data can be used to evaluate the base prediction performance based on the learned predictor parameters. The prediction performance based on this verification data can also be visualized for each traversal. As an indicator of this prediction performance, the prediction accuracy representing the proportion of correct predictions, or its opposite error rate, the value of the error function (loss), etc. can also be used. The predictor parameters obtained through learning in this way are applied to the base predictor. However, as described later, the final judgment of whether to adopt the latest predictor parameters is made in step S199, so the past predictor parameters before the update (before learning in step S194) are stored in the storage unit 809.

[0204] (E) Base prediction (S195)

[0205] The base prediction unit 803 performs base prediction for all clusters using the predictor parameters obtained in step S194, thereby outputting the base sequences of all clusters. The base prediction of Example 1 is applied to this prediction.

[0206] (G) Realignment Processing (S196)

[0207] The learning unit 804 performs alignment processing again on the base sequence obtained in step S195. This alignment processing is exactly the same as that in step S192 except that the input base sequence is different, so detailed description is omitted.

[0208] (F) Update continuation determination (S197)

[0209] Based on the alignment rate and correct base rate obtained in step S196, it is determined whether to continue or terminate the above-mentioned update process of the predictor parameters.

[0210] like Figure 24 As shown, it is considered that by basically repeating the above-mentioned update process of the predictor parameters, the alignment rate and the correct base rate gradually increase, the increase rate gradually decreases and soon saturates, or the learning fails and the increase rate becomes negative.

[0211] exist Figure 25 In the example, the alignment rate and incorrect base rate are plotted for each number of times. Therefore, as a method for determining whether to continue updating the predictor parameters, the following determination method can be used: thresholds are set for the rate of increase in the alignment rate and the rate of increase in the correct base rate, and the predictor parameter update is terminated when these rates of increase fall below the thresholds.

[0212] When the update of the predictor parameters is continued, the process returns to step S193 and the training data is updated using the correct information of the targeted community obtained in step S196.

[0213] (G) Base prediction unit determination (S198)

[0214] When the updating of the predictor parameters is completed in step S198, an optimal parameter is selected from the predictor parameters obtained by repeated updating including the initial base prediction (S191), and a base predictor is determined (S198).

[0215] Examples of criteria for selecting optimal parameters include maximizing the alignment rate or the accuracy rate, etc. However, as the alignment rate increases, there is a possibility that bases that are more difficult to predict will be aligned, so the parameters can also be determined based on criteria such as maximizing the weighted sum of the alignment rate and the accuracy rate.

[0216] As described above, the nucleic acid analysis device of Example 4 generates base sequences using the initial base predictor for a set of captured images provided for training. Correct information is extracted from the populations aligned with the reference sequence through alignment processing to update training data, and this training data is used to learn the predictor parameters. By repeating this process, high-quality training data is extracted from the set of captured images used for training and applied to base predictor training, thereby improving base calling accuracy.

[0217] Example 5

[0218] Example 5 is an example of learning parameters of a base predictor in which, in addition to the ROI image of the cycle as the target of base estimation described in Example 2, ROI images of multiple cycles before and after the cycle are added to the channels of the input image.

[0219] In this embodiment, the ROI image added to the training data as correct information is not as in Figure 22 The learning method is the same as that of Example 4.

[0220] As described above, in Example 5, by creating training data in which multiple ROI images before and after the cycle of the prediction target are added to the input image and learning the predictor parameters, it is possible to achieve highly accurate base prediction taking into account the influence of attenuation.

[0221] Example 6

[0222] In the sixth embodiment, an image obtained by applying image processing to an ROI image included in the training data is added to the training data as a new ROI image.

[0223] exist Figure 26In

[15] , as an example, filtering is applied to the ROI images of the original training data to create moderately blurred images, which are then added to the training data to learn the predictor parameters. This process can improve the robustness of base prediction against focus shifts.

[0224] exist Figure 27 As another example, images obtained by performing a shift process on the original training data ROI image are created and added to the training data to learn the predictor parameters. This process can improve the robustness of base prediction against variations in alignment accuracy.

[0225] Furthermore, in addition to the above-mentioned examples, images to which processing such as rotation, enlargement, or reduction is applied may be added.

[0226] As described above, in Example 6, the robustness of the base predictor can be improved by adding ROI images subjected to various image processing to training data to learn the parameters of the base predictor.

[0227] Example 7

[0228] In the seventh embodiment, the learning process ( Figure 19 ) in the training data update (S193) step to perform the screening of the correct information added to the training data. In Example 4, as shown in FIG. Figure 22 As previously described, the correct information (2100-2104) for all bases aligned in step S192 is included in the training data. However, there is a possibility that the aligned bases contain data that is not intended for use as training data. In this embodiment, the training data is updated after such bases that are not intended for use as training data are identified and removed through a screening process. An example of determining when to remove bases from the training data is described below.

[0229] (1) Community separation

[0230] Figure 28 The concept of stripping the detection community from the flow path chip is shown in the figure. In this figure, an example of aligning the base sequence 2901 called in a certain community to the reference sequence 2900 is shown. In this example, inconsistencies in the reference sequence are caused in the second cycle, the fourth cycle, the fifth cycle, and the sixth cycle. In this figure, 2902 to 2906 represent the signal intensities of four fluorescent images (G, A, T, and C correspond to the fluorescent substances FAM, Cy5, Cy3, and TxR, respectively) at the center position of the community corresponding to the base sequence 2901. In addition, the signal intensity can be obtained directly from the fluorescent image or through a calculation process such as linear transformation based on a pre-measured color transformation matrix.

[0231] In 2902, the signal intensity corresponding to C is high and is also significant compared to the signal intensities of other bases. In contrast, after 2903, all signal intensities become low. In this case, if the fluorescence intensity decreases as a whole after a certain cycle, the following possibility is considered: Figure 2 No fluorescence stripped from the flow chip was obtained during the chemical treatment cycles described in . Since the fluorescence intensity of all fluorescent bodies is, for example, below a threshold, when considering such stripping of the community, regardless of whether it is consistent with the reference arrangement or not, it is removed from the training data.

[0232] (2) Base variation

[0233] Figure 29 This is a conceptual diagram of detecting base variation. The figure shows an example of aligning a base sequence 3001 called in a certain community to a reference sequence 3000. In this example, a discrepancy with the reference sequence occurs in the third cycle. In the figure, 3002 to 3006 represent the signal intensities of four fluorescent images at the center of each community. In the figure, the signal intensities corresponding to the called base sequence 3001 in 3002 to 3006 are significantly higher than those of other bases and are consistent with the reference sequence. Figure 28 Different and the signal strength is also high.

[0234] In particular, in 3004 of the diagram, where a discrepancy occurs in the third cycle, the called base "C" is also more prominent than the other bases. This significant signal intensity suggests that the called base is relatively reliable. Therefore, if the signal intensity of the called base is more prominent than that of the other bases in several cycles before and after the discrepancy, including the cycle where the discrepancy occurs, a mutation occurred in the cycle where the discrepancy occurred, and it is therefore inferred that the base is different from the reference sequence 3000. In other words, the information in the reference sequence 3000 cannot be trusted when such a mutation occurs. Therefore, bases in which such a mutation is detected are removed from the training data.

[0235] As an example of an index indicating how significant the signal intensity of the called base is compared with other bases, the following formula may be used.

[0236] [Formula 5]

[0237] D=I_call / ∑I…(5)

[0238] Here, I_call is the signal intensity of the called base, and the denominator is the sum of the four-color fluorescence intensities I. Such an index D can also be used to determine whether the inconsistent base is a variant.

[0239] As another example, a method of using the likelihood information output by the base prediction unit can be cited. In the base prediction process (S137) in the base prediction unit 803 described in Example 1, Figure 16 In the CNN described in [1], the Softmax part finally outputs the likelihood Yk of each base.

[0240] exist Figure 29 3007-3011 of FIGURE 3 show examples of the likelihood of each base in each cycle. It can be said that the higher the likelihood of a base, the higher the reliability of the base call. Therefore, if the likelihood of the called base is high in several cycles before and after the cycle where the discrepancy occurred, it can be inferred that a mutation occurred in the discrepancy cycle and that the base is different from the reference sequence 3000. Bases with such detected mutations are removed from the training data.

[0241] As described above, in Example 7, when updating training data, the reliability of the base call results is calculated based on information such as the signal intensity and likelihood of the fluorescence image of the aligned bases. This information is then used to determine whether the bases should be added as training data. This improves the quality of the training data during learning and can enhance the prediction accuracy of the base predictor.

[0242] Example 8

[0243] In Example 4, a configuration was described in which the base prediction unit 803 and the learning unit were provided in the nucleic acid analysis device 100. In Example 8, an example of a system configuration in which the nucleic acid analysis device, the base prediction unit, and the learning unit are separated is shown.

[0244] Figures 30A to 30C A plurality of system configuration examples are shown.

[0245] Figure 30A The following system configuration is shown: nucleic acid analysis devices 1 and 2 are equipped with the same base prediction unit, the base prediction unit is learned in the nucleic acid analysis device 2, and the prediction model parameters obtained through learning are sent to the nucleic acid analysis device 1. The training data used during learning uses images captured by the nucleic acid analysis device 2. As a typical application example, the following form is used: a nucleic acid analysis device supplier uses the nucleic acid analysis device 2 it owns to perform measurements to generate prediction model parameters, and then downloads them to the nucleic acid analysis device 1 owned by the user. In cases where the deviation between devices is small, such a configuration example can be applied. As an advantage of this configuration, the user's nucleic acid analysis device 1 does not need to have a learning function, so the device cost can be reduced.

[0246] Figure 30BThe following system structure is shown: the nucleic acid analysis device 1 and the external learning server have the same base prediction unit, and the learning server sends the prediction model parameters obtained by learning the base prediction unit using the image captured by the nucleic acid analysis device 1 to the nucleic acid analysis device 1. As a typical application example, it is as follows: the nucleic acid analysis device supplier provides a computer with only a learning function to the user as a learning server, sends the image captured by the nucleic acid analysis device 1 owned by the user to the learning server, and downloads the prediction model parameters obtained by learning on the server to the nucleic acid analysis device 1. Such a structure is effective when the deviation between devices is so large that it cannot be ignored, and it is preferred to have prediction model parameters for each device. Figure 30A Likewise, since the user's nucleic acid analysis device 1 does not need to have a learning function, the device cost can be reduced. However, network capacity is required to transmit images to the server.

[0247] Figure 30C It is from Figure 30B The structure further transfers the base calling function of the nucleic acid analysis device 1 to an external server. In this structure, the nucleic acid analysis device 1 only captures the colony image, and all subsequent base calling is performed by the external server. Figure 30B This can reduce the device cost, but requires network capacity for transmitting images to the server.

[0248] As described above, in Example 8, by setting a system structure in which the nucleic acid analysis device, base prediction unit, and learning unit are divided, the nucleic acid analysis device, base prediction processing function, and learning processing function provided to users can be reduced in cost.

[0249] Example 9

[0250] In Example 9, several examples of user interfaces in the previously described embodiments are shown. These user interfaces are Figure 6 However, it can also be presented through the monitor screen of an external computer and peripheral devices such as a mouse and a keyboard.

[0251] (1) Selection of multiple predictors

[0252] Figure 31 The structure using multiple base predictors described in Example 3 is shown ( Figure 18) shows an example of a screen for selecting multiple base predictors. This figure shows an example of a screen in which a list of multiple prediction model parameters already learned within the nucleic acid analyzer is presented to the user, from which the user can select a base predictor for prediction. This figure shows, as an example, information such as the learning accuracy and creation date of each prediction parameter. However, this is not necessarily limited to these; various information serving as a reference for selecting a prediction parameter may also be presented.

[0253] (2) Setting the prediction model for each cycle

[0254] Figure 32 The structure using multiple base predictors described in Example 3 is shown ( Figure 18 ) is an example of a setting screen when distinguishing the use of a base predictor for each cycle. In the figure, different prediction models are set at intervals of 50 cycles, and the combination of the prediction models is defined as a new prediction model. As described in Example 2, as the cycle becomes larger, the influence of the attenuation between cycles will accumulate. Therefore, the characteristics of the camera image change according to the number of cycles. Therefore, there is an effective possibility of switching the prediction model for base calling according to the cycle. The example of switching the model for every 50 cycles is shown in the figure, but it can also be further refined, for example, it can also be switched according to each cycle. However, in the model used in each cycle, it is preferably to use the training data obtained in the cycle used.

[0255] (3) Selection of a Data Set to Use in Learning a Base Prediction Model

[0256] Figure 33 An example of a setting screen for adding or deleting a data set for learning a new prediction model or an existing prediction model is shown. The data set for learning can be image data stored in the storage unit 809 of the nucleic acid analyzer or stored in an external computer or storage device.

[0257] exist Figure 33 In the example shown in FIG, a list of data sets to be added to the learning process for a selected prediction model is displayed, prompting the user to select a data set. By pressing the "Add" button, the selected data set can be added.

[0258] The same screen can also display a list of data sets already used for learning the selected prediction model, prompting the user to select image data to be removed from the learning target during retraining. In this case, the "Delete" button is used instead of the "Add" button to delete the selected data set. This also allows the user to change the image data set used for learning based on an existing prediction model and save the prediction model under a new name using the file name setting dialog (not shown).

[0259] Furthermore, the learning parameters described in Example 4 and later can be set for each image data using a learning setting screen (not shown). Several examples of such learning parameter setting items are listed below. However, the present invention is not necessarily limited to these, and various parameters related to the base prediction method described in this embodiment can also be set by the user.

[0260] (A) Application cycle range

[0261] Setting the range of the cycle of image data used for learning: As described above, the influence of attenuation varies depending on the cycle, so it is effective to be able to select which image to use in which cycle.

[0262] (B) Application FOV

[0263] Specifies which FOV of the image data to use. This is because the image characteristics may vary depending on the position of the FOV due to influences such as distortion of the flow channel chip. Even if the entire FOV is not used, it is useful to limit the FOV to a specific FOV for purposes such as learning.

[0264] (C) Input ROI size and number of forward and backward cycles

[0265] The setting may be changed for each prediction model by taking into account the degree of focus of the image and the influence of attenuation.

[0266] (D) CNN network structure

[0267] You can also change the settings of known CNN settings such as the number of network layers, the type of activation function, the presence or absence of a pooling layer, the learning rate, the number of traversals, and the number of mini-batches for each prediction model.

[0268] (E) Choice of additional learning or new learning

[0269] When adding training data to update the base prediction model, it is also possible to set the prediction model to be updated using the immediately previous base prediction model as the initial value, or to reset the prediction model to create a new initial value and learn all the training data again.

[0270] (F) Setting of training data screening

[0271] The filtering conditions when updating the training data described in Example 7 are set, such as the reliability threshold, likelihood threshold, and signal strength threshold used to determine whether the data is not included in the training data.

[0272] The nucleic acid analysis devices and base calling methods described in the above-described embodiments can be used to detect various nucleic acid reactions and perform nucleic acid analysis, such as DNA sequencing. The present invention is not limited to the above-described embodiments but encompasses various variations. For example, the above-described embodiments are provided to facilitate understanding of the present invention and are not intended to necessarily include all of the described configurations. While the nucleic acid analysis devices in the above-described embodiments target DNA fragments for measurement and analysis, they can also target other biologically relevant substances, such as RNA.

[0273] Furthermore, the above-described structures, functions, computers, and the like have been described primarily with the example of creating a program that implements some or all of them. However, it is apparent that some or all of them can also be implemented in hardware, for example, by designing an integrated circuit. In other words, instead of a program, all or part of the functions of the processing unit can be implemented using an integrated circuit such as an ASIC (Application Specific Integrated Circuit) or an FPGA (Field Programmable Gate Array).

[0274] Industrial applicability

[0275] It can be used in nucleic acid analysis for measuring bio-related substances.

Claims

1. A nucleic acid analysis device, characterized in that: have: a base prediction unit configured to perform base prediction using as input a plurality of images obtained by detecting luminescence from DNA disposed on a substrate; a positioning unit for aligning the plurality of images with respect to a reference image; as well as an extraction unit that extracts bright spots from the plurality of images, The plurality of images are images obtained by respectively detecting, by sensors, a plurality of types of luminescence from a plurality of types of fluorescent bodies introduced into the DNA, wherein for each of the plurality of types of luminescence, at least one of the sensors performing the detection and the optical paths to the sensors performing the detection are different, The extraction unit performs a predetermined threshold determination on the image to extract the bright spot. The base prediction unit is composed of a neural network capable of supervised learning, and only the pixels in a predetermined range centered on the position of the extracted bright spot in the multiple images are input as an ROI image. The neural network is used to extract the feature value of the ROI image, and the base is predicted based on the feature value.

2. The nucleic acid analysis device according to claim 1, characterized in that The plurality of images include images acquired in different cycles of a cycle repeated a plurality of times for sequentially determining each base of a base sequence, The base prediction unit receives as input the ROI image of at least one cycle selected from the previous cycle and the next cycle, in addition to the ROI image of the cycle in which prediction is performed.

3. The nucleic acid analysis device according to claim 1, wherein The nucleic acid analysis device includes a plurality of base prediction units to which different base prediction parameters are set as the base prediction units, and the nucleic acid analysis device predicts bases based on prediction results of the plurality of base prediction units.

4. A nucleic acid analysis method for performing base prediction using a plurality of images obtained by detecting luminescence from DNA as input to a base prediction device, characterized in that: The nucleic acid analysis method performs a colony position determination stage and a base sequence determination stage, The plurality of images are images obtained by detecting, by sensors, a plurality of types of luminescence from a plurality of types of fluorescent substances emitted from a colony serving as a DNA sample fixed on a substrate, wherein for each of the plurality of types of luminescence, at least one of the sensor performing the detection and the optical path to the sensor performing the detection is different. The base predictor is composed of a neural network capable of supervised learning, which extracts features from the input image and classifies the bases based on the features. During the group location determination phase, the following steps are performed: Positioning processing, performing position alignment of the multiple images; as well as The colony position determination process extracts bright spots by performing a predetermined threshold determination on the plurality of images to determine the colony position of the DNA. In the base sequence determination stage, only pixels in a predetermined range centered on the position of the extracted bright spot representing the community within the multiple images are input to the base predictor as an ROI image, the feature value of the ROI image is extracted, and the base is predicted based on the feature value.

5. The nucleic acid analysis method according to claim 4, characterized in that The plurality of images include images acquired in different cycles of a cycle repeated a plurality of times for sequentially determining bases, In the base sequence determination stage, a group of a plurality of the ROI images captured at temporally different timings in the different cycles is input to the base predictor.

Citation Information

Patent Citations

  • Substrate for nucleic acid analysis, flow cell for nucleic acid analysis, and analysis method

    JP2020000060A

  • Method and system for recognizing base in nucleic acid

    CN113012757A

  • Basecaller for DNA sequencing using machine learning

    US20150169824A1

  • Machine learning enabled pulse and base calling for sequencing devices

    US20190237160A1

  • Luminescence image coding device, luminescence image decoding device, and luminescence image analysis system

    WO2017203679A1