High density sample gene sequencing method

By employing large-area, wide-beam illumination and imaging techniques, combined with fluorescence signal decoupling and algorithm analysis, the problem of inaccurate sequencing results in high-density samples was solved, achieving high throughput and high precision in high-density gene sequencing.

CN115985392BActive Publication Date: 2026-03-31SHANGHAI JIAOTONG UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-02-02
Publication Date
2026-03-31

AI Technical Summary

Technical Problem

Existing DNA sequencing methods are affected by the diffraction limit of optical imaging systems and neighboring signals under high-density sample conditions, resulting in inaccurate sequencing results and failing to effectively improve sequencing throughput.

Method used

By employing large-area, wide-beam illumination and imaging technology, and through the acquisition and decoupling of fluorescence signals arranged in a two-dimensional array, combined with Gaussian distribution function fitting and orthogonal matching pursuit algorithm, the spatial distribution and signal intensity of sample points are calculated to achieve high-density gene sequencing.

Benefits of technology

Without changing the traditional laser illumination imaging method, the sequencing density was significantly increased to 16000K/mm2, the read error rate was kept below one ten-thousandth, and the sequencing throughput was increased by 10 times.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115985392B_ABST
    Figure CN115985392B_ABST
Patent Text Reader

Abstract

The application discloses a high-density sample gene sequencing method, which comprises the following steps: collecting A, C, T and G four-channel fluorescence collection images of high-density sequencing samples arranged in a two-dimensional dot matrix respectively, determining the diffraction spot center position of each fluorescence signal, calculating the overall spatial position distribution of the sample dot matrix, determining the fluorescence signal calculation image of each sample dot under the condition that the sample dot exists signal alone during imaging, decoupling the coupling image according to the fluorescence signal image, obtaining the signal strength of each sample dot, analyzing the signal strength of each sample dot in different wavelength channels during each imaging, and determining the base arrangement sequence of nucleic acid in each sample dot. On the basis of the existing sequencing instrument, the sample distribution interval can be reduced to 0.45 times of the imaging system resolution, and the reading error rate is maintained at 1 / 10,000. When the sample interval is reduced to 0.36 times of the imaging system resolution, the error rate is maintained at 1 / 1,000.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to a technology in the biomedical field, specifically a technology with a density greater than 16000 K / mm². 2 Gene sequencing methods. Background Technology

[0002] Current DNA sequencing methods segment fluorescence images to obtain the signal intensity at the location of the sequencing signal. This method has certain requirements on the density of the sequencing sample. Due to the diffraction limit of optical imaging systems, when the density of the sequencing sample is high, different base signals can couple together, affecting the sequencing results. In nucleic acid sequencing based on fluorescence imaging, the diffraction limit of the optical system and the influence of neighboring signals limit the reduction of distance between each sample point and the increase of sample point density, thus hindering further improvements in sequencing throughput. Summary of the Invention

[0003] This invention addresses the problem of existing gene sequencing technologies' inability to read signals from high-density samples. It proposes a high-density sample gene sequencing method that utilizes large-area, wide-beam illumination and imaging to perform real-time computational analysis on high-density samples, achieving high-speed, high-throughput gene sequencing. Based on existing sequencing instruments, the sample distribution spacing can be reduced to as low as 0.45 times the imaging system resolution while maintaining a read error rate of 0.01%. Further reduction to 0.36 times the imaging system resolution maintains an error rate of 0.1%.

[0004] This invention is achieved through the following technical solution:

[0005] This invention relates to a high-density sample gene sequencing method. It involves acquiring four-channel fluorescence images (A, C, T, G) of a high-density sequencing sample arranged in a two-dimensional array, one for each channel, and determining the center position of the diffraction spot for each fluorescence signal. This allows for the calculation of the overall spatial distribution of the sample array and the determination of the fluorescence signal image of each sample point when it exists independently during imaging. The coupled images are then decoupled based on these fluorescence signal images to obtain the signal intensity of each sample point. By analyzing the signal intensity of each sample point in different wavelength channels during each imaging session, the base sequence of the nucleic acids in each sample point is determined.

[0006] The aforementioned two-dimensional dot array arrangement refers to: fabricating a spatially distributed array through microfabrication, where each dot in the array is spaced at a fixed distance, and each dot in the array has a groove of a certain radius. An imaging sample containing a DNA template to be sequenced is prepared, and through vapor deposition, the nucleic acid to be sequenced is amplified only within the groove.

[0007] The four-channel fluorescence acquisition image is acquired by setting a prism on one side of the two-dimensional dot array to reflect the incident laser, setting a lens with a numerical aperture greater than 0.5 on the other side of the two-dimensional dot array, and setting an image sensor behind the lens and the lens barrel to obtain a coupled wide-field image of the fluorescence signal associated with the sequence information of each sample point in the two-dimensional dot array.

[0008] The center position of the diffraction spot for each fluorescence signal is obtained as follows: A coupled wide-field image of fluorescence signals with sequential information is preprocessed; then, several pixels to be fitted are obtained by calculating local maxima and thresholding; finally, the point spread function of the diffraction spots is fitted using a Gaussian distribution function to obtain the position coordinates (x, y) of each diffraction spot. i y i And the size of the diffraction spot σ x , σ y .

[0009] Preferably, by repeatedly fitting each candidate point in the image and comparing the size of the diffraction spot generated by a single sample point, the center positions of the diffraction spots coupled by multiple sample points are filtered out, resulting in a series of coordinates (x1, y1), (x2, y2)...(x n y n ).

[0010] The preprocessing includes: mean filtering, Gaussian filtering, low-pass Gaussian filtering, difference of Gaussian function filtering, wavelet filtering, and mean filtering.

[0011] The Gaussian distribution function is: Where: I represents the signal of the four-channel fluorescence acquisition image, I i To calculate the intensity of the image for the fluorescence signal, x i y i σ represents the center position coordinates of the fitted signal. x , σ y The values ​​of the fitted light spot are defined on the x and y axes. These parameters are estimated using either the least squares method or maximum likelihood estimation.

[0012] The decoupling refers to estimating the position of the first point in the field of view based on the center position of the diffraction spot, thereby generating the base coordinate matrix A of the spatial distribution of the sequencing sample, and using the orthogonal matching pursuit algorithm to optimize the problem min||y-Ax||, where: y is the four-channel fluorescence acquisition image obtained in step B, A is the base coordinate matrix, and x is the signal intensity of each position of the final calculated point array, and the component under each coordinate is the signal intensity of the corresponding point of the sample.

[0013] The position of the first point in the field of view refers to the position of the sample point closest to the origin of the image coordinates for calculating the fluorescence signal. Generally, based on the definition of image pixels, the first point is selected from the upper left corner of the image; the coordinates P of the point based on the center position of the diffraction spot. i =x i y i The coordinates of the first point in the field of view satisfy the following: Where: n i m i For positive integers, taking the x-coordinate as an example, the modulo operation on the equation has (x... i -x0)mod d=(n i ·d)mod d, i.e. x i Since mod d = x0, the sequence of located points is converted into an estimate of the first point in the field of view, i.e., P′. i =(x i mod d, y i mod d), calculate the variance of the overall estimate Q = Σ(P0 - P′) i ) 2 By minimizing the overall variance, we can obtain the minimum variance estimate P0 = (P′1 + P′2 + ... + P′) for the first point in the field of view. n ) / n; After determining the position of the first point in the field of view, P0 = (x0, y0), the positions of the remaining points can be calculated using the parameters of the dot matrix: 0 ≤ x0, y0 < d, where d is the spacing used in the dot matrix design.

[0014] The aforementioned base coordinate matrix refers to a one-dimensional vector formed by connecting the end of the previous row and the beginning of the next row of a two-dimensional image, where each column of the base coordinate matrix is ​​a wide-field fluorescence imaging image calculated when each sample point has its own signal. Each column of the base coordinate matrix is ​​a calculated image of the fluorescence signal obtained through simulation.

[0015] The aforementioned orthogonal matching pursuit specifically includes:

[0016] i) Perform an inner product calculation on each column vector of the base coordinate matrix of the one-dimensional fluorescence signal calculation image of each sample point when the signal exists independently during imaging, after one-dimensional processing of the actual coupled fluorescence image of each channel;

[0017] ii) Obtain the value of the signal point with the largest inner product as an estimate of the signal strength of the corresponding point, and subtract the estimated signal from the acquired coupled image signal;

[0018] iii) When the estimated number of sample points is greater than or equal to two, the estimated signal intensity is corrected using the least squares method. Specifically, the optimization equation is: min||A new λ-y||, where Anew Let y be the matrix consisting of the vectors corresponding to the base coordinate matrices from steps i and ii, where y is the estimated signal strength. The least squares method will be used to estimate λ to correct the signal strength estimate of the sample point location in step ii.

[0019] iv) Iterate through steps i to iii until the iterations exceed a certain number or the coupled image approaches 0, then stop.

[0020] The signal intensity analysis refers to determining the base sequence of nucleic acids in each sample point by analyzing the signal intensity of different wavelength channels during each imaging session. The base channel with the strongest signal intensity is used to determine the base category of the measured base. If the signal intensities of the four channels are relatively similar, the signal from that point is discarded, indicating a reading error.

[0021] Technical effect

[0022] This invention utilizes the sparsity of four-channel fluorescence images in sequencing to locate the actual spatial distribution of the sample matrix. By leveraging this spatial distribution information, high-density sample signals are decoupled, thereby increasing sequencing density. Without altering the principles of traditional laser illumination imaging, this invention achieves a 10-fold increase in sequencing density, reaching 16000 K / mm². 2 sequencing density. Attached Figure Description

[0023] Figure 1 This is a flowchart of the present invention;

[0024] Figure 2 This is a schematic diagram of the dot matrix in the embodiment;

[0025] Figure 3 This is a local fluorescence imaging image of the four channels A, C, T, and G;

[0026] Figure 4 This includes the possible offsets and deflections of the sample;

[0027] Figure 5 This is a local result of Gaussian fitting of the fluorescence image spot;

[0028] Figure 6 A partial diagram of the base coordinate matrix.

[0029] Figure 7 Calculation diagram of the corresponding position signals of the four-channel dot matrix A, C, T, and G;

[0030] Figure 8 The results of comparing A, C, T, and G to determine the final base category;

[0031] Figure 9 This is a schematic diagram of the imaging process for an example. Detailed Implementation

[0032] This embodiment relates to an imaging system with a resolution of 550 nanometers, and the sample distribution lattice is as follows: Figure 2 As shown, the center-to-center distance of the samples is 230 nanometers, and the sample radius is 80 nanometers. The system image obtained through the imaging module is as follows. Figure 3 The fluorescence signal acquisition image is shown. Based on the above parameters, the sample density is calculated to be greater than 16000 K / mm². 2 .

[0033] like Figure 1 As shown in this embodiment, a high-density sample gene sequencing method includes the following steps:

[0034] Step 1: Acquire images using four fluorescence channels: A, C, T, and G;

[0035] Step 2: Use Gaussian fitting to determine the position and size of the light spot;

[0036] Step 3: Estimate the position of the first point in the field of view based on the position of the light spot;

[0037] Step 4: Generate the base coordinate matrix of the point distribution;

[0038] Step 5: Calculate the signal strength of the A, C, T, and G channel dot matrix positions using the orthogonal matching pursuit algorithm;

[0039] Step 6: Compare the signal intensities of the four channels A, C, T, and G to determine the base type at the corresponding position in the matrix.

[0040] like Figure 4 As shown, this is the sample deflection method in this embodiment. After the imaging system is properly adjusted, the horizontal and vertical deflection angles are approximately 0 degrees. Therefore, only the magnitude of the horizontal offset needs to be considered, i.e., the position of the first point in the field of view. Figure 5 The image localization process yields a series of spot center coordinates. By analyzing the selected spot center positions, the coordinates of the first point in the field of view are estimated: P0 = (P′1 + P′2 + ... + P′). n ) / n.

[0041] like Figure 6 As shown, the base coordinate matrix A is generated based on the coordinates of the first point in the field of view and the dot matrix spacing. The signal strengths of the four channels A, C, T, and G are iteratively calculated using the generated base coordinate matrix. Figure 7 As shown.

[0042] like Figure 8 As shown, the final base class is calculated based on the signals at the same lattice positions in the four channels A, C, T, and G.

[0043] like Figure 9 As shown, this embodiment achieves imaging in the following way: In the imaging module, the laser is reflected by a prism to achieve uniform illumination over a large area while maintaining the laser power at the output power of a conventional laser. The laser beam is wide and each beam is placed close together. The laser illuminates the sample via total internal reflection, exciting the sample and causing it to emit fluorescence. This fluorescence passes through a high numerical aperture and wide field-of-view lens and a fluorescence filter before being detected by the detector. The system corrects for phase aberration in the lens using a telescope lens. The detector can be a CCD or a CMOS sensor.

[0044] By using actual parameters of the optical detection system, with the system resolution set to 550 nanometers, and simulating different sample spacings, different imaging pixel sizes were adjusted to perform simulation calculations to obtain the error rate of sequencing using the method described in this invention. Figure 9 As shown, when the spacing between adjacent sample points is reduced to 0.4 times the resolution of the imaging system, the calculated sequencing read error rate is still less than one ten-thousandth. Even when the sample spacing is reduced to 0.36 times the resolution of the imaging system, the sequencing error rate remains less than one thousandth.

[0045] Table 1 shows the relationship between sample spacing, pixel size, and read error rate.

[0046]

[0047] Compared with existing technologies, this method achieves a sample density greater than 16000 K / mm² without changing the traditional laser illumination imaging method. 2 This gene sequencing task represents an order of magnitude improvement over existing sequencing densities.

[0048] The above-described specific implementations can be partially adjusted by those skilled in the art in different ways without departing from the principles and purpose of the present invention. The scope of protection of the present invention is defined by the claims and is not limited to the above-described specific implementations. All implementation schemes within the scope of the claims are bound by the present invention.

Claims

1. A high density sample gene sequencing method, characterized by, The overall spatial position distribution of the sample array is calculated by collecting A, C, T and G four-channel fluorescence images of the high-density sequencing sample arranged in a two-dimensional dot array respectively and determining the diffraction spot center position of each fluorescence signal, so as to obtain the fluorescence signal calculation image of each sample point in the case of existing signal alone during imaging, and then decoupling the coupled image according to the fluorescence signal image to obtain the signal intensity of each sample point. The high density refers to that when the density of the sequencing sample is large, different base signals will be coupled together, which will affect the sequencing result. The center position of each diffraction spot of the fluorescent signal is obtained by pre-processing the coupled wide-field image of the fluorescent signal associated with the sequence information, then calculating the local maximum points of the image and screening the points by thresholding to obtain a plurality of pixel points to be fitted, and then fitting the point spread function of the diffraction spot by using a Gaussian distribution function to obtain the position coordinates of each diffraction spot And the size of the diffraction spot By repeatedly fitting each candidate point in the image and comparing the size of the diffraction spot generated by a single sample point, the center positions of the diffraction spots coupled by multiple sample points are filtered out to obtain a series of position coordinates ; The decoupling refers to: according to the estimation of the center position of the diffraction spot, the position of the first point in the field of view is generated, so as to generate the base coordinate matrix A of the spatial distribution of the sequencing sample, and the orthogonal matching pursuit algorithm is used to optimize the problem Wherein: y is a four-channel fluorescence acquisition image, A is a base coordinate matrix, x is the intensity of the final calculated point array position signal, and the component under each coordinate, that is, the signal intensity of the corresponding point of the sample, is obtained.

2. The high density sample genetic sequencing method of claim 1, wherein, The two-dimensional dot array arrangement refers to that the array with spatial distribution is prepared by microprocessing, the distance between each point in the array is fixed, a certain radius groove is processed on each point of the array, and the imaging sample with the DNA template to be sequenced is prepared, and the nucleic acid to be sequenced will only be amplified in the groove by the gas deposition method.

3. The high density sample genetic sequencing method of claim 1, wherein, The four-channel fluorescence acquisition image is obtained by the following method: a prism is arranged on one side of the two-dimensional dot array for reflecting the incident laser, a lens with a numerical aperture greater than 0.5 is arranged on the other side of the two-dimensional dot array, and an image sensor is arranged behind the lens and the lens barrel to obtain the coupled wide-field image of the fluorescence signal of each sample point of the two-dimensional dot array and its sequence information.

4. The high density sample genetic sequencing method of claim 1, wherein, The preprocessing includes mean filtering, Gaussian filtering, low-pass Gaussian filtering, Gaussian difference function filtering, wavelet filtering and mean filtering. The Gaussian distribution function is: wherein I is a signal of a four-channel fluorescence acquisition image, calculating the intensity of the image for the fluorescence signal, a coordinate of a center position of the fitted signal, a size of the fitted spot on the x-axis and the y-axis, and the estimation of the above parameters is realized by a least square method or a maximum likelihood estimation.

5. The high density sample genetic sequencing method of claim 1, wherein, The position of the first point in the field of view refers to the position of the sample point closest to the coordinate origin of the fluorescence signal calculation image, and the first point in the upper left corner of the image is selected according to the definition of the image pixel. Point coordinates based on diffraction spot center position The coordinates of the first point in the field of view satisfy: Wherein: is a positive integer, and the modulo operation of the equation with the x coordinate is That is, Therefore, the sequence of the located points is converted into an estimate of the first point in the field of view, that is, The overall estimated variance is calculated The minimum variance estimate of the first point in the field of view is obtained by minimizing the overall variance When the position of the first point in the field of view is determined The positions of the remaining points are calculated through the parameters of the point array: Wherein: d is the distance adopted in the design of the point array.

6. The high density sample genetic sequencing method of claim 1, wherein, The base coordinate matrix refers to that each column in the base coordinate matrix is the wide-field fluorescence imaging image calculated in the case of each sample point existing signal alone, and the one-dimensional vector formed by connecting the tail of the upper row and the head of the lower row of the two-dimensional image, and each column in the base coordinate matrix is the fluorescence signal calculation image obtained by simulation calculation.

7. The high density sample genetic sequencing method of claim 1, wherein, The orthogonal matching pursuit specifically includes: i) inner product calculation is performed between each channel of the actual collected coupled fluorescence image after one-dimensional processing and each column vector in the base coordinate matrix which is the one-dimensional fluorescence signal calculation image of each sample point in the case of existing signal alone during imaging; ii) the value of the signal point with the maximum inner product is taken as the estimation of the signal intensity of the corresponding point, and the estimated signal is subtracted from the collected coupled image signal; iii) when the estimated number of sample points is greater than or equal to two, the estimated signal intensities are corrected using least squares, by solving the equation , where is a matrix composed of the base coordinate matrix corresponding vectors of the estimated signal positions from steps i and ii, and y is the estimated signal intensities; and the estimated signal intensities from the least squares estimation are used to correct the estimation of the signal intensities at the sample point positions from step ii. iv) the above steps i to iii are iterated until the iteration is greater than a certain number of steps or the coupled image tends to 0.

8. The high density sample genetic sequencing method of claim 1, wherein, The signal intensity analysis refers to that the base arrangement order of the nucleic acid in each sample point is determined by analyzing the signal intensity of each sample point in each imaging at different wavelength channels, the base category of the measured base is determined by comparing the intensity of the four-channel signal and the base channel of the signal with the maximum intensity, and if the signal intensity of the four channels is small, the signal of this point is discarded, and this point is regarded as a waste point, indicating a reading error.

Citation Information

Patent Citations

  • Base fluorescence image capturing system device and method for high-flux genome sequencing

    CN105039147A

  • Efficient nucleic acid detection and gene sequencing method and device thereof

    CN111926065A