Polygene repetitive sequence sequencing method based on fluorescence imaging

By using a fluorescence imaging-based multi-gene repetitive sequence sequencing method, the problems of short read length, low accuracy, high signal-noise ratio, and poor probe specificity in existing multi-gene repetitive sequence sequencing technologies have been solved, achieving efficient and accurate multi-gene repetitive sequence sequencing.

CN121065314APending Publication Date: 2025-12-05SHANGHAI LINGEN BIOTECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511302302.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-09-12
Publication Date
2025-12-05

AI Technical Summary

Technical Problem

Existing sequencing methods suffer from problems such as short read lengths, low accuracy, high signal-to-noise ratios, high costs, weak fluorescent labeling signals, and poor probe specificity in multi-gene repetitive sequence sequencing, making it difficult to meet the demand for accurate sequencing of multi-gene repetitive sequences.

Method used

A fluorescence imaging-based multi-gene repetitive sequence sequencing method is employed, including sample targeted preprocessing, fluorescent probe design, in situ hybridization reaction, multilayer fluorescence imaging acquisition, fluorescence signal analysis and sequence identification, and sequence verification and assembly. By accurately acquiring target samples, designing multiple sets of fluorescent probes, stable hybridization and high-quality imaging, accurate signal analysis and rigorous verification and assembly, efficient and accurate multi-gene repetitive sequence sequencing is achieved.

Benefits of technology

It significantly improves sequencing efficiency and accuracy. By acquiring high-purity, high-concentration target DNA fragments, multilayer fluorescence imaging, and sequence verification assembly, it solves the sequencing challenges of multi-gene repetitive sequences and achieves high-resolution, low-interference multi-gene repetitive sequence sequencing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121065314A_ABST
    Figure CN121065314A_ABST
Patent Text Reader

Abstract

The invention discloses a multi-gene repetitive sequence sequencing method based on fluorescence imaging, and belongs to the technical field of gene sequencing. The method comprises the following steps: sample targeting pretreatment, fluorescent probe synthesis, in-situ hybridization reaction, multi-layer fluorescence imaging acquisition, fluorescence signal analysis and sequence recognition, and sequence verification and splicing, and the sample targeting pretreatment is subjected to microdissection, enzyme digestion and magnetic bead enrichment to obtain high-purity and high-concentration target DNA. The design of the fluorescent probe is guaranteed by multiple groups of fluorescent labels and high purity, in-situ hybridization enhances the combination efficiency by optimizing fixing and washing conditions, a light path and an imaging interval are accurately controlled by multi-layer fluorescent imaging, a high-resolution and low-interference image is obtained, complete signal information is reserved, accurate conversion from the image to a sequence is realized by fluorescent signal analysis, and the detection accuracy is improved. Sequence verification splicing is subjected to secondary hybridization verification and overlapping region integration, short segment limitation is broken through, the final sequence accuracy is improved, and the multi-gene repetitive sequence sequencing requirement is comprehensively met.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of gene sequencing, in particular to a multi-gene repeat sequence sequencing method based on fluorescence imaging. BACKGROUND

[0002] Multi-gene repeat sequence is a DNA fragment formed by repeating arrangement of specific base units in the genome, and the number of repeats and arrangement pattern are closely related to genetic diseases, species evolution and gene expression regulation.

[0003] At present, the commonly used sequencing methods have obvious limitations in multi-gene repeat sequence sequencing: Sanger sequencing has a short read length (usually ≤1000bp), which is difficult to cover long repeat sequences, and the sequencing accuracy of high repeat regions is low; NGS can realize high-throughput sequencing, but it relies on PCR amplification, which is prone to amplification bias due to high GC content or secondary structure of repeat sequences, resulting in loss of sequence information; the third generation single molecule sequencing does not require amplification, but has high signal noise, insufficient ability to distinguish short repeat units (such as 3-6nt), and high sequencing cost.

[0004] In addition, although the traditional fluorescence in situ hybridization technology can realize in situ sequence detection, it has problems such as weak fluorescent label signal, poor probe specificity, and difficulty in distinguishing multiple repeat units, which is difficult to meet the precise sequencing requirements of multi-gene repeat sequences. SUMMARY

[0005] The purpose of the present application is to provide a multi-gene repeat sequence sequencing method based on fluorescence imaging to solve the problems raised in the background art.

[0006] To achieve the above purpose, the present application provides the following technical solution: a multi-gene repeat sequence sequencing method based on fluorescence imaging, comprising: Targeted pretreatment of samples, selecting biological samples containing target multi-gene repeat sequences, obtaining cell populations containing target sequences, releasing genomic DNA in the cell populations and performing enzyme digestion, obtaining DNA fragments containing target repeat sequences, enriching the DNA fragments, and obtaining enriched target DNA fragments; Synthesis of fluorescent probes, designing multiple sets of fluorescent probe groups for characteristic repeat units of target multi-gene repeat sequences and synthesizing the fluorescent probe groups by chemical synthesis method; In situ hybridization reaction, fixing target DNA fragments on a glass slide, dehydrating the DNA fragments to keep them in a linearly stretched state, adding a fluorescent probe group mixture to the glass slide with fixed DNA fragments, performing in situ hybridization reaction, and washing the glass slide with a washing solution to remove unbound free probes after hybridization is completed; Multi-layer fluorescence imaging acquisition, the hybridized slide is placed under a fluorescence microscope, and a laser confocal scanning method is used for imaging to obtain a multi-channel fluorescence image sequence; Fluorescence signal analysis and sequence identification, mark the effective signal region, extract the position information and intensity information of each fluorescence signal in the multi-channel fluorescence image sequence, obtain the distribution position of the fluorescence signal on the DNA fragment, and determine the arrangement order of the repeat units of the preliminary identified sequence; Sequence verification and splicing, designing verification probes for the preliminary identified sequence, performing secondary hybridization and collecting fluorescence images after secondary hybridization, comparing the signal distribution of the fluorescence images after secondary hybridization with the preliminary identified sequence, obtaining the sequencing result, splicing the sequencing results of multiple target DNA fragments and integrating the overlapping sequence parts to obtain a complete multi-gene repeat sequence.

[0007] Further, sample targeted pretreatment, specifically including: Selecting a biological sample containing a target multi-gene repeat sequence, separating and obtaining a cell group containing a target sequence by microdissection; Placing the cell group in a hypotonic solution and incubating at room temperature for 10-15 minutes to release genomic DNA by breaking the cells, wherein the hypotonic solution is a 0.075 mol / L potassium chloride solution; Performing enzyme digestion on the released genomic DNA using a restriction endonuclease, wherein the recognition site of the restriction endonuclease is located in the flanking conserved region of the target repeat sequence, the concentration of genomic DNA in the enzyme digestion system is 50-100 ng / μL, and the amount of restriction endonuclease used is 10-20 U / μg DNA; Incubating the enzyme digestion system at 37°C for 1-2 hours to obtain DNA fragments containing the target repeat sequence; after the enzyme digestion reaction is completed, adding an enzyme inactivator with a final concentration of 0.5 mol / L EDTA to terminate the reaction, and placing it in a 25°C environment for 5 minutes; Taking magnetic beads with single-stranded oligonucleotides complementary to the flanking conserved region of the target repeat sequence on the surface, washing the magnetic beads twice with a binding buffer; mixing the enzyme-digested DNA fragment solution with the magnetic beads, and performing a 30-minute constant-temperature oscillation reaction at 25°C to allow the target DNA fragments to bind to the surface of the magnetic beads through base complementary pairing; Using a magnetic separation frame to adsorb the magnetic beads, discarding the supernatant, washing the magnetic beads with a washing buffer, adding an elution buffer, incubating at 65°C for 5 minutes, collecting the supernatant after magnetic separation, and obtaining the enriched target DNA fragments.

[0008] Further, each group of fluorescence probes includes at least 3 probes with different fluorescent labels, and different fluorescent labels correspond to different repeat unit base combinations; The fluorescent probe is 15-25 nt in length, and its sequence is complementary to the repeating unit of the target repeating sequence. The fluorescent label is located at the 5' end of the probe, and fluorescent quantum dots are used as fluorescent labels. Different fluorescent quantum dots have distinguishable emission wavelengths. During the synthesis of the fluorescent probe set, the purity of the probe was controlled to be ≥95%, and the synthesized probe was purified by high performance liquid chromatography.

[0009] Further, in situ hybridization reactions specifically include: The target DNA fragment was transferred onto a glass slide and fixed with 4% paraformaldehyde fixative at room temperature for 10-15 minutes. After fixation, the slides were subjected to gradient dehydration treatment with 70%, 85%, and 100% ethanol solutions by volume, with each concentration of ethanol solution being treated for 2 minutes to keep the DNA fragments in a linear extended state. After treatment, the slides were allowed to air dry naturally. The fluorescent probe mixture was dropped onto the hybridization region of a glass slide immobilized with the target DNA fragment. The concentration of each probe in the probe mixture was 0.5-1 μmol / L. Cover with a coverslip and seal the edges, then place the slide in a hybridization apparatus to perform in situ hybridization. The in situ hybridization reaction procedure is as follows: denature at 95℃ for 5 minutes to break up the DNA fragment, then cool down to 60℃ for constant temperature hybridization for 12-16 hours to allow the fluorescent probe to bind complementary to the target DNA fragment; After the hybridization reaction was completed, the coverslip was removed, and the slide was placed in citrate buffer containing 0.1% sodium dodecyl sulfate and washed three times at 42°C for 5 minutes each time. Unbound free probes and non-specifically bound probes are removed by washing. After washing, the slide is quickly rinsed with ultrapure water and dried with nitrogen to obtain a slide that has completed hybridization.

[0010] Further, multi-layer fluorescence imaging acquisition specifically includes: Place the hybridized slide on the stage of the fluorescence microscope and adjust the stage position so that the sample area is in the center of the microscope's optical path. Based on the emission wavelength of the fluorescent label used, select a matching multi-channel filter group and install it in the microscope optical path, turn on the microscope light source, adjust the excitation light intensity to the preset initial value, and preheat the microscope system. Adjust the focus using the microscope eyepiece or imaging software to ensure the target DNA fragment is in a clear imaging field, and record the coordinates of the imaging field. Start the laser confocal scanning system, set the scanning resolution to 1024×1024 pixels, and the scanning speed to 1-2 frames per second; The multiple fluorescence imaging is sequentially performed on the same imaging field of view, each imaging corresponding to an excitation wavelength of a fluorescence label, and the fluorescence signal intensity is monitored in real time during the imaging process; The time interval between the adjacent two fluorescence imaging is 10 seconds, the next imaging is started after the signal of the previous fluorescence label is stable, and the image acquisition of all fluorescence channels is sequentially completed to obtain a multi-channel fluorescence image sequence; The acquired multi-channel fluorescence image sequence is named according to the channel type and field of view coordinates and stored in a special image database.

[0011] Further, the fluorescence signal analysis and sequence identification specifically includes: The position information and intensity information of each fluorescence signal in the effective signal region are extracted; The position information is determined by a signal positioning algorithm, and the relative position coordinates of the signal in the extension direction of the DNA fragment are calculated based on the pixel coordinates of the signal center point; The intensity information is calculated by integrating the total fluorescence intensity value in the signal region; The distance between the adjacent two fluorescence signals is also calculated, and the distance calculation is based on the straight-line distance between the signal center points; A corresponding relationship table of the fluorescence label type and the repeating unit base combination is established, and the corresponding repeating unit base combination is matched from the corresponding relationship table according to the label type of the fluorescence signal; According to the position information of the distribution of each fluorescence signal on the target DNA fragment and the distance between the adjacent signals, the signals are sorted in the linear extension direction of the DNA fragment; According to the sorting result and the matched repeating unit base combination, the repeating unit arrangement order of the preliminary identified sequence is determined.

[0012] Further, before extracting the position information and intensity information of each fluorescence signal in the effective signal region, the fluorescence image is preprocessed, specifically including: The multi-channel fluorescence image sequence is subjected to background noise reduction by Gaussian filtering, and the filter kernel size is set to 3*3 pixels; the background signal intensity average value of each channel image is calculated by image analysis, and the region with a fluorescence signal intensity higher than 3 times the background signal intensity is marked as an effective signal region; the image information of the effective signal region is retained, and the regions and noise interference signals below the threshold are excluded.

[0013] Further, sequence verification splicing specifically includes: At least two characteristic fragments in the preliminary identified sequence are selected as verification targets, and each characteristic fragment has a length of 20-30 nt; Designing a verification probe complementary to the feature fragment, the verification probe being 20-30 nt in length, the 5' end of the verification probe being labeled with a verification label, the verification label being a fluorescent quantum dot different from the fluorescent probes in the fluorescent probe set, the emission wavelength of the verification label being different from the emission wavelength of the fluorescent labels in the fluorescent probe set by ≥20 nm, and synthesizing the verification probe; Performing secondary hybridization on the completed hybridization slide using the verification probe, imaging the completed secondary hybridization slide, and obtaining a fluorescent image after secondary hybridization; Extracting the signal distribution position in the secondary hybridization fluorescent image, comparing the predicted position of the corresponding feature fragment in the preliminary identification sequence, and calculating the matching degree, if the matching degree is ≥90%, then confirming that the preliminary identification sequence of this part is correct; Comparing the sequencing results of the plurality of target DNA fragments to identify an overlapping sequence region; Integrating the overlapping sequence region to obtain a complete multi-gene repeat sequence, and correcting the splicing direction based on the verified correct preliminary identification sequence fragment during splicing integration.

[0014] Further, in the fluorescent signal analysis and sequence identification step, further specifically comprising the following steps: Extracting the intensity information and position information of each fluorescent signal in the effective signal region, the intensity information being the total fluorescent intensity value in the signal region calculated by integration, denoted as I i , wherein i represents the i-th fluorescent signal point; the position information being the relative position coordinates of the signal in the extension direction of the DNA fragment calculated based on the signal center point pixel coordinates, denoted as P i ; the distance between adjacent fluorescent signals being calculated by the straight-line distance between the signal center points, denoted as D ij , wherein i and j represent adjacent i-th and j-th fluorescent signal points, respectively; Based on the intensity and position information of the fluorescent signal, calculating the sequence repetition degree R of the target DNA fragment, the sequence repetition degree R being used to represent the distribution regularity of the repeat unit in the target multi-gene repeat sequence, and the calculation formula being: R=Σ(I i ×W i ) / (N×D avg ) , wherein R represents the sequence repetition degree, which is dimensionless; I i represents the total fluorescent intensity value of the i-th fluorescent signal, which is in gray value and ranges from 0 to 65535; W i represents the weight factor of the i-th fluorescent signal, which is defined as W i =1 / (1+|P i -P avg | / P std ), wherein P avgP is the average value of the position coordinates of all fluorescent signals std W is the standard deviation of the position coordinates i N is the total number of fluorescent signal points, ranging from 10 to 1000; D avg D is the average distance between all adjacent fluorescent signals, in pixels, and is calculated by the formula D avg =∑D ij / (N-1); ∑ represents the summation operation on all fluorescent signal points i; By calculating the sequence repetition R, the uniformity of the distribution of fluorescent signals and the periodicity of the repeating unit are evaluated. If the R value is greater than a preset threshold T, it is considered that the target DNA fragment has a high regularity of repeated sequences, and the sequence recognition step is entered. The range of T is 0.8-1.2; If R is less than T, the fluorescent signal is corrected again, which specifically includes adjusting the excitation light intensity of the fluorescence microscope to 1.2 times the initial value, reacquiring a multi-channel fluorescence image sequence, and reextracting signal intensity and position information. The R value is repeatedly calculated until R≥T or the maximum correction number 3 times is reached; In the sequence recognition step, the sequence repetition R is combined with the corresponding relationship table of the fluorescent label type to optimize the matching accuracy of the repeating unit base combination. Specifically, the fluorescent signals are grouped according to the R value, and the signal group with a higher R value is preferentially matched with a high-confidence repeating unit base combination, and the signal group with a lower R value is verified after matching. The sequence recognition result optimized by the R value is used in the subsequent sequence verification and splicing step to improve the accuracy of the sequencing result.

[0015] Further, after the sample targeting pretreatment step, it further specifically includes the following steps: After the target DNA fragments are enriched, an aliquot of the enriched target DNA fragment solution is taken, the sample volume is 5-10 μL, a nucleic acid quantitative dye is added, the nucleic acid quantitative dye is a double-stranded DNA binding dye with high specificity, and the dye concentration is 0.1-0.5 μg / mL; the sample is placed in a microspectrophotometer, the absorbance value A260 at 260 nm wavelength and the absorbance value A280 at 280 nm wavelength are measured, the A260 / A280 ratio is calculated, and the purity of the target DNA fragments is determined. If the A260 / A280 ratio is in the range of 1.8-2.0, it is considered that the target DNA fragment purity is qualified, and the subsequent step is entered; if the A260 / A280 ratio is not in the range, the magnetic bead enrichment step is performed again until the purity is qualified or the maximum repeated enrichment number 3 times is reached; The concentration of the target DNA fragment is determined by a fluorescent quantitative PCR technology, specifically, a reaction system containing the target DNA fragment, primers, fluorescent dyes and polymerase is prepared, the primer sequence is complementary to the flanking conserved region of the target repeat sequence, the length of the primer is 18-22 nt, and the concentration of the primer is 0.2-0.5 μmol / L; the reaction system is placed in a real-time fluorescent quantitative PCR instrument, and the PCR reaction program is as follows: 95°C pre-denaturation for 3 minutes, then 35 cycles of 95°C denaturation for 30 seconds, 60°C annealing for 30 seconds and 72°C extension for 30 seconds; the concentration of the target DNA fragment is calculated by monitoring the fluorescent signal intensity in real time and combining a standard curve, the standard curve is established by using DNA standard samples with known concentrations, and the concentration range is 0.1-100 ng / μL; if the concentration of the target DNA fragment is in the range of 5-50 ng / μL, it is considered that the concentration is qualified, and the subsequent in situ hybridization reaction step is entered; if the concentration is unqualified, the volume of the elution buffer in the enrichment step is adjusted, and the concentration is re-enriched and determined until it is qualified; The length of the target DNA fragment is analyzed, specifically, 1 μL of the enriched target DNA fragment solution is taken and added into an agarose gel electrophoresis system, the concentration of the agarose gel is 1.5%, the electrophoresis buffer is 1×TAE buffer, and the electrophoresis condition is 100 V voltage for 30 minutes; after the electrophoresis is completed, the electrophoresis strip is observed by using an ultraviolet transmission instrument, and the length of the main band is recorded, and the length of the band is determined by comparison with a DNA ladder with known molecular weight; if the length of the main band is in the range of 200-1000 bp, it is considered that the length of the target DNA fragment is qualified, and the subsequent step is entered; if the length is unqualified, the amount of the restriction endonuclease or the enzyme digestion time in the enzyme digestion step is optimized, and the enzyme digestion and enrichment are re-performed until the length is qualified; through the quantitative quality control step, it is ensured that the purity, concentration and length of the target DNA fragment meet the requirements of subsequent in situ hybridization and fluorescent imaging, and the reliability and accuracy of sequencing are improved.

[0016] Compared with the prior art, the beneficial effects of the present application are: 1. The sample targeted pretreatment and fluorescent probe design of the present application can significantly improve the sequencing efficiency and accuracy by accurately obtaining the target sample and specific probe combination, the steps of microscopic cutting to obtain a target cell group, releasing DNA in a low permeability solution, target cutting by a restriction endonuclease and magnetic bead enrichment can realize high-purity and high-concentration acquisition of the target DNA fragment, solve the problems of low content of multi-gene repeat sequences and easy interference, provide high-quality templates for subsequent reactions, the fluorescent probe design adopts multiple fluorescent labels, combines specific length and high-purity requirements, ensures that the probe can accurately bind to the target sequence, and the fluorescent signal is clear and distinguishable, reduces non-specific hybridization interference, and improves the signal-to-noise ratio of the hybridization signal.

[0017] 2. The in situ hybridization and multi-layer fluorescence imaging collection of the application provide clear signals for sequence analysis by stable hybridization and high-quality imaging, DNA fixation and dehydration treatment ensure its linear expansion, reasonable probe concentration and hybridization and washing conditions improve the binding efficiency of the probe and the target DNA, reduce the background signal, and multi-layer fluorescence imaging acquires high-resolution, low-interference multi-channel fluorescence images by precisely adjusting the light path, selecting matching filters, and controlling the imaging interval, avoids signal crosstalk and fluorescence bleaching, and completely retains the signal position and intensity information.

[0018] 3. The fluorescence signal analysis and sequence verification splicing of the application realize complete and accurate sequence determination by accurately analyzing the signal and strictly verifying the splicing, the fluorescence signal analysis converts the fluorescence image into readable sequence information, solves the problem of correspondence between fluorescence signal and sequence, the sequence verification splicing ensures the accuracy of the sequence by secondary hybridization verification, integrates the overlapping region based on the verified correct fragment to form a complete sequence, breaks through the limitation of short fragment sequencing, improves the accuracy of the final sequence, and realizes efficient and accurate sequencing of multiple gene repeat sequences. BRIEF DESCRIPTION OF DRAWINGS

[0019] Figure 1 The figure is a flowchart of the multiple gene repeat sequence sequencing method of the application. DETAILED DESCRIPTION

[0020] The technical solutions in the embodiments of the application will be described clearly and completely below with reference to the drawings in the embodiments of the application. Obviously, the described embodiments are only part of the embodiments of the application, not all the embodiments. Based on the embodiments in the application, all other embodiments obtained by those skilled in the art without creative labor are within the protection scope of the application.

[0021] Please refer to Figure 1 The application provides the following technical solutions: The multiple gene repeat sequence sequencing method based on fluorescence imaging comprises: Sample targeting pretreatment, selecting a biological sample containing a target multiple gene repeat sequence, obtaining a cell population containing a target sequence, releasing genomic DNA in the cell population and performing enzyme cutting, obtaining DNA fragments containing a target repeat sequence, enriching the DNA fragments, and obtaining enriched target DNA fragments; Synthesizing fluorescence probes, designing multiple groups of fluorescence probe groups for characteristic repeat units of the target multiple gene repeat sequence, and synthesizing the fluorescence probe groups by a chemical synthesis method; In situ hybridization reaction, the target DNA fragment is fixed on the glass slide, the DNA fragment is kept in a linear stretched state by dehydration treatment, the fluorescent probe mixture is added to the glass slide fixed with the DNA fragment, and the in situ hybridization reaction is carried out, and after the hybridization is completed, the glass slide is washed with the washing solution to remove the unbound free probe; Multi-layer fluorescence imaging acquisition, the hybridized glass slide is placed under a fluorescence microscope, and imaging is performed in a laser confocal scanning mode to obtain a multi-channel fluorescence image sequence; Fluorescent signal analysis and sequence identification, label the effective signal region, extract the position information and intensity information of each fluorescent signal in the multi-channel fluorescence image sequence, obtain the distribution position of the fluorescent signal on the DNA fragment, and determine the arrangement order of the repeat units of the preliminary identified sequence; Sequence verification and splicing, designing verification probes for the preliminary identified sequence, performing secondary hybridization and collecting fluorescence images after secondary hybridization, comparing the signal distribution of the fluorescence images after secondary hybridization with the preliminary identified sequence, obtaining the sequencing result, splicing the sequencing results of multiple target DNA fragments and integrating the overlapping sequence parts to obtain a complete multi-gene repeat sequence.

[0022] Sample targeting pretreatment, specifically including: Selecting a biological sample containing a target multi-gene repeat sequence, separating and obtaining a cell group containing the target sequence by microdissection; Placing the cell group in a hypotonic solution and incubating at room temperature for 10-15 minutes to release genomic DNA by breaking the cells, and the hypotonic solution is 0.075 mol / L potassium chloride solution; Enzymatic digestion of the released genomic DNA with a restriction endonuclease, the recognition site of the restriction endonuclease being located in the flanking conserved region of the target repeat sequence, the genomic DNA concentration in the enzyme digestion system being 50-100 ng / μL, and the amount of restriction endonuclease being 10-20 U / μg DNA; Incubating the enzyme digestion system at 37°C for 1-2 hours to obtain DNA fragments containing the target repeat sequence; after the enzyme digestion reaction is completed, adding enzyme inactivator with a final concentration of 0.5 mol / L EDTA to terminate the reaction, and placing it in a 25°C environment for 5 minutes; Taking magnetic beads modified with single-stranded oligonucleotides complementary to the flanking conserved region of the target repeat sequence, washing them twice with binding buffer; mixing the enzyme-digested DNA fragment solution with the magnetic beads and performing 25°C constant temperature oscillation reaction for 30 minutes to allow the target DNA fragments to bind to the surface of the magnetic beads through base complementary pairing; Using a magnetic separation rack to adsorb the magnetic beads, discarding the supernatant, washing the magnetic beads with washing buffer, adding elution buffer, incubating at 65°C for 5 minutes, collecting the supernatant after magnetic separation, and obtaining the enriched target DNA fragments.

[0023] In the above embodiment, sample target pretreatment precisely acquires target cell groups through microdissection, avoids irrelevant cell interference, improves subsequent sequencing specificity, and releases genomic DNA gently through hypotonic solution treatment, reducing the risk of DNA fragmentation; restriction endonuclease target cleavage ensures the acquisition of complete target repeat sequence fragments by cutting the flanking conserved region; magnetic bead enrichment specifically binds target DNA using base complementary pairing, significantly increasing target fragment concentration, and terminating the reaction in a timely manner after enzyme digestion while strictly controlling the temperature to ensure DNA fragment stability; multi-step washing and elution operations reduce impurity residues to provide high-purity templates for subsequent hybridization, solving the problem of difficult sequencing of multi-gene repeat sequences due to low content and susceptibility to interference.

[0024] Each group of fluorescent probes includes at least 3 different fluorescently labeled probes, and different fluorescent labels correspond to different repeat unit base combinations; The fluorescent probe is 15-25 nt in length, and its sequence is complementary to the repeat unit of the target repeat sequence; the fluorescent label is located at the 5' end of the probe, and a fluorescent quantum dot is used as the fluorescent label, and different fluorescent quantum dots have distinguishable emission wavelengths; The purity of the probe is controlled to be greater than or equal to 95% during the synthesis of the fluorescent probe group, and the synthesized probe is purified by high-performance liquid chromatography.

[0025] In the above embodiment, the design and synthesis of the fluorescent probe use a multi-group fluorescent labeling strategy, with at least 3 different fluorescent probes in each group corresponding to different repeat units to achieve the identification and recognition of various repeat sequences, and the probe length of 15-25 nt balances the binding specificity and hybridization efficiency, the sequence complementary to the repeat unit ensures the accuracy of targeted binding, the fluorescent quantum dot label has the advantages of high fluorescence intensity and good stability, and the distinguishable emission wavelengths avoid signal overlap; the probe purity of greater than or equal to 95% and high-performance liquid chromatography purification reduce non-specific hybridization, and the signal-to-noise ratio of the hybridization signal is improved by 3-5 times, providing a clear and distinguishable signal basis for subsequent fluorescence imaging, and solving the problems of weak signal and poor specificity of traditional probe labeling.

[0026] The in situ hybridization reaction specifically includes: The target DNA fragments are transferred to the surface of a glass slide, and fixed with a 4% polyformaldehyde fixing solution at room temperature for 10-15 minutes; After fixation, the glass slide is subjected to gradient dehydration treatment with ethanol solutions with volume concentrations of 70%, 85%, and 100%, respectively, for 2 minutes each, so that the DNA fragments remain in a linearly stretched state, and then naturally air-dried after treatment; The fluorescent probe mixture is added dropwise to the hybridization area of the glass slide on which the target DNA fragments are fixed, and the concentration of each probe in the probe mixture is 0.5-1 μmol / L; Cover the coverslip and seal the edges, place the slide in the hybridization instrument for in situ hybridization reaction; The in situ hybridization reaction procedure is: denaturation at 95°C for 5 minutes to unwind the DNA fragments, then reduce the temperature to 60°C for hybridization for 12-16 hours to allow the fluorescent probe to bind to the target DNA fragment; After the hybridization reaction is completed, remove the coverslip and place the slide in a citrate buffer containing 0.1% sodium dodecyl sulfate, and wash at 42°C for 3 times, each for 5 minutes; Remove the unbound free probe and non-specifically bound probe by washing, and after washing, quickly rinse the slide with ultrapure water, and obtain the completed hybridization slide after nitrogen blowing.

[0027] In the above embodiment, the in situ hybridization reaction is fixed by 4% paraformaldehyde and dehydrated by gradient ethanol, which linearly stretches and firmly fixes the DNA fragments, avoids folding to affect hybridization and imaging, and controls the probe concentration at 0.5-1 μmol / L to balance the binding efficiency and the risk of non-specific binding; 95°C denaturation ensures DNA unwinding, and 60°C constant temperature hybridization for 12-16 hours ensures sufficient probe binding; the washing solution containing 0.1% sodium dodecyl sulfate is accurately washed at 42°C, which effectively removes the free probe, reduces the background signal, improves the binding efficiency of the probe and the target DNA, and reduces the non-specific signal, thereby providing a high-specificity hybridization template for fluorescence imaging.

[0028] Multi-layer fluorescence imaging acquisition, specifically comprising: Place the completed hybridization slide on the objective stage of the fluorescence microscope, adjust the position of the objective stage to make the sample area in the center of the microscope light path; According to the emission wavelength of the used fluorescent marker, select a matched multi-channel filter set and install it in the microscope light path, turn on the microscope light source, adjust the excitation light intensity to the preset initial value, and preheat the microscope system; Adjust the focal length through the microscope eyepiece or imaging software to make the target DNA fragment in the clear imaging field, and record the coordinate position of the imaging field; Start the laser confocal scanning system, set the scanning resolution to 1024x1024 pixels, and the scanning speed to 1-2 frames per second; Perform multiple fluorescence imaging on the same imaging field in turn, each imaging corresponding to the excitation wavelength of one fluorescent marker, and monitor the fluorescence signal intensity in real time during imaging; The time interval between adjacent two fluorescence imaging is 10 seconds, and the next imaging is started after the signal of the previous fluorescent marker is stable, and the image acquisition of all fluorescence channels is completed in turn to obtain a multi-channel fluorescence image sequence; Name the collected multi-channel fluorescence image sequence according to the channel type and field coordinates, and store it in a special image database.

[0029] In the above embodiment, the multi-layer fluorescence imaging acquisition is realized by precisely adjusting the objective table and the optical path to ensure that the sample is in the best imaging position; the multi-channel filter set matched with the fluorescence label avoids signal crosstalk, the laser confocal scanning realizes high-resolution imaging (1024x1024 pixels), 10-second interval imaging reduces fluorescence bleaching, real-time monitoring of signal intensity guarantees image quality; the storage according to the channel and field of view naming facilitates subsequent analysis, the conversion of fluorescence signal into clear digital image retains signal position and intensity information, improves the resolution of imaging mode, provides reliable data support for sequence identification, and solves the problems of blurred imaging and signal loss of repeated sequences.

[0030] The fluorescence signal analysis and sequence identification specifically includes: The multi-channel fluorescence image sequence is subjected to background noise reduction through Gaussian filtering, and the filter kernel size is set to 3x3 pixels; the background signal intensity average of each channel image is calculated through image analysis, and the region with fluorescence signal intensity higher than 3 times the background signal intensity is marked as an effective signal region; the image information of the effective signal region is retained, and the regions below the threshold and noise interference signals are excluded; The position information and intensity information of each fluorescence signal in the effective signal region are extracted; The position information is determined by a signal positioning algorithm, and the relative position coordinates of the signal in the DNA fragment extension direction are calculated based on the signal center point pixel coordinates; The intensity information is calculated by integrating the total fluorescence intensity value in the signal region; The distance between adjacent two fluorescence signals is also calculated, and the distance is calculated based on the straight-line distance between the signal center points; A corresponding relationship table of fluorescence label type and repeated unit base combination is established, and the corresponding repeated unit base combination is matched from the corresponding relationship table according to the label type of the fluorescence signal; According to the position information of the fluorescence signal distributed on the target DNA fragment and the distance between adjacent signals, the signals are sorted in the linear extension direction of the DNA fragment; According to the sorting result and the matched repeated unit base combination, the arrangement order of the repeated units in the preliminary identified sequence is determined.

[0031] In the above embodiment, the fluorescence signal analysis and sequence identification are realized by Gaussian filtering and signal threshold screening, which effectively eliminates noise, accurately obtains signal position by positioning algorithm, and realizes the ordered arrangement of signals on the DNA fragment by combining intensity information and adjacent distance; the corresponding relationship table of fluorescence label and repeated unit directly converts the fluorescence signal into base sequence information, and through the logical chain of image preprocessing-information extraction-sequence deduction, the fluorescence image is converted into readable sequence information, which solves the core problem of correspondence between fluorescence signal and sequence.

[0032] The sequence verification splicing specifically comprises: At least two characteristic fragments in the preliminary identified sequence are selected as verification targets, and each characteristic fragment has a length of 20-30 nt; A verification probe complementary to the characteristic fragment is designed, the verification probe has a length of 20-30 nt, the 5' end of the verification probe is labeled with a verification label, the verification label is a fluorescent quantum dot different from the fluorescent probes in the fluorescent probe group, the emission wavelength of the verification label is different from the emission wavelength of the fluorescent label in the fluorescent probe group by ≥20 nm, and the verification probe is synthesized; The verification probe is used for secondary hybridization on the completed hybridization slide, the slide after the secondary hybridization is imaged, and a fluorescent image after the secondary hybridization is obtained; The signal distribution position in the secondary hybridization fluorescent image is extracted, the predicted position of the corresponding characteristic fragment in the preliminary identified sequence is compared, and the matching degree is calculated, if the matching degree is ≥90%, it is confirmed that the preliminary identified sequence of this part is correct; The sequencing results of the plurality of target DNA fragments are compared, and an overlapping sequence region is identified; The overlapping sequence region is integrated, and a complete multi-gene repeat sequence is obtained, and the splicing direction is corrected based on the verified correct preliminary identified sequence fragment during the integration.

[0033] In the above embodiment, the sequence verification splicing is performed by designing an independent verification probe (with a fluorescent difference of ≥20 nm from the original probe), secondary hybridization verification is performed to avoid single hybridization error, a matching degree of ≥90% is used to ensure sequence accuracy, the splicing direction is corrected based on the verified correct fragment to reduce splicing errors, overlapping regions are identified and integrated to solve the problem of incomplete sequence caused by the length limitation of the target DNA fragment. The step forms a closed loop of “preliminary identification-independent verification-splicing integration”, improves the accuracy of the final complete sequence, and realizes long fragment sequencing through overlapping region integration, thereby breaking through the length limitation of short fragment sequencing.

[0034] In the fluorescent signal analysis and sequence identification step, the following steps are further specifically included: The intensity information and position information of each fluorescent signal in the effective signal region are extracted, the intensity information is calculated by integrating the total fluorescent intensity value in the signal region, and is denoted as I i , wherein i represents the i-th fluorescent signal point; the position information is calculated based on the pixel coordinates of the signal center point, and the relative position coordinates of the signal in the extension direction of the DNA fragment are calculated, and are denoted as P i ; the distance between adjacent fluorescent signals is calculated by the straight line distance between the signal center points, and is denoted as D ij , wherein i and j represent adjacent i-th and j-th fluorescent signal points, respectively; Based on the intensity and position information of the fluorescence signal, the sequence repetition degree R of the target DNA fragment is calculated, which is used to represent the distribution regularity of the repeat unit in the target multi-gene repeat sequence, and the calculation formula is: R =∑(I i ×W i ) / (N×D avg ) Wherein, R represents the sequence repetition degree, which is dimensionless; I i represents the total fluorescence intensity value of the i-th fluorescence signal, which is in gray value, ranging from 0 to 65535; W i represents the weight factor of the i-th fluorescence signal, which is defined as W i =1 / (1+|P i -P avg | / P std ), wherein P avg is the average value of all fluorescence signal position coordinates, P std is the standard deviation of the position coordinates, W i is used to represent the degree of signal position deviation from the average position, ranging from 0 to 1; N represents the total number of fluorescence signal points, ranging from 10 to 1000; D avg represents the average distance between all adjacent fluorescence signals, which is in pixels, and the calculation formula is D avg =∑D ij / (N-1); ∑ represents the summation operation on all fluorescence signal points i; By calculating the sequence repetition degree R, the distribution uniformity of the fluorescence signal and the periodicity of the repeat unit are evaluated, and if the R value is greater than a preset threshold T, it is considered that the repeat sequence of the target DNA fragment has high regularity, and enters the sequence recognition step; T ranges from 0.8 to 1.2; If the R value is less than T, the fluorescence signal is corrected again, which specifically includes adjusting the excitation light intensity of the fluorescence microscope to 1.2 times the initial value, reacquiring the multi-channel fluorescence image sequence, and reextracting the signal intensity and position information, repeating the calculation of R value, until R≥T or the maximum correction number 3 times is reached; In the sequence recognition step, the sequence repetition degree R is combined with the fluorescence label type corresponding relationship table to optimize the matching accuracy of the repeat unit base combination, specifically: according to the R value, the fluorescence signal is grouped, the signal group with higher R value is preferentially matched with the repeat unit base combination with high confidence, and the signal group with lower R value is matched and verified; The sequence recognition result optimized by R value is used in the subsequent sequence verification and splicing step to improve the accuracy of the sequencing result.

[0035] In the above embodiments, the fluorescent signal analysis and sequence recognition steps are aimed at converting the fluorescent signal into the repeat sequence information of the target DNA fragment by analyzing the signal characteristics in the multi-channel fluorescent image sequence, ensuring the accuracy of the sequencing results. The intensity information (denoted as I i ) of the fluorescent signal is obtained by integrating the pixel gray value in the effective signal area. Specifically, the sum of the pixel gray values of each signal area is calculated using image processing software (such as ImageJ), and the gray value range is 0-65535, the unit is dimensionless gray value, which reflects the brightness of the fluorescent signal. The position information (denoted as P i ) is calculated based on the pixel coordinates of the signal center point by the signal positioning algorithm (such as the centroid method). The specific steps are as follows: first, apply Gaussian filtering (kernel size 3x3 pixels) to the signal area to remove background noise, then calculate the weighted average coordinates of the pixel gray value in the signal area to determine the signal center point, and convert it into relative position coordinates along the extension direction of the DNA fragment (usually the horizontal axis of the image), the unit is pixel. The distance (denoted as D ij ) between adjacent fluorescent signals is obtained by calculating the Euclidean straight line distance between the center points of the two signals, the unit is pixel, and the calculation formula is D ij =√((x i -x j ) 2 +(y i -y j ) 2 ), where (x i , y i ) and (x j , y j ) are the center point coordinates of the i-th and j-th signals, respectively. Based on the above information, the sequence repeat degree R is calculated to represent the distribution regularity of the repeat units, and the formula is R=Σ(I i ×W i ) / (N×D avg ). Wherein, W i is the weight factor, defined as W i =1 / (1+|P i -P avg | / P std ), P avg is the average value of all signal position coordinates (calculated by dividing the sum of all P i by the total number of signals N), P std is the standard deviation of the position coordinates (calculated by the standard deviation formula √(Σ(P i -P avg ) 2 / N)), W i ranges from 0 to 1, reflecting the degree of signal position deviation from the average position; D avg is the average distance between adjacent signals, and the calculation formula is Davg =ΣD ij / (N-1), where N is the total number of fluorescence signals (range 10-1000). The R value is obtained by applying I to all signal points. i and W i Weighted summation divided by N and D avg The values ​​are dimensionless. If R ≥ T (T is a preset threshold, ranging from 0.8 to 1.2, set based on empirical values), the sequence is considered to have high regularity, and sequence identification proceeds directly. If R < T, the excitation light intensity of the fluorescence microscope is adjusted to 1.2 times the initial value (the light source power is adjusted via the microscope software; the initial value is usually 50-100 mW), the image is reacquired, and R is recalculated, with a maximum of 3 corrections. In sequence identification, a correspondence table is established between fluorescent label types (such as quantum dot emission wavelengths) and repeating unit base combinations (pre-generated using probe design software, based on target sequence characteristics). Signals are grouped according to R values ​​(higher R values ​​are prioritized for matching high-confidence base combinations), and the signal positions are sorted before matching base combinations to generate a preliminary identification sequence.

[0036] In a real-world case, sequencing of a specific STR (short tandem repeat) region in the human genome involved acquiring multi-channel fluorescence image sequences (1024×1024 pixels, 3 channels corresponding to 3 types of quantum dot labels). These images were imported into ImageJ software, and 3×3 Gaussian filtering was applied for noise reduction. The average background signal intensity was calculated to be 5000 grayscale values. Regions with a label intensity ≥15000 grayscale values ​​were considered valid signal regions, resulting in the identification of 50 signal points. For each signal point, the total fluorescence intensity Ii was calculated (e.g., Ii for signal point 1). i =20000 grayscale value), the center point coordinates Pi are determined by the centroid method (e.g., P of signal point 1). i =100 pixels (along the DNA extension direction), and calculate the distance D between adjacent signals. ij (e.g., D12 = 10 pixels). Further calculate P. avg =120 pixels, P std =15 pixels, resulting in W i (e.g., W at signal point 1) i =1 / (1+|100-120| / 15)=0.43), D avg =ΣD ij / 49=12 pixels, substituting into the formula R=Σ(I i ×W i ) / (N×D avg), R = 0.95 is obtained. If T = 0.8 is preset, and R > T, the sequence recognition is directly entered; if R = 0.7 < T, the excitation light intensity is increased from 100 mW to 120 mW by the microscope software, the image is re-collected, and the calculation is repeated until R ≥ 0.8 or the correction is repeated for 3 times. In the sequence recognition, according to the corresponding relationship table (for example, the wavelength of 550 nm corresponds to the AT repeat unit, and the wavelength of 580 nm corresponds to the GC repeat unit), the signal group with high R value (such as R ≥ 0.9) is preferentially matched with the high-confidence base combination (such as AT), and the signal group with low R value (such as R < 0.9) is matched and verified by secondary hybridization, and finally the preliminary sequence (such as AT-GC-AT-AT) is generated, which provides high-accuracy data for subsequent verification and splicing. The method optimizes the signal grouping and matching through the R value, significantly improves the recognition accuracy of complex repeat sequences, and solves the sequence derivation error problem caused by uneven signal distribution.

[0037] After the sample target pretreatment step, further specifically comprising the following steps: After the target DNA fragment is enriched, an aliquot sample of the enriched target DNA fragment solution is taken, the sample volume is 5-10 μL, a nucleic acid quantitative dye is added, the nucleic acid quantitative dye is a double-stranded DNA binding dye with high specificity, and the dye concentration is 0.1-0.5 μg / mL; the sample is placed in a microspectrophotometer, the absorbance value A260 at 260 nm wavelength and the absorbance value A280 at 280 nm wavelength of the sample are measured, the A260 / A280 ratio is calculated, the purity of the target DNA fragment is judged, if the A260 / A280 ratio is in the range of 1.8-2.0, it is considered that the purity of the target DNA fragment is qualified, and the subsequent step is entered; if the A260 / A280 ratio is not in the range, the magnetic bead enrichment step is re-performed until the purity is qualified or the maximum repeated enrichment number of 3 times is reached; The concentration of the target DNA fragment is measured by the fluorescence quantitative PCR technology, specifically: a reaction system containing the target DNA fragment, primers, fluorescent dyes and polymerase is prepared, the primer sequence is complementary to the flanking conserved region of the target repeat sequence, the primer length is 18-22 nt, and the primer concentration is 0.2-0.5 μmol / L; the reaction system is placed in a real-time fluorescence quantitative PCR instrument, and the PCR reaction program is: 95°C pre-denaturation for 3 minutes, followed by 35 cycles, each cycle including 95°C denaturation for 30 seconds, 60°C annealing for 30 seconds, and 72°C extension for 30 seconds; the concentration of the target DNA fragment is calculated by real-time monitoring of the fluorescence signal intensity combined with the standard curve, and the standard curve is established by a known concentration of DNA standard sample, and the concentration range is 0.1-100 ng / μL; if the concentration of the target DNA fragment is in the range of 5-50 ng / μL, it is considered that the concentration is qualified, and the subsequent in-situ hybridization reaction step is entered; if the concentration is not qualified, the elution buffer volume in the enrichment step is adjusted, and the concentration is re-enriched and measured until it is qualified. The target DNA fragment is subjected to fragment length analysis, specifically: 1 μL of the enriched target DNA fragment solution is taken and added to an agarose gel electrophoresis system, the agarose gel concentration is 1.5%, the electrophoresis buffer is 1×TAE buffer, the electrophoresis condition is 100 V voltage, and the running time is 30 minutes; after the electrophoresis is completed, the electrophoresis strip is observed using an ultraviolet transmission instrument, the length of the main strip is recorded, and the strip length is determined by comparison with a known molecular weight DNA ladder; if the length of the main strip is in the range of 200-1000 bp, it is considered that the length of the target DNA fragment is qualified, and the subsequent step is entered; if the length is not qualified, the amount of the restriction endonuclease in the enzyme digestion step or the enzyme digestion time is optimized, and the enzyme digestion and enrichment are performed again until the length is qualified; through the quantitative quality control step, it is ensured that the purity, concentration and length of the target DNA fragment meet the requirements of the subsequent in situ hybridization and fluorescence imaging, and the reliability and accuracy of sequencing are improved.

[0038] In the above embodiments, after sample target pre-treatment, quantitative quality control steps are used to ensure that the enriched target DNA fragments meet the subsequent in situ hybridization and fluorescence imaging requirements in purity, concentration and length, and to improve sequencing reliability. Purity detection is determined by a microspectrophotometer to measure the absorbance of the sample at 260 nm (A260) and 280 nm (A280) wavelengths, and the A260 / A280 ratio is calculated. A260 reflects DNA concentration, and A280 reflects protein impurities. An ideal ratio of 1.8-2.0 indicates high-purity DNA. The specific steps are as follows: take 5-10 μL of enriched DNA solution, add nucleic acid quantitative dye (such as SYBR Green, concentration 0.1-0.5 μg / mL, which enhances absorbance at 260 nm after binding to double-stranded DNA), place it in a spectrophotometer (such as NanoDrop), set the light path length to 1 mm, measure A260 and A280, and calculate the ratio. If the ratio is not within the range of 1.8-2.0, it indicates that there is protein or RNA contamination, and the magnetic bead enrichment needs to be repeated (adjust the salt concentration of the elution buffer, such as from 150 mM to 200 mM), a maximum of 3 times. Concentration detection is performed by quantitative PCR (qPCR), which includes target DNA, primers (sequences complementary to the flanking conserved regions of the target repeat sequence, lengths of 18-22 nt, concentrations of 0.2-0.5 μmol / L, generated by primer design software such as Primer3), fluorescent dye (such as SYBR Green), and Taq polymerase, with a total volume of 20 μL. The qPCR program is 95°C pre-denaturation for 3 minutes, 35 cycles of 95°C denaturation for 30 seconds, 60°C annealing for 30 seconds, and 72°C extension for 30 seconds. The fluorescence intensity is monitored by a real-time fluorescence quantitative PCR instrument (such as Bio-Rad CFX96), and the DNA concentration is calculated by combining the standard curve (constructed by 0.1-100 ng / μL known concentration DNA standard sample, linear regression of Ct value and concentration logarithm), with a target range of 5-50 ng / μL. If the concentration is not qualified, adjust the elution buffer volume (such as from 50 μL to 30 μL) to re-enrich. Fragment length analysis is performed by 1.5% agarose gel electrophoresis, taking 1 μL of DNA solution and adding it to the electrophoresis system (1x TAE buffer, 100 V, 30 minutes). After electrophoresis, the band length is determined by observing the band with a UV transilluminator and comparing it with a DNA ladder (100-2000 bp), with a target range of 200-1000 bp. If the length is not qualified, optimize the enzyme digestion step (adjust the amount of endonuclease, such as from 10 U / μg to 15 U / μg, or the enzyme digestion time from 1 hour to 1.5 hours), and re-digest and enrich. These quality control steps use quantitative data from spectrophotometry, qPCR and electrophoresis to ensure DNA sample quality.

[0039] In one practical case, when sequencing a VNTR (variable number of tandem repeat) region in human genome, 8 μL of the enriched DNA fragment solution was taken, 0.2 μg / mL SYBR Green dye was added, and the solution was placed in a NanoDrop spectrophotometer. The A260=0.5 and A280=0.27 were measured, and the A260 / A280=1.85 was calculated, which was in the range of 1.8-2.0, indicating that the purity was qualified. If the ratio was 1.6, it indicated protein contamination. The NaCl concentration of the magnetic bead elution buffer was adjusted to 200 mM, and the enrichment was repeated. The ratio was measured again until it was qualified or repeated for 3 times. To determine the concentration, a qPCR reaction system (containing 0.3 μmol / L primers, the sequences were 5'-CTGATCGTACG-3' and 5'-AGCTAGCTAGC-3', which were complementary to the conserved regions flanking the VNTR) was prepared, and a Bio-Rad CFX96 was run. The Ct value of 20 was compared with the standard curve, and the concentration was calculated to be 25 ng / μL, which met the requirement of 5-50 ng / μL. If the concentration was 2 ng / μL, the elution buffer volume was reduced to 30 μL, and the enrichment was repeated for determination. For length analysis, 1 μL of the DNA solution was taken and added to a 1.5% agarose gel, and electrophoresis was performed at 100 V for 30 minutes. Under the observation of an ultraviolet transmission instrument, the main band was about 500 bp, which was compared with a 100-2000 bp DNA ladder, and the length was confirmed to be qualified. If the band was 1500 bp, the amount of enzyme used was adjusted to 15 U / μg DNA, and the enzyme was cut for 1.5 hours. The enrichment and electrophoresis were repeated until the length was in the range of 200-1000 bp. These quality control steps ensured high purity (no protein contamination) of the DNA fragments, appropriate concentration (convenient for hybridization), and suitable length (suitable for imaging), which significantly reduced the problems of weak signal or sequence error caused by insufficient sample quality in subsequent steps, and laid a foundation for high-precision sequencing.

[0040] The above merely provides the preferred embodiments of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art, according to the technical solution and inventive concept of the present application, can make equivalent replacement or change within the technical range disclosed by the present application, which should be covered within the protection scope of the present application.

Claims

1. A method for multiple gene repeat sequence sequencing based on fluorescence imaging, characterized in that, The method comprises the following steps: Sample target pretreatment: selecting a biological sample containing a target multi-gene repeat sequence, obtaining a cell group containing the target sequence, releasing genomic DNA in the cell group and performing enzyme digestion, obtaining DNA fragments containing the target repeat sequence, enriching the DNA fragments, obtaining the enriched target DNA fragments, determining the purity, concentration and fragment length of the target DNA fragments, adjusting the sample target pretreatment parameters when the determination result is insufficient, and re-enriching until the determination result is qualified; Synthesis of fluorescent probes: designing multiple groups of fluorescent probes for the characteristic repeat units of the target multi-gene repeat sequence and synthesizing the groups of fluorescent probes by chemical synthesis; In situ hybridization reaction: fixing the target DNA fragments on a glass slide, dehydrating the DNA fragments to keep them in a linear stretched state, adding a fluorescent probe mixture to the glass slide with the fixed DNA fragments, and performing in situ hybridization reaction, and then washing the glass slide with a washing solution to remove unbound free probes; Multi-layer fluorescence imaging acquisition: placing the hybridized glass slide under a fluorescence microscope, performing imaging by laser confocal scanning, and obtaining a multi-channel fluorescence image sequence; Fluorescent signal analysis and sequence identification: marking effective signal regions, extracting the position information and intensity information of each fluorescent signal in the multi-channel fluorescence image sequence, obtaining the distribution position of the fluorescent signal on the DNA fragments, and determining the repeat unit arrangement order of the preliminary identified sequence; Sequence verification and splicing: designing verification probes for the preliminary identified sequence, performing secondary hybridization and collecting fluorescence images after secondary hybridization, comparing the signal distribution of the fluorescence images after secondary hybridization with the preliminary identified sequence, obtaining sequencing results, splicing the sequencing results of multiple target DNA fragments, and integrating overlapping sequence parts to obtain a complete multi-gene repeat sequence.

2. The method of claim 1, wherein the method is a method of sequencing multiple genetic repeats based on fluorescence imaging. Sample target pretreatment, specifically comprising: Selecting a biological sample containing a target multi-gene repeat sequence, separating and obtaining a cell group containing the target sequence by microdissection; Placing the cell group in a hypotonic solution and incubating at room temperature for 10-15 minutes to release genomic DNA, wherein the hypotonic solution is a 0.075 mol / L potassium chloride solution; Performing enzyme digestion on the released genomic DNA using a restriction endonuclease, wherein the recognition site of the restriction endonuclease is located in the flanking conserved region of the target repeat sequence, the concentration of genomic DNA in the enzyme digestion system is 50-100 ng / μL, and the amount of restriction endonuclease used is 10-20 U / μg DNA; Incubating the enzyme digestion system at 37°C for 1-2 hours to obtain DNA fragments containing the target repeat sequence; after the enzyme digestion reaction is completed, adding enzyme inactivation agent with a final concentration of 0.5 mol / L EDTA to terminate the reaction, and placing it in a 25°C environment for 5 minutes; Taking magnetic beads with single-stranded oligonucleotides complementary to the flanking conserved region of the target repeat sequence on the surface, washing them twice with a binding buffer; mixing the enzyme-digested DNA fragment solution with the magnetic beads and performing 25°C constant temperature oscillation reaction for 30 minutes to allow the target DNA fragments to bind to the surface of the magnetic beads through base complementary pairing; The magnetic separation frame is used to adsorb the magnetic beads, the supernatant is discarded, the magnetic beads are washed with a washing buffer, an elution buffer is added, and the mixture is incubated at 65 DEG C for 5 minutes; the supernatant is collected after magnetic separation to obtain the target DNA fragments.

3. The multi-gene repetitive sequence sequencing method based on fluorescence imaging as described in claim 1, characterized in that, Each group of fluorescent probes comprises at least three different fluorescently labeled probes, and different fluorescent labels correspond to different base combinations of repeat units; The length of the fluorescent probe is 15-25 nt, the sequence of the fluorescent probe is complementary to the repeat unit of the target repeat sequence, the fluorescent label is located at the 5' end of the probe, and a fluorescent quantum dot is used as the fluorescent label, and different fluorescent quantum dots have distinguishable emission wavelengths; The purity of the probe is controlled to be greater than or equal to 95% during the synthesis of the group of fluorescent probes, and the synthesized probe is purified by high-performance liquid chromatography.

4. The multi-gene repetitive sequence sequencing method based on fluorescence imaging as described in claim 1, characterized in that, The in situ hybridization reaction specifically comprises the following steps: The target DNA fragments are transferred to the surface of a glass slide, and fixed at room temperature for 10-15 minutes by using a 4% paraformaldehyde fixing solution; After the fixing is completed, the glass slide is subjected to gradient dehydration treatment by using ethanol solutions with volume concentrations of 70%, 85% and 100% in sequence, each concentration of ethanol solution is treated for 2 minutes, the DNA fragments are kept in a linearly stretched state, and the glass slide is naturally air-dried after the treatment; The mixed solution of the group of fluorescent probes is added dropwise to the hybridization area of the glass slide on which the target DNA fragments are fixed, and the concentration of each probe in the probe mixture is 0.5-1 μmol / L; A cover glass is covered and the edge is sealed, and the glass slide is placed in a hybridization instrument to perform in situ hybridization reaction; The in situ hybridization reaction program is as follows: the DNA fragments are denatured at 95 DEG C for 5 minutes to make the DNA fragments unbound, and then the temperature is lowered to 60 DEG C for hybridization for 12-16 hours to make the fluorescent probes and the target DNA fragments complementary to each other; After the hybridization reaction is completed, the cover glass is removed, the glass slide is placed in a citrate buffer containing 0.1% sodium dodecyl sulfate, and washed at 42 DEG C for 3 times, each time for 5 minutes; The unbound free probes and non-specifically bound probes are removed by washing, the glass slide is quickly rinsed with ultrapure water after the washing is completed, and the glass slide on which the hybridization is completed is obtained after nitrogen blowing.

5. The multi-gene repetitive sequence sequencing method based on fluorescence imaging as described in claim 1, characterized in that, Multi-layer fluorescence imaging acquisition specifically comprises the following steps: The glass slide on which the hybridization is completed is placed on the objective table of a fluorescence microscope, the position of the objective table is adjusted so that the sample area is in the center of the light path of the microscope; According to the emission wavelengths of the used fluorescent labels, a matched multi-channel filter set is selected and installed in the light path of the microscope, the light source of the microscope is turned on, the excitation light intensity is adjusted to a preset initial value, and the microscope system is preheated; The focal length is adjusted through the ocular lens of the microscope or imaging software, so that the target DNA fragments are in a clear imaging field of view, and the coordinate position of the imaging field of view is recorded; The laser confocal scanning system is started, and the scanning resolution is set to 1024x1024 pixels, and the scanning speed is 1-2 frames per second; The same imaging field of view is sequentially subjected to multiple fluorescence imaging, each imaging corresponds to the excitation wavelength of one fluorescent label, and the fluorescence signal intensity is monitored in real time during the imaging process; The time interval between adjacent two fluorescence imaging is 10 seconds, the next imaging is started after the signal of the previous fluorescent label is stable, and the image acquisition of all fluorescence channels is sequentially completed to obtain a multi-channel fluorescence image sequence. The collected multi-channel fluorescence image sequence is named according to the channel type and field of view coordinates, and stored in a special image database.

6. The fluorescence imaging-based multiple gene repeat sequence sequencing method of claim 1, wherein, The fluorescence signal analysis and sequence identification specifically includes: The position information and intensity information of each fluorescence signal in the effective signal region are extracted; The position information is determined by a signal positioning algorithm, and the relative position coordinates of the signal in the extension direction of the DNA fragment are calculated based on the pixel coordinates of the signal center point; The intensity information is calculated by integrating the total fluorescence intensity value in the signal region; The distance between two adjacent fluorescence signals is also calculated, and the distance is calculated based on the straight-line distance between the signal center points; A corresponding relationship table of fluorescence label type and repeat unit base combination is established, and the corresponding repeat unit base combination is matched from the corresponding relationship table according to the label type of the fluorescence signal; The positions of the fluorescence signals on the target DNA fragment are combined with the distance between adjacent signals, and the signals are sorted according to the linear extension direction of the DNA fragment; According to the sorting result and the matched repeat unit base combination, the repeat unit arrangement order of the preliminary identified sequence is determined.

7. The multi-gene repetitive sequence sequencing method based on fluorescence imaging as described in claim 6, characterized in that, Before extracting the position information and intensity information of each fluorescence signal in the effective signal region, the fluorescence image is preprocessed, specifically including: The multi-channel fluorescence image sequence is denoised by Gaussian filtering, and the filter kernel size is set to 3*3 pixels; the background signal intensity average value of each channel image is calculated by image analysis, and the region with fluorescence signal intensity higher than 3 times the background signal intensity is marked as the effective signal region; the image information of the effective signal region is retained, and the regions below the threshold and the noise interference signals are removed.

8. The fluorescence imaging-based multiple gene repeat sequence sequencing method of claim 1, wherein, Sequence verification and splicing, specifically including: At least two characteristic fragments in the preliminary identified sequence are selected as verification targets, and each characteristic fragment has a length of 20-30 nt; A verification probe complementary to the characteristic fragment is designed, the verification probe has a length of 20-30 nt, the 5' end of the verification probe is labeled with a verification label, the verification label is a fluorescent quantum dot different from the fluorescent probes in the fluorescent probe group, the emission wavelength of the verification label is different from the emission wavelength of the fluorescent labels in the fluorescent probe group by ≥20 nm, and the verification probe is synthesized; The verification probe is used for secondary hybridization of the completed hybridization slide, the slide after secondary hybridization is imaged, and the fluorescence image after secondary hybridization is obtained; The signal distribution position in the secondary hybridization fluorescence image is extracted, compared with the predicted position of the corresponding characteristic fragment in the preliminary identified sequence, and the matching degree is calculated, if the matching degree is ≥90%, it is confirmed that the preliminary identified sequence of this part is correct; The sequencing results of multiple target DNA fragments are compared, and the overlapping sequence region is identified; The overlapping sequence region is integrated to obtain a complete multi-gene repeat sequence, and the correct preliminary identified sequence fragment is used as a reference to correct the splicing direction during the splicing and integration process.

9. The multi-gene repetitive sequence sequencing method based on fluorescence imaging as described in claim 1, characterized in that, In the fluorescence signal analysis and sequence identification step, the further specific steps include: The intensity information of each fluorescent signal in the effective signal region is extracted, and the position information is extracted. The intensity information is calculated by integrating the total fluorescent intensity value in the signal region, denoted as I i where i represents the i-th fluorescent signal point. The position information is based on signal center point pixel coordinates, and the relative position coordinates of the signal in the DNA fragment extension direction are calculated and recorded as P i The distance between adjacent fluorescent signals is calculated by the straight-line distance between signal center points and recorded as D ij where i and j represent the i th and j th adjacent fluorescent signal points, respectively. Based on the intensity and position information of the fluorescence signal, the sequence repetition degree R of the target DNA fragment is calculated, and the sequence repetition degree R is used to represent the distribution regularity of the repeat unit in the target multi-gene repeat sequence, and the calculation formula is: R = ∑(I i × W i ) / (N × D avg ) wherein R represents the sequence repetition degree, dimensionless; I i represents the total fluorescence intensity value of the i-th fluorescent signal, unit: gray value, range: 0-65535; W i represents the weight factor of the i-th fluorescent signal, defined as W i =1 / (1+|P i -P avg | / P std ), wherein P avg is the average value of all fluorescent signal position coordinates, P std is the standard deviation of the position coordinates, W i is used to represent the degree of signal position deviation from the average position, range: 0-1; N represents the total number of fluorescent signal points, range: 10-1000; D avg represents the average distance between all adjacent fluorescent signals, unit: pixels, calculation formula: D avg =ΣD ij / (N-1); Σ represents the summation operation on all fluorescent signal points i; The uniformity of the fluorescence signal distribution and the periodicity of the repeat unit are evaluated by calculating the sequence repetition degree R. If the R value is greater than a preset threshold T, it is considered that the repeat sequence of the target DNA fragment has high regularity, and the sequence recognition step is entered. The range of T is 0.8-1.2; If the R value is less than T, the fluorescence signal is corrected again, specifically including adjusting the excitation light intensity of the fluorescence microscope to 1.2 times the initial value, reacquiring the multi-channel fluorescence image sequence, and reextracting the signal intensity and position information. The R value is repeatedly calculated until R≥T or the maximum correction number 3 times is reached. In the sequence recognition step, the sequence repetition degree R is combined with the fluorescence labeling type corresponding relationship table to optimize the matching accuracy of the repeat unit base combination. Specifically, the fluorescence signals are grouped according to the R value. The signal group with a higher R value is preferentially matched with a high-confidence repeat unit base combination, and the signal group with a lower R value is verified after matching. The sequence recognition result optimized by the R value is used for the subsequent sequence verification and splicing step.

10. The multi-gene repetitive sequence sequencing method based on fluorescence imaging as described in claim 1, characterized in that, After the sample targeting pretreatment step, the following steps are further included: After the target DNA fragment is enriched, an aliquot of the enriched target DNA fragment solution is taken, the sample volume is 5-10 μL, a nucleic acid quantitative dye is added, the nucleic acid quantitative dye is a double-stranded DNA binding dye with high specificity, and the dye concentration is 0.1-0.5 μg / mL. The sample is placed in a microspectrophotometer, the absorbance value A260 at 260 nm and the absorbance value A280 at 280 nm are measured, the A260 / A280 ratio is calculated, and the purity of the target DNA fragment is determined. If the A260 / A280 ratio is in the range of 1.8-2.0, it is considered that the target DNA fragment is of qualified purity, and the subsequent step is entered. If the A260 / A280 ratio is not in the range, the magnetic bead enrichment step is performed again until the purity is qualified or the maximum repeated enrichment number 3 times is reached. The concentration of the target DNA fragment is measured by fluorescence quantitative PCR technology. Specifically, a reaction system containing the target DNA fragment, primers, fluorescent dyes, and polymerase is prepared. The primer sequence is complementary to the flanking conserved region of the target repeat sequence, the primer length is 18-22 nt, and the primer concentration is 0.2-0.5 μmol / L. The reaction system is placed in a real-time fluorescence quantitative PCR instrument. The PCR reaction program is as follows: 95°C pre-denaturation for 3 minutes, followed by 35 cycles, each cycle including 95°C denaturation for 30 seconds, 60°C annealing for 30 seconds, and 72°C extension for 30 seconds. The concentration of the target DNA fragment is calculated by real-time monitoring of the fluorescence signal intensity combined with a standard curve. The standard curve is established by known concentration DNA standard samples, and the concentration range is 0.1-100 ng / μL. If the concentration of the target DNA fragment is in the range of 5-50 ng / μL, it is considered to be qualified, and the subsequent in situ hybridization reaction step is entered. If the concentration is not qualified, adjust the elution buffer volume in the enrichment step, re-enrich and measure the concentration until it is qualified. The target DNA fragment is subjected to fragment length analysis, specifically: 1 μL of the enriched target DNA fragment solution is taken and added to an agarose gel electrophoresis system, the agarose gel concentration is 1.5%, the electrophoresis buffer is 1×TAE buffer, the electrophoresis condition is 100 V voltage, and the running time is 30 minutes; after the electrophoresis is completed, the electrophoresis strip is observed by using an ultraviolet transmission instrument, the length of the main strip is recorded, and the strip length is determined by comparison with a known molecular weight DNA ladder; if the length of the main strip is in the range of 200-1000 bp, it is considered that the length of the target DNA fragment is qualified, and the subsequent step is entered; if the length is not qualified, the amount of the restriction endonuclease or the enzyme digestion time in the enzyme digestion step is optimized, and the enzyme digestion and enrichment are performed again until the length is qualified.