Abnormal data positioning method of gene sequencing system, computer equipment and computer readable storage medium
By mapping and comparing gene sequencing data, the method of marking abnormal spots solves the problem of the inability to accurately locate low-quality or abnormal data in existing technologies, and enables refined analysis and improvement of sequencing systems.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- SIKUN LIFE SCIENCE CO LTD
- Filing Date
- 2026-01-23
- Publication Date
- 2026-05-05
AI Technical Summary
Existing gene sequencing system evaluation methods cannot accurately locate low-quality or abnormal data, nor can they identify and pinpoint specific problems in the sequencing system.
By acquiring gene sequencing data, mapping and comparing gene fragment sequences of sample libraries with reference sequences, marking the positions of abnormal gene fragment sequences in sequencing images, generating sequencing images marked with abnormal light spots, grouping and statistically analyzing the data, and establishing an abnormality indicator heatmap, the precise location of abnormal data can be achieved.
It enables fine-grained localization of low-quality or abnormal gene fragment sequences, helping researchers analyze and improve the stability and accuracy of sequencing systems.
Smart Images

Figure CN121983147A_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of gene sequencing technology, and more specifically, to a method for locating abnormal data in a gene sequencing system, a computer device, and a computer-readable storage medium. Background Technology
[0002] Nucleic acid sequencing technology can determine the sequence of genetic material and is widely used in molecular biology research, genetic breeding, clinical diagnosis, drug development, and other fields. Gene sequencing technology can simultaneously analyze millions or even billions of sample libraries (e.g., nucleic acid fragments), achieving high-throughput sequencing. Sequencing systems used for gene sequencing are complex in structure and highly integrated. High-quality data output requires high stability for the sequencing chip, reagent reactions, temperature control system, and optical system as a whole. Simultaneously, the sequencing system must possess high sensitivity to identify gene fragment sequences from different sample libraries. Systematic evaluation and improvement of sequencing systems based on gene sequencing data are essential for enhancing their stability. However, current evaluation methods often only reflect the overall quality of the sequencing system and cannot accurately pinpoint problems. Summary of the Invention
[0003] This disclosure provides at least one method for locating abnormal data in gene sequencing data, a computer device, and a storage medium.
[0004] In a first aspect, embodiments of this disclosure provide a method for locating abnormal data in a gene sequencing system, including: Acquire gene sequencing data by sequencing a sample library of one or more samples using a gene sequencing system. The gene sequencing data includes gene fragment sequences of the sample library, imaging unit information of the gene fragment sequences, and position information of the gene fragment sequences in the sequencing image corresponding to the imaging unit. The gene fragment sequences of the sample library are mapped and aligned with the reference sequences to obtain the mapping and alignment results of the gene fragment sequences. Based on the mapping and alignment results, abnormal gene fragment sequences are identified from the gene fragment sequences. Based on gene sequencing data, the imaging unit where the abnormal gene fragment sequence is located and the position of the abnormal gene fragment sequence in the sequencing image corresponding to the imaging unit are determined. In the sequencing image corresponding to the imaging unit, abnormal light spots at the position of the abnormal gene fragment sequence are marked to obtain a sequencing image marked with abnormal light spots.
[0005] Optionally, determining the abnormal gene fragment sequence from the gene fragment sequence based on the comparison result includes at least one of the following: If the alignment result of the gene fragment sequence indicates that the alignment was unsuccessful, the gene fragment sequence will be identified as the abnormal gene fragment sequence. If the alignment result of the gene fragment sequence indicates a successful alignment, the number of unaligned bases in the gene fragment sequence and the reference genome sequence is counted. If the number of unaligned bases is greater than a preset data threshold, the gene fragment sequence is identified as the abnormal gene fragment sequence.
[0006] Optionally, determining the abnormal gene fragment sequence from the gene fragment sequence based on the mapping and alignment results includes at least one of the following: If the mapping and alignment results of the gene fragment sequence indicate that the alignment was unsuccessful, the gene fragment sequence will be identified as the abnormal gene fragment sequence. If the mapping and alignment results of the gene fragment sequence indicate a successful alignment, the number of unaligned bases in the gene fragment sequence and the reference sequence is counted. If the number of unaligned bases is greater than a preset threshold, the gene fragment sequence is identified as the abnormal gene fragment sequence.
[0007] Optionally, the gene sequencing data may also include at least one of the following information regarding the origin of the gene fragment sequence: the upper and lower surfaces of the sequencing chip, the flow channels of the sequencing chip, the side columns of the flow channels, and the camera used to take the picture.
[0008] Optionally, the method further includes: Before mapping and aligning the gene fragment sequences of the sample library with the reference sequences, the gene fragment sequences are grouped using at least one source information as a dimension for grouping and summarizing, resulting in multiple groups of gene fragment sequences. The process involves mapping and aligning the gene fragment sequences of the sample library with reference sequences to obtain the mapping and alignment results of the gene fragment sequences. Based on the alignment results, abnormal gene fragment sequences are identified from the gene fragment sequences as follows: The gene fragment sequences in each group are mapped and compared with the reference sequence to obtain the mapping and comparison results of the gene fragment sequences in each group. Based on the mapping and comparison results, the abnormal gene fragment sequences in each group are identified.
[0009] Optionally, the method further includes: After mapping and comparing the gene fragment sequences of the sample library with the reference sequences, at least one source information is used as a grouping and summarizing dimension to group the abnormal gene fragment sequences, resulting in multiple groups of abnormal gene fragment sequences. The process involves determining the imaging unit containing the abnormal gene fragment sequence and its position in the corresponding sequencing image based on gene sequencing data. Then, in the sequencing image corresponding to the imaging unit, abnormal light spots are marked at the positions of the abnormal gene fragment sequence, resulting in a sequencing image marked with abnormal light spots. Based on the gene sequencing data, the imaging unit where the abnormal gene fragment sequence in each group is located and the position of the abnormal gene fragment sequence in each group in the sequencing image corresponding to the imaging unit are determined. In the sequencing image corresponding to the imaging unit, abnormal light spots at the positions of the abnormal gene fragment sequences in each group are marked to obtain sequencing images marked with abnormal light spots.
[0010] Optionally, marking the location of the abnormal gene fragment sequence in the sequencing image corresponding to the imaging unit to obtain a sequencing image marked with the abnormal light spots includes: In the case where the abnormal gene fragment sequence is not successfully aligned, one sequencing image is selected from all sequencing images of all sequencing cycles corresponding to the imaging unit. In the selected sequencing image, the abnormal light spot at the position of the abnormal gene fragment sequence in the imaging unit is marked to obtain a sequencing image marked with the abnormal light spot. or, If the number of bases in the abnormal gene fragment sequence that were successfully aligned but not aligned is greater than a preset threshold, the sequencing image of the sequencing cycle with the erroneous bases is selected from all sequencing images of all sequencing cycles corresponding to the imaging unit. In the selected sequencing image, the abnormal light spot at the position of the abnormal gene fragment sequence in the imaging unit is marked to obtain a sequencing image marked with the abnormal light spot.
[0011] Optionally, the method further includes: The sequencing image marked with abnormal light spots is divided into multiple sub-images; According to different anomaly types, the number or proportion of unaligned abnormal gene fragment sequences and successfully aligned but unaligned bases in the sub-images are statistically analyzed. An anomaly indicator heatmap is established based on the number or proportion of abnormal gene fragment sequences of different anomaly types in the sub-image; wherein, the anomaly indicator heatmap includes anomaly indicator information corresponding to each sub-image; the anomaly indicator information characterizes the number or proportion of unaligned abnormal gene fragment sequences and successfully aligned but unaligned bases in the sub-image through pixel values.
[0012] Optionally, the method further includes: According to different abnormality types, the number or proportion of abnormal gene fragment sequences present in the imaging unit is counted; The number or proportion of abnormal gene fragment sequences is compared with a set threshold, and the overall data quality of the imaging unit is determined based on the comparison results.
[0013] In a second aspect, an optional implementation of this disclosure also provides a computer device, a processor, and a memory, wherein the memory stores machine-readable instructions executable by the processor, and the processor is configured to execute the machine-readable instructions stored in the memory. When the machine-readable instructions are executed by the processor, the steps of the first aspect above, or any possible implementation of the first aspect, are performed.
[0014] Thirdly, an optional implementation of this disclosure also provides a computer-readable storage medium storing a computer program that, when run, performs the steps of the first aspect or any possible implementation of the first aspect.
[0015] This embodiment utilizes gene sequencing data obtained from gene sequencing of a sample library to accurately locate low-quality or abnormal gene fragment sequences in sequencing images, resulting in sequencing images marked with abnormal spots. These sequencing images can reflect the distribution of abnormal gene fragment sequences, facilitating researchers to analyze the reasons for low sequence quality or abnormality based on the phenomena observed in different sequencing cycles or positions of the sequenced images. This achieves refined localization of low-quality gene fragment sequences.
[0016] To make the above-mentioned objects, features and advantages of this disclosure more apparent and understandable, preferred embodiments are described below in detail with reference to the accompanying drawings. Attached Figure Description
[0017] To more clearly illustrate the technical solutions of the embodiments of this disclosure, the accompanying drawings used in the embodiments will be briefly described below. These drawings are incorporated in and constitute a part of this specification. They illustrate embodiments conforming to this disclosure and, together with the specification, serve to explain the technical solutions of this disclosure. It should be understood that the following drawings only show some embodiments of this disclosure and should not be considered as limiting the scope. Those skilled in the art can obtain other related drawings based on these drawings without creative effort.
[0018] Figure 1 A schematic diagram of a sequencing system provided by some embodiments of this disclosure is shown; Figure 2 A schematic diagram of a sequencing chip provided in some embodiments of this disclosure is shown; Figure 3A schematic diagram of an optical detection system provided by some embodiments of the present disclosure is shown; Figure 4 A flowchart is shown below illustrating a method for locating abnormal data in a gene sequencing system provided in some embodiments of this disclosure; Figure 5 This illustration shows one specific example of gene sequencing data carried in a fastq format sequence data file, as provided in some embodiments of this disclosure. Figure 6 This is a second specific example of a chip provided in some embodiments of this disclosure; Figure 7 This is a third specific example of a sequencing image included in an image block provided in some embodiments of this disclosure; Figure 8 This illustrates a fourth specific example of comparing a gene fragment sequence with a reference gene fragment sequence, provided in some embodiments of this disclosure. Figure 9 This is a fifth specific example of a sequencing image marked with anomalous light spots provided in some embodiments of the present disclosure; Figure 10 This is the sixth specific example of an anomaly indication heatmap provided in some embodiments of the present disclosure; Figure 11 A schematic diagram of an anomaly data localization system for a gene sequencing system provided in some of the disclosed embodiments is shown. Figure 12 A schematic diagram of a computer device provided with some of the disclosed embodiments is shown. Detailed Implementation
[0019] To make the objectives, technical solutions, and advantages of the embodiments of this disclosure clearer, the technical solutions of the embodiments of this disclosure will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of this disclosure, and not all of them. The components of the embodiments of this disclosure described and shown herein can generally be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of this disclosure is not intended to limit the scope of the claimed disclosure, but merely represents selected embodiments of this disclosure. All other embodiments obtained by those skilled in the art based on the embodiments of this disclosure without inventive effort are within the scope of protection of this disclosure.
[0020] To facilitate understanding of the technical solutions disclosed herein, the technical terms used in the embodiments of this disclosure will first be explained: Library Construction: The genomic DNA or RNA molecules to be sequenced are broken down using physical or chemical methods, such as by ultrasound. After breaking, DNA or RNA fragments are formed. First, enzymes are used to fill in the ends of the DNA or RNA fragments. Then, specific enzymes are used to ligate specific DNA or RNA sequences (usually, this specific DNA or RNA sequence is also called a linker) to the ends of the fragments, forming a mixture of DNA or RNA. This mixture of DNA or RNA is also known in the industry as a library.
[0021] To save sequencing costs, multiple samples are typically sequenced simultaneously within the same sequencing system. To differentiate the sequencing results of different samples, when preparing libraries for different samples, a DNA or RNA sequence (usually containing 6-8 bases) is included in the adapters to identify the sample's origin. This DNA or RNA sequence can also be called a sample tag (or index, barcode). Understandably, each sample's library adapter contains its own unique sample tag.
[0022] Typically, library construction is performed outside the sequencing system, for example, by obtaining the library through experimental procedures in the laboratory.
[0023] Amplification reaction: Taking the amplification of a DNA library via bridge PCR (polymerase chain reaction) as an example, after constructing the library, it can be seeded onto a sequencing chip for amplification. The adapters at both ends of the library are complementary to the first type of amplification primer on the sequencing chip; therefore, complementary hybridization allows the library to be seeded onto the sequencing chip.
[0024] After the library is seeded onto the sequencing chip, it can be used as a template for amplification. For example, the amplification process can begin by adding dNTPs and polymerase to the sequencing chip. The polymerase will synthesize a new DNA strand along the template strand, starting with the first amplification primer. This new DNA strand is completely complementary to the template strand and is therefore called the complementary strand of the template strand. This complementary strand is covalently linked to the sequencing chip. Next, NaOH solution is added to the sequencing chip for rinsing. In the presence of NaOH solution, the template strand and the complementary strand unwind, and the template strand is washed away with the alkali solution, while the complementary strand covalently linked to the sequencing chip is retained. Next, a neutral liquid is added to the sequencing chip to neutralize the NaOH solution. The entire environment within the sequencing chip becomes neutral, allowing the other end of the complementary strand to continue complementary hybridization with the second amplification primer on the sequencing chip. Then, dNTPs and polymerase are added. The polymerase synthesizes a completely new DNA strand, starting from the second amplification primer and continuing along the complementary strand. At this point, the new DNA strand is completely complementary to the complementary strand and identical to the template strand. NaOH solution is then added again to untie the two strands. This yields two covalently linked and complementary strands for the sequencing chip. Repeating this process will result in an exponential increase in the number of DNA strands.
[0025] After amplification, the sequencing chip retains two identical DNA double strands, one matching the template strand and the other the complementary strand. Specific reaction reagents are then added to the chip to cleave the DNA strand synthesized from one of the amplification primers. For example, the DNA strand matching the complementary strand is cleaved, leaving the strand matching the template strand. NaOH solution is then added to the chip for rinsing. The alkali solution unwinds the DNA double strands, and the cleaved strands are washed away, leaving only single-stranded DNA on the chip. The number of single-stranded DNA strands on the chip at this point is exponentially higher than at the start of amplification, forming DNA clusters. All single-stranded DNA strands within a cluster are identical. Adding a neutral solution allows for sequencing of all single-stranded DNA within the cluster under neutral conditions.
[0026] It should be noted that the amplification reaction described above can be performed outside the sequencing system, such as amplifying the library in the laboratory through experimental operations, or it can be performed inside the sequencing system. When the amplification reaction is performed inside the sequencing system, the library and various reaction reagents, such as deoxyribonucleoside triphosphate (dNTP), polymerase, NaOH alkaline solution, neutral solution, etc., can be added to the sequencing chip through the sequencing system's liquid circuit system. In addition, the above amplification reaction is only exemplary, and this application is not limited to the bridge PCR amplification method. Other amplification methods can also be used, such as loop-mediated isothermal amplification (LAMP), nucleic acid-dependent amplification (NASBA), rolling circle amplification (RCA), multiplex probe amplification (MPA), etc.
[0027] Sequencing reaction: Taking the sequencing-by-synthesis principle as an example, during sequencing, four dNTPs with fluorescent groups are added to the sequencing chip through a liquid circuit system. Each dNTP can only be synthesized with one of the four bases ATCG, and the 3' end of each dNTP is blocked by a blocking group (including but not limited to azide). Then, polymerase is added to the sequencing chip through the liquid circuit system. Through the action of polymerase, one of the four dNTPs will be synthesized with a complementary base on the single strand being sequenced. Since the 3' end of the dNTP is blocked by a blocking group, only one dNTP can be extended on the single strand being sequenced at a time. After synthesis, specific chemical reagents are added to the sequencing chip through the liquid circuit system to flush away excess dNTPs and polymerase. Next, the fluorescent groups of the dNTPs synthesized on the single strands can be excited by the optical detection system, causing the fluorescent groups to emit fluorescent signals. Since the fluorescent groups of the dNTPs on each single strand in a cluster will emit the same fluorescent signal, the fluorescent signal is amplified. Therefore, the optical detection system can collect the fluorescent signal and generate sequencing images.
[0028] By processing and analyzing sequencing images using a computer system, it can be determined which type of dNTP was synthesized onto the sequenced single strand. Then, based on the complementarity principle, it can be deduced which base on the sequenced single strand was synthesized with the dNTP. This completes one sequencing cycle.
[0029] Next, specific chemical reagents are added to the sequencing chip through a liquid circuit system to remove the blocking and fluorescent groups, thereby exposing the 3'-terminal hydroxyl groups of the dNTPs.
[0030] Next, proceed to the next sequencing cycle and repeat the above process.
[0031] As we can understand, one sequencing cycle can detect one base, and through multiple sequencing cycles, multiple bases in the sequenced single strand can be detected. Specifically, the number of sequencing cycles can be determined based on the set sequencing read length, for example, 150 or 300 sequencing cycles.
[0032] Of course, it should also be noted that the above sequencing reactions are merely illustrative, and this application is not limited to using the principle of sequencing by synthesis; other sequencing principles may also be used.
[0033] Sequencing system: See Figure 1 As shown, the sequencing system includes: a sequencing chip 10, a chip platform 20, a reagent storage container 30, a liquid path system (or flow distribution system) 40, an optical detection system 50, a computer system 60, a waste liquid storage container 70, and an electronic control system 80. Among them: Sequencing chip 10 is configured to provide reaction regions for amplification and sequencing reactions; Chip platform 20, configured to fix and support sequencing chip 10; The chip platform 20 can also drive the sequencing chip 10 to move along the X and Y axes, so that the optical detection system 50 can perform optical detection on different regions of the sequencing chip 10.
[0034] The reagent storage container 30 is configured to store one or more mixed sample libraries and one or more reagents; The liquid circuit system 40 is configured to controllably deliver one or more mixed sample libraries and one or more reagents from the reagent storage container 30 to the sequencing chip 10 for amplification and sequencing reactions in the sequencing chip 10, and to controllably deliver the waste liquid after the reaction from the sequencing chip 10 to the waste liquid storage container 70. The optical detection system 50 is configured to excite and acquire fluorescence signals during the sequencing reaction, and generate sequencing images based on the fluorescence signals; Computer system 60 is configured to acquire sequencing images from optical detection system 50 and identify gene fragment sequences of sample libraries based on the sequencing images; Waste liquid storage container 70 is configured to store waste liquid generated after the reaction; The electronic control system 80 is configured to control the operation and function of the chip platform 20, the liquid circuit system 40, and the optical detection system 50 according to the instructions of the computer system.
[0035] Sequencing chip: Sequencing chips, serving as carriers for amplification and sequencing reactions, provide reaction regions, known as channels. Typically, a sequencing chip 10 may contain one or more channels 11 (e.g., 2, 4, 6, or 8 channels), which are isolated from each other. See also... Figure 2 As shown, taking four channels as an example, each channel 11 has a small hole 12 at each end for the flow of fluids (e.g., biological samples, reaction reagents) into and out. The upper and lower surfaces of each channel are chemically modified and covalently seeded with two amplification primers. These two amplification primers are complementary to the adapters at both ends of the library to achieve amplification of the library. When acquiring sequencing images, the sample library is distributed in multiple sealed channels of the sequencing chip. Each sealed channel may have the sample library distributed on a single surface or on both the upper and lower surfaces. Each surface contains multiple side columns, and each side column contains multiple imaging units. The imaging field of view of the camera corresponds to one imaging unit. That is, the camera can take a picture of one imaging unit at a time to obtain a sequencing image.
[0036] Optical inspection system: See Figure 3 As shown, taking a dual-color channel as an example, the optical detection system 50 includes at least a light source assembly and an imaging assembly. The light source assembly includes at least a light source (5101, 5102), a field aperture (5201, 5202), a dichroic mirror 5301, and a filter element 540.
[0037] In one embodiment, the light source (5101, 5102) can be a light-emitting diode (LED), and the LED can be an aspherical mirror to diffuse the LED point light source into parallel light. Of course, other forms of point or area light sources can also be used besides LEDs. In another embodiment, a collimating element can also be placed after the light source (5101, 5102). Figure 3 (Not shown in the image) is used to collimate the light emitted by the light source (5101, 5102) into a parallel beam. The collimating element may include one or more lenses, including but not limited to any one or any combination of single lenses, cemented lenses, spherical lenses, and aspherical lenses.
[0038] Parallel light beams are emitted through field apertures (5201, 5202), which limit the field of view of the light emitted from the light source (5101, 5102), thereby limiting the field of view of the excitation light illuminating the sequencing chip 10. Light passing through the field apertures (5101, 5102) reaches the dichroic mirror 5301. The dichroic mirror 5301 can transmit light emitted from one of the light sources (5101, 5102) and reflect light emitted from the other light source. The light passing through the dichroic mirror 5301 is further filtered by the light filter element 540. The light filter element 540 allows light of a certain wavelength emitted by the light source (5101, 5102) to pass through and be used as excitation light, while blocking light of other wavelengths from passing through. For example, it blocks light of the same wavelength as the fluorescence emitted by the fluorescent group from passing through, ensuring that the fluorescence emitted by the fluorescent group does not contain stray light introduced by the light source (5101, 5102), which helps to improve the optical imaging effect.
[0039] In one embodiment, the dichroic mirror 5301 can be fixed at a certain angle; for example, the dichroic mirror 5301 is fixed by dispensing adhesive. In another implementation, the field stop (5201, 5202) can be, but is not limited to, a rectangular or circular stop; the field stop (5201, 5202) can be a single-aperture or multi-aperture stop; the field stop (5201, 5202) can be made of an opaque material; for example, it can be a metal sheet.
[0040] It should be noted that in the optical detection system 50, the light sources (5101, 5102) emit light alternately, not simultaneously. By operating the light sources (5101, 5102) in a time-division manner (i.e., the light wavelengths emitted by light sources 5101 and 5102 are different, and they are turned on alternately during use), only one light source is turned on at a time, which reduces the optical power of the excitation light illuminating the chip and helps protect the fluorescence lifetime of the phosphor.
[0041] Furthermore, in order to obtain the optimal sequencing image through the imaging component, it is also necessary to focus the imaging component. The light source component also includes a light source 5103, a field aperture 5203, an attenuator 550, and a dichroic mirror 5302. In one embodiment, considering that using LED light as excitation light for focusing would damage the DNA strands to be sequenced on the sequencing chip 10 and affect the subsequent sequencing quality, the light source 5103 is, for example, a semiconductor laser (LD) and emits laser light for focusing the imaging component.
[0042] When focusing the imaging component, the laser emitted by the semiconductor laser passes through the field stop 5203. The field stop 5203 defines the field of view of the laser emitted by the semiconductor laser, thereby defining the field of view of the excitation light illuminating the sequencing chip 10. The field stop 5203 can be a single-aperture or multi-aperture aperture. The field stop 5203 can be made of an opaque material, for example, a metal sheet. The laser can only pass through the light-transmitting portion of the corresponding field stop, while the laser is blocked from the non-light-transmitting portions of the field stop. The laser passing through the field stop 5203 reaches the attenuator 550. After being attenuated by the attenuator 550, the laser reaches the dichroic mirror 5302, which can transmit the laser and reflect the LED light. The dichroic mirror 5302 can be fixed at a certain angle, for example, by adhesive dispensing. The laser light transmitted through the dichroic mirror 5302 or the LED light reflected by it reaches the convex lens 560, which collimates the laser light or LED light into a parallel beam.
[0043] The imaging assembly includes at least a dichroic mirror 5303, an objective lens 570, a tube lens 580, and an image sensor 590. Laser and LED light collimated into parallel beams pass through the dichroic mirror 5303, which reflects the laser and LED light towards the sequencing chip 10. The reflected laser and LED light pass through the objective lens 570 and illuminate the sequencing chip 10, exciting a fluorescent group to emit a fluorescence signal. The objective lens 570 collects the fluorescence signal, which then passes through the objective lens 570 to the dichroic mirror 5303. The dichroic mirror 5303 transmits the fluorescence signal to the tube lens 580, which projects the fluorescence signal onto the image sensor 590. The image sensor 590 acquires the fluorescence signal and generates a sequencing image. The computer system 60 identifies gene fragment sequences based on the sequencing image. In one embodiment, a filter element may be disposed between the tube lens 580 and the dichroic mirror 5303. Figure 3 (Not shown in the image), a filter element may be disposed between the tube lens 580 and the image sensor 590. Figure 3 (Not shown in the image). In one embodiment, the image sensor 590 may be, exemplarily, an industrial camera. In another embodiment, the collimating element may include one or more lenses, including but not limited to any one or any combination of a single lens, a cemented lens, a spherical lens, and an aspherical lens.
[0044] To achieve focusing of the imaging components, a motor 600 is configured on one side of the objective lens 570. In one implementation, the motor can be, for example, a voice coil motor or other linear motor. Based on the image quality of the sequencing image identified by the computer system 60, the voice coil motor can be controlled by a driver to move the objective lens 570 up and down, so that the objective lens 570 reaches the optimal focal plane and ultimately obtains the best sequencing image.
[0045] Gene sequencing technology has been applied in tumor-related, reproductive genetics, infection-related, and scientific research fields, especially next-generation sequencing (NGS). Taking clinical tumor-related research as an example, compared with traditional methods, NGS has many advantages in accuracy, sensitivity, and speed. NGS can evaluate multiple genes in a single test, eliminating the need to order multiple tests to identify pathogenic mutations. This multi-gene approach shortens the time to obtain results, provides a more economical solution, and reduces the risk of depleting precious clinical samples. The entire sequencing system of second- or third-generation sequencing can be divided into three main control directions: 1. Temperature control, which controls the various reactions during the sequencing process; 2. Reagent control, which uses various reagents to control the synthesis of dNTPs and the cleavage of blocking and fluorescent groups; 3. Optical detection control, with mainstream sequencing systems relying on high-resolution cameras. On the surface of the sequencing chip in mainstream second-generation sequencing systems, numerous oligonucleotide chains (i.e., short nucleotide chains) complementary to P5 and P7 in the library adapters are interleaved and fixed. After a single-stranded sample library (e.g., a DNA fragment) enters the flow channel of the sequencing chip, it can bind to the oligonucleotide chains on its surface, thus entering the sequencing process. Different sequencing systems differ in their sequencing principles and system integration details, but the ultimate goal is the same: to obtain high-quality gene fragment sequences.
[0046] Sequencing systems are complex in structure and highly integrated. High-quality data output requires high overall stability from the sequencing chip, reagents, temperature control system, and optical detection system. Simultaneously, sequencing systems must possess high sensitivity to gene fragment sequences from different sample libraries. Therefore, the stability of a sequencing system depends not only on the hardware modules but also on the characteristics of the reagents and the complexity of the sample library used. Thus, systematically evaluating and improving sequencing systems based on the generated gene sequencing data is essential for enhancing their stability. This helps identify and address potential problems in the sequencing system, promoting improvements in both accuracy and stability.
[0047] Common evaluation methods include the following: 1. Quality Check of Raw Sequencing Data: In addition to producing gene fragment sequences from the sample library, the sequencing system assigns a quality value to each base in the sequence to measure sequencing accuracy. This quality value is generally expressed in Phred form: Q = -10log10(1-A). Here, A represents the probability that the base was correctly detected, i.e., the accuracy of that base. Q is the Phred value; when the base accuracy is 99%, Q is 20; when the base accuracy is 99.9%, Q is 30. The Q values of the produced sequencing data are statistically analyzed to assess the quality of the generated sequencing data and, consequently, to evaluate the sequencing system. For example, the higher the proportion of bases with a Q value greater than 20 or 30 (≥Q20 / Q30), the higher the accuracy of the generated sequencing data, and thus the easier it is to give the sequencing system a higher evaluation.
[0048] 2. Utilize reference sequence (e.g., reference genome sequence) databases: Map and align the fragment sequences produced by the sequencing system with reference sequences. Evaluate the sequencing system by comparing the alignment rate or similarity with known reference sequences. A higher alignment rate or similarity indicates better sequencing data quality and a higher overall sequencing system quality.
[0049] 3. Data Detection Accuracy and Sensitivity: Standard samples are used to verify the overall quality of the sequencing system. Different standard libraries are constructed or purchased, including but not limited to human whole-genome DNA standards and E. coli genome standards. The data obtained from sequencing these standard libraries needs to be evaluated for quality using different bioinformatics analysis software, according to the specific analytical requirements of each standard. Taking human whole-genome DNA standards as an example, key indicators include GC content, sequencing coverage, variant detection accuracy and sensitivity, especially SNP accuracy, SNP sensitivity, Indel accuracy, Indel sensitivity, SV accuracy, and SV sensitivity. For this type of data analysis, the closer the data indicators are to the given indicators of the standard, the more stable the sequencing system is.
[0050] 4. Other indicators: such as nitrogen (N) base content and repetitive sequence ratio, can supplement the evaluation of the generated sequencing data. The lower the N base content, the more accurate the sequencing system; the lower the repetitive sequence ratio, the more stable the sequencing system; and the lower the repetitive sequence ratio, the more accurate the sequencing results, thus making it easier to give the sequencing system a higher evaluation.
[0051] The above-mentioned technologies currently provide a rather general evaluation of sequencing systems, which can only reflect the overall quality of the sequencing system and cannot finely locate the low-quality sequences produced. For example, they cannot analyze and judge the reasons for the generation of low-quality gene fragment sequences, nor can they pinpoint specific aspects of the sequencing system such as optical detection systems, reagents, and sequencing chips.
[0052] Based on the above research, this disclosure provides a method for locating abnormal data in a gene sequencing system. By using gene sequencing data obtained from gene sequencing of a sample library, low-quality or abnormal gene fragment sequences are accurately located in the sequencing image, resulting in a sequencing image marked with abnormal spots. This sequencing image marked with abnormal spots can reflect the distribution of abnormal gene fragment sequences, allowing researchers to pinpoint the anomaly to a specific sequencing cycle and imaging unit based on the sequencing image marked with abnormal spots, thus achieving fine-grained localization of low-quality gene fragment sequences.
[0053] The shortcomings of the above solutions are the result of the inventor's practical experience and careful research. Therefore, the discovery process of the above problems and the solutions proposed in this disclosure below should be considered as the inventor's contribution to this disclosure.
[0054] It should be noted that similar labels and letters in the following figures indicate similar items. Therefore, once an item is defined in one figure, it does not need to be further defined and explained in subsequent figures.
[0055] To facilitate understanding of this embodiment, a detailed description of the abnormal data localization method for a gene sequencing system disclosed in this disclosure embodiment is provided first. The execution entity of the abnormal data localization method for a gene sequencing system provided in this disclosure embodiment is generally a computer device with certain computing capabilities. This computer device may include, for example, the computer system 30 within the gene sequencing system, or it may be other computer devices independent of the gene sequencing system. This computer device can utilize the output of the gene sequencing system and employ the method described in this disclosure embodiment to detect abnormalities in the gene sequencing system. In some possible implementations, the abnormal data localization method for the gene sequencing system can be implemented by a processor calling computer-readable instructions stored in memory.
[0056] The following describes the method for locating abnormal data in a gene sequencing system provided in the embodiments of this disclosure.
[0057] See Figure 4 The diagram shows a flowchart of an abnormal data localization method for a gene sequencing system provided in this embodiment of the present disclosure. The method includes steps S401 to S403, wherein: S401: Obtain gene sequencing data obtained by sequencing a sample library of one or more samples using a gene sequencing system. The gene sequencing data includes gene fragment sequences of the sample library, imaging unit information where the gene fragment sequences are located, and position information of the gene fragment sequences in the sequencing image corresponding to the imaging unit.
[0058] In practice, sample libraries may be derived from species with reference genomes or species with known gene fragment sequences.
[0059] Specifically, when sequencing a sample library, a gene sequencing system typically generates a sequencing data file in fastq format. This sequencing data file is used to carry the gene sequencing results obtained from sequencing the sample library.
[0060] like Figure 5 The example shown illustrates a specific instance of gene sequencing results carried by a FastQ format sequencing data file. In this file, the gene sequencing data for each sample library consists of four lines: a header (s1), a gene fragment sequence (s2), a "+" separator (s3), and the quality values (s4) of each base in the gene fragment sequence (s2) (e.g., Q20, Q30, etc.).
[0061] The title s1 typically includes the following information: the front and back of the sequencing chip containing the gene fragment sequence, the flow channel, the side column in the flow channel, the imaging unit in the side column, the position of the imaging unit in the sequencing image, and the camera used to take the picture.
[0062] Therefore, the gene sequencing data obtained from gene sequencing of a sample library includes at least the following three pieces of information: a1 to a3: a1: The gene fragment sequence corresponding to the sample library.
[0063] a2: Information about the imaging unit where the gene fragment sequence is located.
[0064] a3: The location information of the gene fragment sequence in the sequencing image corresponding to the imaging unit.
[0065] The gene fragment sequence s2 is the sequence consisting of individual bases in the sample library obtained by gene sequencing of the sample library.
[0066] Sequencing images are generated by acquiring fluorescence signals from sample libraries. After processing through a series of algorithms, these images can be converted into sequencing data files in FASTQ format. Gene sequencing systems typically partition the sequencing chip and, within a sequencing cycle, capture images of each imaging unit according to its specific imaging unit, thus generating image patches (i.e., sequencing images) corresponding to each unit. After multiple sequencing cycles, each imaging unit corresponds to a set of image patches (i.e., a set of sequencing images).
[0067] like Figure 6 The image shows a specific example of the chip. The chip includes four flow channels, each flow channel includes four side columns, and each side column includes 24 imaging units; that is, each flow channel includes 96 imaging units.
[0068] like Figure 7 The diagram illustrates an example of the correspondence between imaging units and sequencing images. A set of image blocks constitutes a set of sequencing images. The number of sequencing images in a set is the length of a single gene fragment sequence (one base in the gene fragment sequence corresponding to each sequencing cycle) multiplied by 2 or 4 (the difference between two-color and four-color sequencing fluorescence systems). For example, in a tag sequence mode without sequencing: a 100bp single-end sequencing fragment has a length of 100, and the number of sequencing images in a set is 100×2 or 100×4; similarly, a 75bp paired-end sequencing fragment has a total length of 150, and the number of sequencing images in a set is 150×2 or 150×4.
[0069] Following the above step S401, after acquiring gene sequencing data, the abnormal data localization method for the gene sequencing system provided in this embodiment of the disclosure further includes: S402: Map and align the gene fragment sequences of the sample library with the reference sequences to obtain the mapping and alignment results of the gene fragment sequences, and determine the abnormal gene fragment sequences from the gene fragment sequences based on the mapping and alignment results.
[0070] In practice, when locating abnormal data in gene sequencing data, the abnormal data location processing can be performed on the full gene sequencing data.
[0071] In another embodiment, a single gene sequencing operation may sequence tens of thousands or even hundreds of thousands of sample libraries, resulting in gene fragment sequences in the sequencing data reaching tens or even hundreds of thousands. Furthermore, when any part of the gene sequencing system—optical detection, sequencing chip, reagents, fluidic circuitry, or mechanics—experiences anomalies, it typically affects the gene fragment sequences of a large area of the sample library within the imaging unit. This manifests as anomalies in gene fragment sequences detected from a specific region within the imaging unit. Therefore, by performing anomaly localization processing on a portion of the gene sequencing data, the specific source of the anomaly in the gene sequencing system can be pinpointed. Thus, to reduce the computational load required for anomaly localization, anomaly localization processing can be performed on a portion of the full gene sequencing data.
[0072] Furthermore, in this embodiment of the disclosure, the gene sequencing data used for anomaly localization in the gene sequencing system may include, for example, at least one of the following: b1: Based on the flow channel, the sequencing images are sampled, and the gene sequencing data of the fluorescent spots in the sequencing images of each sampled flow channel are used as gene sequencing data for anomaly localization.
[0073] Here, for example, sequencing images can be sampled along the channel dimension. That is, for each channel, a sequencing image corresponding to that channel can be sampled, and anomaly localization processing can be performed on the gene sequencing data corresponding to the sampled sequencing images. In this way, the sampled sequencing images can cover all channels. When an anomaly exists in any region of a channel, such as the presence of air bubbles, the appearance of the region with air bubbles projected onto the sequencing image differs significantly from the fluorescence projection of normal clusters. Thus, based on the results of anomaly localization, the problem in the channel can be identified.
[0074] b2: Based on the sampling of sequencing images by the camera, the gene sequencing data of the fluorescent spots in the sequencing images of each camera are used as gene sequencing data for anomaly localization.
[0075] Here, for example, gene sequencing data can be sampled at the camera level. That is, for each camera, the sequencing image corresponding to that camera is sampled, and anomaly localization processing is performed on the gene sequencing data of the sampled sequencing images. Specifically, for each camera, for example, multiple imaging units are captured in each sequencing cycle of multiple sequencing cycles, resulting in sequencing images of multiple imaging units in each sequencing cycle; for each camera, the sequencing images captured by each camera are sampled to obtain a sampled sequencing image corresponding to each camera. In this way, if an anomaly exists in a certain camera, the anomaly will be manifested in all sequencing images captured by that camera, resulting in abnormal light spots in the sequencing images captured by that camera, thus allowing the anomaly to be located at a specific camera; if a component in the optical system has a problem, this problem may affect some or all cameras, resulting in abnormal light spots in the sequencing images captured by these cameras, thus allowing the anomaly to be located at a specific component in the optical system.
[0076] b3: The sequence of at least one end of a gene fragment from pair-end sequencing (PE) data.
[0077] Furthermore, the gene fragment sequences used for locating abnormal data in the gene sequencing system can also be gene fragment sequences from other sources or in different groupings, and this disclosure does not limit them.
[0078] When mapping and aligning gene fragment sequences included in gene sequencing data with reference sequences to obtain the mapping and alignment results of each gene fragment sequence, the reference sequence can be a reference genome sequence containing all the genetic information of a species, or it can be a known sequence synthesized from one or more segments. For example, the human reference genome sequence contains the complete DNA sequences of 22 autosomes and a pair of sex chromosomes, as well as mitochondrial DNA and other genetic material; another example is the reference genome sequence of bacteriophage phix174, which contains the complete sequence of a circular DNA molecule. The specific sequence can be determined according to actual needs. For example, open-source software such as bwa, bowtie, blast, and soap2, commercial software such as sentieon and dragon, or other mapping and alignment algorithms can be used during mapping and alignment. This disclosure does not limit the specific algorithms used.
[0079] The gene fragment sequences used for anomaly localization in gene sequencing systems retain only the optimal mapping and alignment results; that is, only the gene fragment sequences that have successfully aligned with the sequence and have the least degree of difference are retained. For example... Figure 8The image shows a specific example of mapping and aligning a portion of a gene fragment sequence with a reference sequence; in this example, "reference" represents the reference sequence; there are 10 gene fragment sequences, namely: Read1~Read10. Among them, Read1~Read7 were all successfully aligned, while Read8~Read10 were all unaligned.
[0080] After obtaining the mapping and alignment results of each gene fragment sequence, at least one of the following methods c1 to c2 can be used to determine the abnormal gene fragment sequence from the gene fragment sequence based on the mapping and alignment results: c1: If the mapping and alignment results of the gene fragment sequence indicate that the alignment was unsuccessful, the gene fragment sequence is identified as the abnormal gene fragment sequence.
[0081] Here, if the alignment result indicates that the alignment was unsuccessful, it means that there is a problem with the identification of the gene fragment sequence, and therefore it is regarded as an abnormal gene fragment sequence.
[0082] c2: If the mapping and alignment results of the gene fragment sequence indicate that the alignment is successful, count the number of unaligned bases in the gene fragment sequence and the reference sequence. If the number of unaligned bases is greater than a preset threshold, the gene fragment sequence is identified as the abnormal gene fragment sequence. Here, the alignment result indicates that the alignment was successful, meaning that the gene fragment sequence was successfully mapped to a certain position in the reference sequence. However, if the difference between the gene fragment sequence and the reference sequence is too large, it also proves that the gene fragment sequence may also have some abnormalities. Therefore, this type of gene fragment sequence will also be regarded as an abnormal gene fragment sequence.
[0083] Here, when determining the differences between a gene fragment sequence and a reference sequence, the following methods can be used, for example: The first base in the gene fragment sequence is taken as the base site in the reference sequence. Starting from this first base site, N bases are read from the reference sequence to form a subsequence; N represents the number of bases in the gene fragment sequence. The subsequence and the gene fragment sequence are aligned position by position to obtain the number of differential bases between the subsequence and the gene fragment sequence; Based on the number of differing bases, the difference information between the gene fragment sequence and the reference sequence is determined.
[0084] Here, the number of different bases can be directly used as differential information; alternatively, the proportion of the number of different bases to the total number of bases in the gene fragment sequence can also be used as differential information.
[0085] After obtaining the difference information between the gene fragment sequence and the reference sequence, the difference information is compared with a preset difference threshold; if the difference information is greater than or equal to the difference threshold, the gene fragment sequence is regarded as an abnormal gene fragment sequence.
[0086] like Figure 9 In the example shown, it is assumed that the number of differential bases is used as differential information and the differential threshold is 3. The number of differential bases between Read2 and the reference sequence is 3, which is equal to the differential threshold. The number of differential bases between Read6 and the reference sequence is 4, which is greater than the differential threshold. The number of differential bases between Read1, Read4, Read7 and the reference sequence is 2 each. The number of differential bases between Read3 and the reference sequence is 1, which is less than the differential threshold. Therefore, Read2 and Read6 are considered as abnormal gene fragment sequences.
[0087] Furthermore, since individual genes can contain mutations, differences in base pairs between gene fragment sequences and reference sequences could be caused by sequencing errors or gene mutations in the individuals from which the sample library was derived. Therefore, to minimize the impact of gene mutations on the evaluation results of the gene sequencing system, in this embodiment, the sample library can be derived from species with low mutation rates or fixed mutation sites, such as phix174.
[0088] In addition to any of the above-mentioned c1~c2, other mapping and alignment indicators can be selected and thresholds can be set as standards for judging the quality of sequence sequencing, such as the identity value in the mapping and alignment result file of BLAST software, the mapQ value in the mapping and alignment result file of BWA software, etc.
[0089] Furthermore, in another embodiment of this disclosure, before mapping and comparing the gene fragment sequences of the sample library with the reference sequence, at least one source information is used as a dimension for grouping and summarizing to group the gene fragment sequences, thereby obtaining multiple groups of gene fragment sequences.
[0090] Here, the gene sequencing data also includes at least one of the following information regarding the source of the gene fragment sequence: the upper and lower surfaces of the sequencing chip, the flow channels of the sequencing chip, the side columns of the flow channels, and the camera used to take the picture.
[0091] After grouping gene fragment sequences using at least one source of information as a grouping and summarizing dimension, multiple groups of gene fragment sequences can be obtained. Then, by mapping and aligning the gene fragment sequences of the sample library with reference sequences to obtain mapping and alignment results, and based on these alignment results, abnormal gene fragment sequences can be identified from the gene fragment sequences. This can be done by mapping and aligning the gene fragment sequences in each group with the reference sequence, obtaining mapping and alignment results for each group, and then identifying the abnormal gene fragment sequences from each group based on these mapping and alignment results.
[0092] Here, gene fragment sequences are pre-grouped before the process of identifying abnormal gene fragment sequences is carried out. This ensures that the abnormal gene fragment sequences are already grouped after identification, allowing subsequent sequencing images marked with abnormal spots to be generated based on different grouping dimensions. This facilitates the determination of specific abnormalities in the gene sequencing system from different dimensions.
[0093] Following S402 above, the abnormal data localization method for the gene sequencing system provided in this embodiment of the disclosure further includes: S403: Based on the gene sequencing data, determine the imaging unit where the abnormal gene fragment sequence is located and the position of the abnormal gene fragment sequence in the sequencing image corresponding to the imaging unit. In the sequencing image corresponding to the imaging unit, mark the abnormal light spot at the position of the abnormal gene fragment sequence to obtain a sequencing image marked with the abnormal light spot.
[0094] In practical implementation, sequencing images marked with abnormal light spots can be generated, for example, in the following manner: Using at least one source of information as a grouping and summarizing dimension, the abnormal gene fragment sequences are grouped to obtain multiple groups of abnormal gene fragment sequences; Based on the gene sequencing data, the imaging unit where the abnormal gene fragment sequence in each group is located and the position of the abnormal gene fragment sequence in each group in the sequencing image corresponding to the imaging unit are determined. In the sequencing image corresponding to the imaging unit, abnormal light spots at the positions of the abnormal gene fragment sequences in each group are marked to obtain sequencing images marked with abnormal light spots.
[0095] In practice, each sequencing image may have information from multiple sources, such as the top and bottom surfaces of the sequencing chip, the flow channels of the sequencing chip, the side columns of the flow channels, the imaging units of the side columns, the camera that took the picture, the red-green channel (or the blue-green channel), and at least one of the sequencing cycles.
[0096] Specifically, in each sequencing cycle, the sample library needs to undergo a series of chemical reactions with reagents before sequencing images are acquired. When acquiring sequencing images, the sample library is distributed in multiple sealed channels of the sequencing chip, and each sealed channel may have the sample library distributed on a single surface or on both the upper and lower surfaces. Each surface contains multiple side columns, and each side column contains multiple imaging units.
[0097] When grouping abnormal gene fragment sequences based on at least one of the above-mentioned source information, at least one of the following d1~d6 information can be obtained first: d1: Chip surface information: The chip has two surfaces, top and bottom. The chip surface information is used to indicate which surface of the chip the camera photographed to produce the sequencing image.
[0098] d2: Flow channel information of the chip: used to indicate which flow channel in the chip the camera photographed to generate the sequencing image.
[0099] d3: Lateral column information of the flow channel, used to indicate which lateral column of the flow channel the sequencing image was taken from to produce the imaging unit.
[0100] d4: Lateral imaging unit information, used to indicate which imaging unit in the lateral column the camera photographed to produce the sequencing image.
[0101] d5: Camera information: Used to indicate which camera was used to take the image.
[0102] d6: Red-green channel information: Used to indicate which excitation light channel (red or green) was used to capture the sequencing image.
[0103] d7: Sequencing cycle number: Used to indicate in which sequencing cycle the sequencing image was taken.
[0104] The aforementioned information can be carried in the imaging number of each sequencing image, which is used to uniquely identify the specific source of a sequencing image. For example, at least one of the numbers d1 to d7 can be used to uniquely identify the sequencing image by comprehensively including, for example, the camera information, chip surface information, channel side column information, and side column imaging unit information.
[0105] Furthermore, other dimensions can also be used when grouping abnormal gene fragment sequences.
[0106] In the field of gene sequencing, flanking is an advanced concept related to mass spectrometry technology, which is based on liquid chromatography-mass spectrometry (LC-MS) platforms. Its core principle is to divide the entire mass-to-charge ratio (m / z) range into a series of continuous, narrow windows (commonly called flanking windows) during mass spectrometry detection. During LC separation, compounds in the sample sequentially enter the mass spectrometer along with the mobile phase. The mass spectrometer then collects and detects ions within each flanking window at set time intervals.
[0107] When grouping abnormal gene fragment sequences, the abnormal gene fragment sequences corresponding to different side columns are classified into different categories according to the side column windows.
[0108] When grouping anomalous gene fragment sequences, they can be grouped based on a grouping dimension that includes at least one source of information. This grouping of anomalous gene fragment sequences under different dimensions allows for analysis of the causes of anomalous gene fragment sequences generated by the gene sequencing system from different perspectives. This grouping dimension can be determined based on the analytical needs of the gene sequencing system. For example, to determine the impact of a camera on the gene sequencing system, anomalous gene fragment sequences can be grouped based on different cameras; similarly, to determine the impact of sequencing channels on the gene sequencing system, anomalous gene fragment sequences can be grouped based on different sequencing channels.
[0109] In addition, abnormal gene fragment sequences can be grouped based on multiple grouping dimensions to obtain grouping results corresponding to each dimension. That is, for each grouping dimension, there is a corresponding grouping result for one type of abnormal gene fragment sequence.
[0110] In this way, by setting multiple grouping dimensions, the impact of different structures in different gene sequencing systems on gene sequencing can be analyzed separately, so as to confirm the exact cause of abnormal gene fragment sequences in the gene sequencing system.
[0111] Subsequently, multiple sets of abnormal gene fragment sequences can be traversed, and based on the imaging unit information corresponding to each abnormal gene fragment sequence in each traversed abnormal gene fragment sequence, the position information of the gene fragment sequence in the imaging unit, and the correspondence between each set of abnormal gene fragment sequences and the image group, the light spot corresponding to the abnormal gene fragment sequence in each traversed abnormal gene fragment sequence in the sequencing image included in the corresponding image group is marked.
[0112] This disclosure also provides a specific method for marking abnormal light spots at the locations of the abnormal gene fragment sequences in a sequencing image corresponding to an imaging unit, thereby obtaining a sequencing image marked with abnormal light spots, including: In the case where the abnormal gene fragment sequence is not successfully aligned, one sequencing image is selected from all sequencing images of all sequencing cycles corresponding to the imaging unit. In the selected sequencing image, the abnormal light spot at the position of the abnormal gene fragment sequence is marked to obtain a sequencing image marked with the abnormal light spot. or, If the number of bases in the abnormal gene fragment sequence that were successfully aligned but not aligned is greater than a preset threshold, the sequencing image of the sequencing cycle with the erroneous bases is selected from all sequencing images of all sequencing cycles corresponding to the imaging unit. In the selected sequencing image, the abnormal light spot at the position of the abnormal gene fragment sequence is marked to obtain a sequencing image marked with the abnormal light spot.
[0113] Specifically, when labeling the light spots, for each abnormal gene fragment sequence, if the abnormal gene fragment sequence is successfully aligned, all sequencing images corresponding to the corresponding imaging unit can be selected based on its imaging unit information. Sequencing images from one sequencing cycle can be randomly selected for labeling the abnormal light spots. Here, for multiple unaligned abnormal gene fragment sequences corresponding to the same imaging unit, the abnormal light spots corresponding to these multiple unaligned abnormal gene fragment sequences can be labeled onto the same sequencing image.
[0114] If the number of bases in an abnormal gene fragment sequence that were successfully aligned but not aligned is greater than a preset threshold, assuming that the corresponding imaging unit includes sequencing images corresponding to N sequencing cycles, and the i-th and j-th bases in the abnormal gene fragment sequence were not aligned successfully, i.e., an error occurred, then the abnormal spot corresponding to the abnormal gene fragment sequence will be marked in the sequencing images of the i-th and j-th sequencing cycles.
[0115] like Figure 9 The image shows a specific example of a sequencing image marked with abnormal spots. In this image, "·" represents a spot from the original sequencing image, and each spot represents one sequence or one base. Abnormal gene fragment sequences are marked with "×". Different colors or shapes can be used to mark abnormal gene fragment sequences determined by different methods on a single sequencing image for easy differentiation; or abnormal gene fragment sequences determined by different methods can be marked on two separate sequencing images. For example, abnormal gene fragment sequences determined by method c1 can be marked on one sequencing image, and abnormal gene fragment sequences determined by method c2 can be marked on another sequencing image, making classification and interpretation easier.
[0116] In this way, based on the image patch identifiers and location information of abnormal gene fragment sequences, sequences with mapping failures and mismatches meeting the threshold condition are marked in the sequencing image. This maps general indicators such as mapping rate and base error rate onto the sequencing image, allowing for intuitive, clear, and specific localization of low-quality sequencing results. Locating mismatched bases in each sequencing cycle to the corresponding cycle's sequencing image allows for the localization of sequencing quality for each cycle. Based on these localizations, it is helpful to analyze and locate potential problems in related modules of the sequencing system.
[0117] In another embodiment of this disclosure, in the acquired sequencing image, abnormal light spots are marked at the positions of the abnormal gene fragment sequences in the imaging unit to obtain a sequencing image marked with abnormal light spots, including: For each of the multiple imaging units, a labeled sequencing image corresponding to that imaging unit is determined; wherein, the labeled sequencing image includes: a sequencing image of at least one sequencing cycle corresponding to that imaging unit; In the labeled sequencing image, abnormal light spots are marked at the positions of the abnormal gene fragment sequences in the imaging unit to obtain a sequencing image with labeled abnormal light spots.
[0118] Here, each imaging unit undergoes multiple sequencing cycles, resulting in sequencing images corresponding to each cycle. Analysis of these sequencing images from multiple cycles yields the gene fragment sequences of each sample library within the imaging unit. When marking abnormal spots, at least one labeled sequencing image is selected from the multiple sequencing images corresponding to the imaging unit. This labeled sequencing image is then used to mark the abnormal spots, describing the specific distribution of gene fragment sequences in each sample library corresponding to the imaging unit.
[0119] For example, for a gene fragment sequence that failed to align, any image can be randomly selected from the sequencing images corresponding to each sequencing cycle to be used as a labeled sequencing image for marking abnormal spots. As another example, for a gene fragment sequence that aligned successfully but contains a certain number of incorrect bases, all sequencing images containing incorrect bases can be used as labeled sequencing images. Alternatively, labeled sequencing images can be determined from the sequencing images corresponding to each sequencing cycle based on image quality, such as sharpness and signal-to-noise ratio. The specific method of determination is not limited in the embodiments of this disclosure.
[0120] In this way, multiple sequencing images of anomalous spots can be generated for each group dimension, with each image unit as the unit.
[0121] In another embodiment of this disclosure, after obtaining the sequencing image marked with abnormal light spots, the sequencing image marked with abnormal light spots can be divided into multiple sub-images. The sequencing image marked with abnormal light spots is divided into multiple sub-images; According to different anomaly types, the number or proportion of unaligned abnormal gene fragment sequences and successfully aligned but unaligned bases in the sub-images are statistically analyzed. An anomaly indicator heatmap is established based on the number or proportion of abnormal gene fragment sequences of different anomaly types in the sub-image; wherein, the anomaly indicator heatmap includes anomaly indicator information corresponding to each sub-image; the anomaly indicator information characterizes the number or proportion of unaligned abnormal gene fragment sequences and successfully aligned but unaligned bases in the sub-image through pixel values.
[0122] In specific implementation, the abnormal type of abnormal gene fragment sequence is, for example, the abnormal type corresponding to any of the abnormal gene fragment sequence determination methods in c1~c2 mentioned above.
[0123] For example, the abnormal types of abnormal gene fragment sequences determined based on c1 above include: abnormal gene fragment sequences that failed to align; the abnormal types of abnormal gene fragment sequences determined based on c2 above include: sequences that successfully align but the number of mismatches meets the threshold condition.
[0124] Sequencing images marked with abnormal spots are partitioned, and the partition size can be changed as needed. Within each partition (corresponding to the spot-marked sub-image), the number of abnormal gene fragment sequences of different abnormal types or their proportion of the total number of spots in that partition (including spots corresponding to abnormal gene fragment sequences and spots corresponding to non-abnormal gene fragment sequences) are counted, and abnormality indicator heatmaps representing abnormal base sequences of different abnormal types are plotted.
[0125] like Figure 10 The image shows a specific example of an anomaly indicator heatmap. Thus, in sequencing images marked with anomalous spots, it is immediately clear which regions have a high number of failed mapping sequences, which regions have a high number of mismatches meeting the threshold, and which regions have a high number of mismatched bases, based on the specific location of the mismatches. After this step, the specific region where the low-quality sequence is located in the sequencing image marked with anomalous spots can be clearly identified.
[0126] Next, the anomaly indicator heatmap can be visualized and displayed to the user. Based on this heatmap, the user can calculate the proportion of abnormal gene fragment sequences of different anomaly types and the proportion of mismatched bases in the gene fragment sequences belonging to each image block, thus obtaining the overall sequencing quality of the image block region. Based on the above anomaly indicator heatmap, the image block regions with sequencing anomalies are output. At the same time, based on the different modules indicated by the classification dimension in the gene sequencing system, it is determined whether there are abnormalities in the reaction of some regions of the chip flow channel, abnormal differences in sequencing cycles, differences or anomalies in different flow channels or cameras, etc., thereby locating phenomena such as sequencing chip anomalies, camera imaging anomalies, impurities in sequencing chips or sequencing reagents, and air bubbles introduced into the liquid path pipelines.
[0127] In this way, sequences with mapping failures and mismatches meeting the threshold, or erroneous bases in each sequencing cycle, are labeled onto the sequencing image. The sequencing image is then partitioned, and the number or proportion of sequences with mapping failures and mismatches meeting the threshold, or erroneous bases in each sequencing cycle, are counted within each partition. A heatmap is then created, providing a more intuitive representation of the location of low-quality sequences in the sequencing image. Other statistical plotting techniques can also be used to statistically analyze and plot low-quality sequences or bases, achieving a more intuitive understanding of the correlation between data quality and sequencing images.
[0128] The abnormal data localization method for a gene sequencing system provided in this disclosure involves acquiring gene sequencing data obtained by sequencing a sample library of one or more samples using a gene sequencing system. The gene sequencing data includes gene fragment sequences from the sample library, imaging unit information of the gene fragment sequences, and position information of the gene fragment sequences within the imaging units. The method maps and aligns the gene fragment sequences in the gene sequencing data with a reference genome sequence to obtain mapping and alignment results. Based on these results, abnormal gene fragment sequences are identified from the gene fragment sequences. The method then determines the imaging unit of the abnormal gene fragment sequence and its position in the sequencing image corresponding to the imaging unit based on the gene sequencing data. Abnormal light spots are marked at the positions of the abnormal gene fragment sequences in the sequencing image corresponding to the imaging unit, resulting in a sequencing image marked with abnormal light spots. In this way, low-quality or abnormal gene fragment sequences can be accurately located in sequencing images using gene sequencing data obtained from gene sequencing of sample libraries. This results in sequencing images marked with abnormal spots, which can reflect the distribution of abnormal gene fragment sequences. This allows researchers to analyze the reasons for low sequence quality or abnormality based on the phenomena observed in different sequencing cycles or positions in the sequencing images, thus achieving fine-grained localization of low-quality gene fragment sequences.
[0129] Furthermore, the abnormal data localization method for gene sequencing systems provided in this disclosure uniquely combines traditional or conventional sequencing data evaluation indicators with sequencing images to achieve accurate backtracking and localization of low-quality sequencing data. This method can be applied not only to species with reference genomes but also supports customizable specific sequences.
[0130] Considering that the evaluation and improvement of sequencing systems are affected by multiple module factors, the abnormal data localization method for gene sequencing systems provided in this disclosure can test and initially screen the data produced by different sequencing systems, different library types, and different library throughput sequencing systems. It can not only preliminarily evaluate the sequencing system, but also comprehensively and effectively evaluate the sequencing system accurately and locate potential problem modules, thus having stronger reliability and universality.
[0131] The abnormal data localization method for gene sequencing systems provided in this disclosure is not only applicable to next-generation sequencing systems that perform base identification based on sequencing images, but also to other sequencing systems that perform base identification based on sequencing images, requiring only the sequencing images produced by the sequencing system. Low-quality sequence data or bases produced by the sequencing system can be directly mapped onto the sequencing images, helping R&D or quality assessment personnel quickly locate problems in the sequencing system, promoting targeted improvements to the sequencing system, saving manpower costs for problem identification, improving the production efficiency of the sequencing system, and having no requirements on sequencing libraries, sequencing throughput, sequence length, etc., thus possessing greater practicality and a wider range of applications.
[0132] See Figure 11 As shown in the embodiments of this disclosure, an abnormal data localization system for a gene sequencing system is also provided. The system includes: a sequence mapping and alignment module, an image localization module, and a visualization module.
[0133] The sequence mapping and alignment module is connected to the sequencing system module, which is also the gene sequencing system. It is used to output gene sequencing data and sequencing images obtained by gene sequencing of sample libraries using the gene sequencing system to the sequence mapping and alignment module.
[0134] The sequence mapping and alignment module is used to align the gene fragment sequences included in the gene sequencing data with the reference genome sequence, obtain the alignment results corresponding to each gene fragment sequence, and determine the abnormal gene fragment sequences from the gene fragment sequences based on the alignment results.
[0135] The image localization module is used to mark the light spots corresponding to the abnormal gene fragment sequences in the sequencing image based on the image block identifiers corresponding to each abnormal gene fragment sequence and the position information of the light spots corresponding to the abnormal gene fragment sequences in the sequencing image, thereby obtaining a sequencing image marked with abnormal light spots.
[0136] The visualization module is used to visualize sequencing images marked with abnormal light spots and / or abnormality indicator heatmaps, so that users can determine the cause of abnormalities in the gene sequencing system.
[0137] Those skilled in the art will understand that, in the above-described method of the specific implementation, the order in which each step is written does not imply a strict execution order and does not constitute any limitation on the implementation process. The specific execution order of each step should be determined by its function and possible internal logic.
[0138] This disclosure also provides a computer device, such as... Figure 12 The diagram shown is a schematic representation of a computer device structure provided in an embodiment of this disclosure, including: Processor 121 and memory 122; the memory 122 stores machine-readable instructions executable by processor 121, and processor 121 executes the machine-readable instructions stored in memory 122. When the machine-readable instructions are executed by processor 121, processor 121 performs the following steps: Acquire gene sequencing data by sequencing a sample library of one or more samples using a gene sequencing system. The gene sequencing data includes gene fragment sequences of the sample library, imaging unit information of the gene fragment sequences, and position information of the gene fragment sequences in the imaging units. The gene fragment sequences in the gene sequencing data are compared with the reference genome sequence to obtain the comparison results of the gene fragment sequences. Based on the comparison results, abnormal gene fragment sequences are identified from the gene fragment sequences. Based on gene sequencing data, the imaging unit of the abnormal gene fragment sequence and the position of the abnormal gene fragment sequence in the sequencing image corresponding to the imaging unit are determined. In the sequencing image corresponding to the imaging unit, abnormal light spots at the position of the abnormal gene fragment sequence are marked to obtain a sequencing image marked with abnormal light spots.
[0139] The aforementioned memory 122 includes a main memory 1221 and an external memory 1222; the main memory 1221, also known as internal memory, is used to temporarily store the computational data in the processor 121, as well as the data exchanged with external memory 1222 such as a hard disk. The processor 121 exchanges data with the external memory 1222 through the main memory 1221.
[0140] The specific execution process of the above instructions can be referred to the steps of the abnormal data localization method of the gene sequencing system described in the embodiments of this disclosure, and will not be repeated here.
[0141] This disclosure also provides a computer-readable storage medium storing a computer program. When executed by a processor, the computer program performs the steps of the abnormal data localization method for the gene sequencing system described in the above-described method embodiments. The storage medium can be a volatile or non-volatile computer-readable storage medium.
[0142] This disclosure also provides a computer program product carrying program code. The program code includes instructions that can be used to execute the steps of the abnormal data localization method of the gene sequencing system described in the above method embodiments. For details, please refer to the above method embodiments, which will not be repeated here.
[0143] The aforementioned computer program product can be implemented through hardware, software, or a combination thereof. In one optional embodiment, the computer program product is specifically embodied in a computer storage medium; in another optional embodiment, the computer program product is specifically embodied in a software product, such as a software development kit (SDK), etc.
[0144] Finally, it should be noted that the above-described embodiments are merely specific implementations of this disclosure, used to illustrate the technical solutions of this disclosure, and not to limit it. The protection scope of this disclosure is not limited thereto. Although this disclosure has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that any person skilled in the art can still modify or easily conceive of changes to the technical solutions described in the foregoing embodiments, or make equivalent substitutions for some of the technical features, within the scope of the technology disclosed in this disclosure; and these modifications, changes, or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this disclosure, and should all be covered within the protection scope of this disclosure. Therefore, the protection scope of this disclosure should be determined by the protection scope of the claims.
Claims
1. A method for locating abnormal data in a gene sequencing system, characterized in that, include: Acquire gene sequencing data by sequencing a sample library of one or more samples using a gene sequencing system. The gene sequencing data includes gene fragment sequences of the sample library, imaging unit information of the gene fragment sequences, and position information of the gene fragment sequences in the sequencing image corresponding to the imaging unit. The gene fragment sequences of the sample library are mapped and aligned with the reference sequences to obtain the mapping and alignment results of the gene fragment sequences. Based on the mapping and alignment results, abnormal gene fragment sequences are identified from the gene fragment sequences. Based on gene sequencing data, the imaging unit where the abnormal gene fragment sequence is located and the position of the abnormal gene fragment sequence in the sequencing image corresponding to the imaging unit are determined. In the sequencing image corresponding to the imaging unit, abnormal light spots at the position of the abnormal gene fragment sequence are marked to obtain a sequencing image marked with abnormal light spots.
2. The method according to claim 1, characterized in that, The step of determining the abnormal gene fragment sequence from the gene fragment sequence based on the mapping and alignment results includes at least one of the following: If the mapping and alignment results of the gene fragment sequence indicate that the alignment was unsuccessful, the gene fragment sequence will be identified as the abnormal gene fragment sequence. If the mapping and alignment results of the gene fragment sequence indicate a successful alignment, the number of unaligned bases in the gene fragment sequence and the reference sequence is counted. If the number of unaligned bases is greater than a preset threshold, the gene fragment sequence is identified as the abnormal gene fragment sequence.
3. The method according to claim 1 or 2, characterized in that, The gene sequencing data also includes at least one of the following information regarding the origin of the gene fragment sequence: the top and bottom surfaces of the sequencing chip, the flow channels of the sequencing chip, the side columns of the flow channels, and the camera used to take the picture.
4. The method according to claim 3, characterized in that, The method further includes: Before mapping and aligning the gene fragment sequences of the sample library with the reference sequences, the gene fragment sequences are grouped using at least one source information as a dimension for grouping and summarizing, resulting in multiple groups of gene fragment sequences. The step of mapping and aligning the gene fragment sequences of the sample library with reference sequences to obtain the mapping and alignment results of the gene fragment sequences, and determining abnormal gene fragment sequences from the gene fragment sequences based on the alignment results, includes: The gene fragment sequences in each group are mapped and compared with the reference sequence to obtain the mapping and comparison results of the gene fragment sequences in each group. Based on the mapping and comparison results, the abnormal gene fragment sequences in each group are identified.
5. The method according to claim 3, characterized in that, The method further includes: After mapping and comparing the gene fragment sequences of the sample library with the reference sequences, at least one source information is used as a grouping and summarizing dimension to group the abnormal gene fragment sequences, resulting in multiple groups of abnormal gene fragment sequences. The step of determining the imaging unit where the abnormal gene fragment sequence is located and the position of the abnormal gene fragment sequence in the sequencing image corresponding to the imaging unit based on gene sequencing data, and marking the abnormal light spot at the position of the abnormal gene fragment sequence in the sequencing image corresponding to the imaging unit to obtain a sequencing image marked with the abnormal light spot includes: Based on the gene sequencing data, the imaging unit where the abnormal gene fragment sequence in each group is located and the position of the abnormal gene fragment sequence in each group in the sequencing image corresponding to the imaging unit are determined. In the sequencing image corresponding to the imaging unit, abnormal light spots at the positions of the abnormal gene fragment sequences in each group are marked to obtain sequencing images marked with abnormal light spots.
6. The method according to any one of claims 1-5, characterized in that, The step of marking abnormal light spots at the locations of the abnormal gene fragment sequences in the sequencing image corresponding to the imaging unit to obtain a sequencing image marked with abnormal light spots includes: In the case where the abnormal gene fragment sequence is not successfully aligned, one sequencing image is selected from all sequencing images of all sequencing cycles corresponding to the imaging unit. In the selected sequencing image, the abnormal light spot at the position of the abnormal gene fragment sequence is marked to obtain a sequencing image marked with the abnormal light spot. or, If the number of bases in the abnormal gene fragment sequence that were successfully aligned but not aligned is greater than a preset threshold, the sequencing image of the sequencing cycle with the erroneous bases is selected from all sequencing images of all sequencing cycles corresponding to the imaging unit. In the selected sequencing image, the abnormal light spot at the position of the abnormal gene fragment sequence is marked to obtain a sequencing image marked with the abnormal light spot.
7. The method according to any one of claims 2-6, characterized in that, The method further includes: The sequencing image marked with abnormal light spots is divided into multiple sub-images; According to different anomaly types, the number or proportion of unaligned abnormal gene fragment sequences and successfully aligned but unaligned bases in the sub-images are statistically analyzed. An anomaly indicator heatmap is established based on the number or proportion of abnormal gene fragment sequences of different anomaly types in the sub-image; wherein, the anomaly indicator heatmap includes anomaly indicator information corresponding to each sub-image; the anomaly indicator information characterizes the number or proportion of unaligned abnormal gene fragment sequences and successfully aligned but unaligned bases in the sub-image through pixel values.
8. The method according to any one of claims 2-6, characterized in that, The method further includes: According to different abnormality types, the number or proportion of abnormal gene fragment sequences present in the imaging unit is counted; The number or proportion of abnormal gene fragment sequences is compared with a set threshold, and the overall data quality of the imaging unit is determined based on the comparison results.
9. A computer device, characterized in that, include: The processor and the memory, wherein the memory stores machine-readable instructions executable by the processor, the processor is configured to execute the machine-readable instructions stored in the memory, and when the machine-readable instructions are executed by the processor, the processor performs the steps of the abnormal data localization method of the gene sequencing system as described in any one of claims 1 to 8.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program, which, when executed by a computer device, performs the steps of the abnormal data localization method for the gene sequencing system as described in any one of claims 1 to 8.