Base identification method based on breeding chip, scanning equipment and storage medium
By acquiring fluorescence images of the breeding chip, extracting probe scatter data features, and using a machine learning model to calibrate base identification, the problems of low probe signal-to-noise ratio and bias in the breeding chip were solved, and the accuracy and stability of base discrimination were improved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- SHENZHEN SALUS BIOMED CO LTD
- Filing Date
- 2026-02-03
- Publication Date
- 2026-05-12
AI Technical Summary
Existing breeding chips suffer from problems such as low probe signal-to-noise ratio, bias, and blurred boundaries between heterozygous and homozygous bases during base recognition, which affect the accuracy of detection sites.
By acquiring fluorescence images of the breeding chip, extracting the scatter data features of each type of probe, forming a pre-trained base recognition model, and using machine learning models such as neural networks and deep learning, calibrating the correspondence between brightness signals and base types, correcting probe-specific noise, and improving the accuracy and stability of base discrimination.
It significantly improves the accuracy, stability, and applicability of base identification, with an accuracy increase of 4-10 percentage points compared to traditional methods.
Smart Images

Figure CN122023918A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of breeding technology, and in particular to a base identification method, scanning device, and computer-readable storage medium based on a breeding chip. Background Technology
[0002] With the rapid development of molecular biology techniques, high-throughput molecular marker detection has become one of the core technologies in modern agricultural breeding. Breeding chips, as a key tool, can detect tens of thousands of single nucleotide polymorphism (SNP) sites at once, greatly accelerating the efficiency of breeding processes such as genotyping, genetic mapping, locating genes for important traits, and genome-wide selection.
[0003] In practical applications of breeding chips, specialized scanning equipment is needed to acquire fluorescence images after the chip hybridization reaction. These images contain fluorescent signal dots (hereinafter referred to as "fluorescent dots"). The intensity information of each fluorescent dot directly corresponds to a specific molecular marker and its genotype information. Therefore, rapid and accurate localization and quantification of each fluorescent dot in the image is the primary prerequisite and key technical step to ensure accurate and reliable genotyping results.
[0004] However, in the entire process of preparing and detecting agricultural breeding chips, the probe signals are subject to some instabilities, such as: 1) some probes have low signal-to-noise ratios; 2) some probes have certain biases; 3) the boundary between heterozygous and homozygous is relatively blurred. These factors can affect the accuracy of base identification at the detection sites. Summary of the Invention
[0005] To address the existing technical problems, this invention provides a base identification method scanning device and computer-readable storage medium based on breeding chips, which can significantly improve the accuracy, stability and applicability of base identification.
[0006] In a first aspect, a base identification method based on a breeding chip is provided, comprising: acquiring a fluorescence image of the breeding chip including multiple different base channels; based on the fluorescence image, acquiring the brightness value of each target probe scatter point in each base channel of each type of probe; extracting the scatter point data features of each type of probe based on the brightness value of each target probe scatter point in each base channel of each type of probe; forming the input of a pre-trained base identification model based on the scatter point data features of each type of probe, and outputting the base data of the target detection site detected by each type of probe.
[0007] In a second aspect, a scanning device is provided, including a memory and a processor. The memory stores a computer program, which, when executed by the processor, causes the processor to perform the steps of the base recognition method based on a breeding chip provided in the embodiments of this application.
[0008] Thirdly, a computer-readable storage medium is provided, storing a computer program that, when executed by a processor, causes the processor to perform the steps of the base recognition method based on a breeding chip provided in the embodiments of this application.
[0009] This application acquires fluorescence images, and based on these images, obtains the brightness values of each target probe scatter point in each base channel for each type of probe. Based on the brightness values of each target probe scatter point in each base channel for each type of probe, it extracts the scatter data features of each type of probe. The scatter data features represent the base brightness features of the target detection sites captured by that type of probe. Based on the scatter data features of each type of probe, it forms the input and output base data of the target detection sites for a pre-trained base recognition model. The core advantage is that it deeply binds the specific features of the probe with the base recognition of the site, greatly improving the accuracy, stability and applicability of base discrimination. Attached Figure Description
[0010] Figure 1 This is a schematic diagram of the scatter distribution of the two probes; Figure 2 This is an application environment diagram of a base recognition method based on a breeding chip in one embodiment; Figure 3 This is a flowchart of a base recognition method based on a breeding chip in one embodiment; Figure 4 This is a scatter plot of probe points before and after preprocessing in one embodiment; Figure 5 This is a comparative diagram showing the process before and after removing outlier data points in one embodiment; Figure 6 This is a schematic diagram of a base recognition device based on a breeding chip in one embodiment; Figure 7 This is a schematic diagram of the structure of a scanning device in one embodiment. Detailed Implementation
[0011] The technical solution of the present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments.
[0012] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains. The terminology used herein in the specification of this invention is for the purpose of describing particular embodiments only and is not intended to limit the scope of the invention. The term "and / or" as used herein includes any and all combinations of one or more of the associated listed items.
[0013] In the following description, the expression “some embodiments” refers to a subset of all possible embodiments. However, it should be understood that “some embodiments” can be the same subset or different subsets of all possible embodiments and can be combined with each other without conflict.
[0014] Breeding microarrays (such as the Illumina Infinium HD microarray) are high-throughput single nucleotide polymorphism (SNP) genotyping tools. They accelerate breeding processes by simultaneously detecting hundreds of thousands to millions of SNP sites, enabling rapid and accurate genotyping of plant or animal genomes. These microarrays are based on the principles of hybridization and fluorescence detection. Each SNP site consists of two allele-specific probes (e.g., corresponding to A and T, respectively) and one site-specific probe (as a control). These probes are immobilized at specific locations on the breeding microarray, forming a tiny array of dots. For allele-specific probes, each SNP site has two probes, each perfectly complementary to one of the two alleles. For example, for a given SNP site (A / T), there are probe A (labeled Cy3) and probe T (labeled Cy5). Site-specific probes are complementary to sequences near the SNP site and are used to verify successful hybridization but do not distinguish between alleles. Each probe (including allele-specific and site-specific probes) exists on the microarray as a repeating array. For example, each type of probe has 500 repeating dots, which are randomly distributed on the chip to eliminate positional deviations.
[0015] The working principle of breeding chips is mainly based on specific adsorption and genotyping technologies. For specific adsorption, each SNP site on the chip corresponds to a specific probe molecule. When DNA molecules from the sample flow through the chip, they specifically adsorb onto the probe molecules. This adsorption is based on the base pairing principle: A pairs with T, and C pairs with G. Therefore, only DNA fragments perfectly complementary to the probe are adsorbed onto the chip. For genotyping, through specific adsorption, the breeding chip can capture DNA fragments at the target SNP site in the sample. These DNA fragments are then labeled and detected using techniques such as fluorescent staining or chemiluminescence. By comparing the signal intensity or type at the same SNP site in different samples, their genotypes (e.g., AC, AA, or BC) can be determined. Breeding chips offer advantages such as high throughput, high sensitivity, and high specificity, enabling the simultaneous detection of thousands or even millions of SNP sites with minimal sample DNA. This makes breeding chips a promising area for applications in plant and animal breeding, disease diagnosis, and personalized medicine.
[0016] Breeding microarray technology mainly includes the following steps: discovery and design (before entering the laboratory), microarray fabrication (probe immobilization), experimental detection (hybridization), and signal interpretation (genotype determination). For discovery and design, thousands of SNP loci are first identified in the target species population using technologies such as DNA sequencing. Then, for each SNP locus to be detected, a DNA probe capable of specifically recognizing it is designed. For microarray fabrication (probe immobilization), these designed probes are precisely immobilized at specific locations on the microarray, forming a tiny array of dots. Therefore, the microarray contains a collection of various probes. For experimental detection (hybridization), genomic DNA is extracted from the plant or animal sample to be tested (such as corn leaves or pig blood). This DNA contains all the naturally occurring SNP loci. The sample DNA is processed and added to the microarray. If a fragment of the sample DNA (containing an SNP locus) is perfectly complementary to the probe sequence at a specific point on the microarray, they will bind together like a zipper (hybridization). For signal interpretation (genotype determination), fluorescence scanning identifies which probe has captured the sample DNA. Because each probe corresponds to a specific SNP site and allele, by detecting which probe emitted the signal, we can deduce the genotype of the corresponding SNP site in the sample genome.
[0017] In existing technologies, there are differences in probe synthesis efficiency during the probe synthesis stage. The in-situ synthesis efficiency of probes on the chip is not uniform but strongly depends on the sequence characteristics of the probe itself. This difference mainly stems from the following: 1. Sequence dependence of coupling chemistry: In each nucleotide addition cycle, the efficiency of chemical coupling is affected by the preceding and following sequence context. For example, the coupling rate of some dinucleotide combinations may be lower than that of other combinations, resulting in differences in synthetic yield.
[0018] 2. Decrease in deprotection efficiency: The photosensitive chemical groups used to protect nucleotides cannot achieve 100% deprotection efficiency. For longer or specific sequence probes, this efficiency loss has a cumulative effect, leading to an increase in the proportion of truncated probes and a decrease in the abundance of full-length probes.
[0019] 3. Interference from secondary sequence structures: During synthesis, the elongating DNA strand may form intramolecular secondary structures. Sequences rich in GC bases are particularly prone to forming stable hairpin structures, which can sterically hinder subsequent synthetic reactions, thereby reducing the final yield of the full-length probe.
[0020] The factors mentioned above collectively lead to inherent differences in the effective anchoring density of different probes on the chip. In other words, even with the same design concentration, the absolute number of full-length hybridizable probes on the chip will differ for different probe sites.
[0021] During the hybridization and detection stages, systematic biases exist in the signal. The non-uniformity of probe density generated during the synthesis stage is amplified in subsequent hybridization signal scanning stages, specifically manifested as: 1. Signal-to-noise ratio differences: Probes with higher synthesis efficiency have higher signal-to-noise ratios, resulting in more reliable data; 2. Noise bias: 1) Sequence-specific signal gain / loss: Probes with specific sequence characteristics (such as high GC content, repetitive sequences, and a tendency to form secondary structures) will exhibit consistently higher or lower signal intensities, deviating from their true biological signal strength; 2) Hybridization kinetic differences: Different probes have different melting temperatures, leading to varying hybridization strictness under standardized hybridization and elution conditions, resulting in systematic biases in signal intensity.
[0022] Furthermore, the boundary between homozygotes and heterozygotes is not clearly defined: due to non-specific hybridization and cross-hybridization, even if there are single or multiple base mismatches between the probe and the target DNA sequence, binding will still occur to some extent under certain hybridization conditions (temperature, salt concentration). This means that homozygotes may still produce a signal higher than the background level. Different probes have different melting temperatures due to variations in sequence length and GC content. This inherent difference in binding efficiency between probes can lead to a systematic distortion of signal intensity, causing a shift in the 1:1 ratio.
[0023] In summary, the data obtained from actual experiments may have the following problems: 1) some probes have low signal-to-noise ratios; 2) some probes have certain biases; 3) the boundary between heterozygous and homozygous is relatively blurred.
[0024] Existing identification technologies, such as those from Lasso, cluster each microbead by determining its category based on the polar coordinate position of each probe's scattered points (primarily the polar coordinate angle and the brightness value of the microbeads). This method, which ultimately classifies each probe, cannot correct for inherent probe biases and errors. Each microbead is treated as an independent entity, lacking a holistic probe perspective. This is because some microbeads inevitably have low signal-to-noise ratios (SNRs) (even outnumbering those with high SNRs) and abnormal signals. Low SNR microbeads are more likely to be misidentified. Identifying each microbead as an independent entity means that, on the one hand, the probe's final result may be influenced by the majority of low SNR microbeads, leading to errors; on the other hand, it is difficult to eliminate the influence of abnormal microbeads on the result.
[0025] Another existing method uses linear fitting, principal component analysis, and multiple linear regression, which essentially treats the scatter points of a probe as a straight line distributed in four-dimensional space for fitting analysis. While this method can analyze a probe holistically, allowing for probe identification based on the distribution trend of the scatter points and better avoiding interference from low signal-to-noise ratios and outliers, it still has certain drawbacks. It cannot correct inherent probe biases and errors; the extracted probe information is too simplistic, focusing only on the slope information of the probe scatter points. Probe scatter points are generally spindle-shaped, with variations in length, width, and other morphological characteristics among different probes. For example, as shown in the figure below... Figure 1 The scatter distribution of the two probes clearly shows differences in length, slope, thickness, and divergence. Simply using slope information alone is detrimental to further differentiation. Therefore, the purpose of this application is to learn the inherent preferences of each probe and to consider various morphological information of the probes during probe classification, thereby improving the accuracy of base identification at detection sites.
[0026] like Figure 2 As shown, Figure 2 This diagram illustrates the application environment of a base identification method based on a breeding chip in one embodiment. The application environment includes a scanning device 20, which acquires fluorescence images of the breeding chip, including multiple different base channels. The acquired fluorescence images are then analyzed to determine the base data of the detected heterozygous sample. The scanning device 20 can be a device with computing capabilities, such as a computer or server. The scanning device 20 is specifically designed for scanning breeding chips.
[0027] Please see Figure 3 This is a flowchart of a base recognition method based on a breeding chip according to an embodiment of this application. The base recognition method based on a breeding chip is applied in a scanning device and includes the following steps: S10. Acquire fluorescence images of the breeding chip, including multiple different base channels.
[0028] In this embodiment, the scanning device is equipped with lasers of different wavelengths. For example, the excitation wavelength corresponding to Cy3 is approximately 532 nm (green laser), and that corresponding to Cy5 is approximately 635 nm (red laser). The laser beam is focused into an extremely fine point and scanned line by line on the chip surface through a precise optical system. When the laser irradiates the probe for hybridization with fluorescently labeled DNA, the dye molecules are excited and emit fluorescence of a specific wavelength. The emitted fluorescence is captured by a highly sensitive detector, which converts the light signal into an electrical signal and amplifies it. Therefore, by exciting different fluorescent dyes with lasers of different wavelengths and collecting and digitizing the fluorescence signals channel by channel, a fluorescence image including multiple base channels is finally generated. For example, when photographing a breeding chip, a fluorescence image of four base channels is obtained, namely, the fluorescence image of the adenine (A) base channel, the fluorescence image of the cytosine (C) base channel, the fluorescence image of the guanine (G) base channel, and the fluorescence image of the thymine (T) base channel.
[0029] S11. Based on the fluorescence image, obtain the brightness value of each target probe scatter point in each base channel for each type of probe.
[0030] In this embodiment, at least one type of probe is configured on the breeding chip, with each type of probe comprising multiple probes at different locations. Each probe scatter point corresponding to a probe at a specific location on the breeding chip corresponds to a probe of that type. Each type of probe has its own probe sequence, which is used to hybridize with the target detection site. A probe scatter point represents the brightness value of the target detection site captured by the probe at a given location on each base channel; that is, one probe scatter point represents one fluorescent spot. The brightness values of the probe scatter point include the brightness values on the A, C, G, and T channels. Each probe at each location on the breeding chip corresponds to a detection site, such as a SNP site. The probe at each location specifically captures and detects the processed DNA target sequence containing the corresponding SNP site in the sample through the principle of complementary base pairing. Each type of probe is distributed at different locations on the breeding chip, which can be a random chip or an array chip; the chip type is not limited here. Therefore, each type of probe corresponds to multiple fluorescent spots, which correspond to probes at different locations on the breeding chip; that is, one fluorescent spot corresponds to one probe scatter point. Each type of probe can be used to detect homozygous or heterozygous samples. Homozygous samples are AA, CC, TT, GG, while heterozygous samples are AC, AT, etc.
[0031] The fluorescence image includes probe scatter points of various probes. Since the positions of various probes on the breeding chip can be determined after the chip is fabricated, the probe scatter points of each type of probe can be obtained based on their positions on the chip. After deleting abnormal scatter points, the target probe scatter points of each type of probe can be obtained, thus obtaining the brightness value of each target probe scatter point in each base channel of each type of probe.
[0032] S12. Based on the brightness values of each target probe scatter point in each base channel of each probe type, extract the scatter point data features of each probe type.
[0033] In this embodiment, the scatter data features represent the base brightness characteristics of the target detection sites captured by the probe. The scatter data features can be data formed by sampling probe scatter points according to preset rules, or features obtained by statistical analysis based on the target probe scatter points.
[0034] S13. Based on the scatter data features of each type of probe, form the input of the pre-trained base recognition model and output the base data of the target detection site detected by each type of probe.
[0035] In this embodiment, the base recognition model is trained on a training dataset. During training, the model learns the mapping relationship between the scatter plot features of each type of probe in the training dataset and the base type of the corresponding detection site. Different types of probes have inherent differences in base binding efficiency, channel brightness response, and noise preference. For example, a certain type of probe may have a strong C channel brightness response and high noise. The model learns the scatter plot distribution pattern of this type of probe and automatically calibrates the correspondence between the brightness signal and the base type, thereby filtering out probe-specific noise and improving the accuracy of base recognition. Using the scatter plot features of each type of probe as input to the pre-trained base recognition model, the output is the base data of the target detection site. The core advantage is that it deeply binds the specific features of various probes with the base recognition of the site, significantly improving the accuracy, stability, and applicability of base discrimination.
[0036] In this embodiment, the machine learning model includes, but is not limited to, any of the following: neural networks, deep learning, similar to the lightGBM principle (XGboost, CatBoost), support vector machines, K-nearest neighbors, etc.
[0037] Optionally, for any type of probe, a preset number of probe scatter points can be directly sampled as input to the machine learning model. For example, the number of scatter points measured by different types of probes may vary and is not fixed. Then, a certain sampling method (e.g., completely random, or dividing the scatter points into intervals based on brightness, and sampling a fixed number of points from each interval) can be used to extract a preset number of scatter points from these probe scatter points as input to the model. For example, a total of 10340×4 scatter point data can be sampled into 200×4 scatter point data as input to the model.
[0038] In this embodiment, for a type of probe, the scatter data features of the probe are used as the input of the base recognition model. The output of the base recognition model includes multiple categories, such as AA, CC, GG, TT, AC, AG, AT, CG, CT, and GT. For a type of probe, the base recognition model outputs the category prediction data of the probe, and the category with the highest confidence in the category prediction data is used as the base data of the target detection site detected by the probe.
[0039] In the above embodiments, fluorescence images are acquired, and based on the fluorescence images, the brightness values of each target probe scatter point in each base channel of each type of probe are acquired. Based on the brightness values of each target probe scatter point in each base channel of each type of probe, the scatter data features of each type of probe are extracted. The scatter data features represent the base brightness features of the target detection sites captured by the probe of that type. Based on the scatter data features of each type of probe, the base data of the target detection sites for the input and output of the pre-trained base recognition model are formed. The core advantage is that the specific features of the probe are deeply bound to the base recognition of the site, which greatly improves the accuracy, stability and applicability of base discrimination. The accuracy is 4-10 percentage points higher than that of using slope and brightness to classify.
[0040] In some embodiments, obtaining the brightness value of each target probe scatter point in each base channel based on the fluorescence image includes: Based on the fluorescence images, preprocessed brightness data for each type of probe is obtained; Based on the preprocessed brightness data of each type of probe, abnormal brightness data in each type of probe is removed to obtain the target brightness data of each type of probe. Based on the target brightness data of each type of probe, the brightness value of each target probe scatter point in each base channel is obtained.
[0041] In this embodiment, the fluorescence image is preprocessed to obtain preprocessed brightness data for each type of probe. The preprocessing includes, but is not limited to, at least one of the following: channel crosstalk correction for each type of probe, dimensional difference processing between channels of each type of probe, etc. The preprocessed brightness data for one type of probe includes the brightness values of the preprocessed probe scatter points on each base channel. For example... Figure 4 As shown, Figure 4 This is a scatter plot of probe points before and after preprocessing in one embodiment.
[0042] Abnormal brightness data in one type of probe represents outlier probe points within that type of probe; these are individual data points whose distribution trend or range differs significantly from the overall probe scatter. They can be understood as points that deviate from the majority of the probe scatter.
[0043] Optionally, the step of removing abnormal brightness data from each type of probe based on the preprocessed brightness data of each probe type to obtain the brightness value of each target probe scatter point on each base channel in each type of probe includes at least one of the following: Based on the preprocessed brightness data of any type of probe, the brightness distribution of the probe is estimated, the distance from each probe point in the probe to the center of the brightness distribution is calculated, and the brightness data of probe points with a distance greater than or equal to a preset distance threshold are regarded as abnormal brightness data and deleted, so as to obtain the brightness value of each target probe point in each base channel in the probe. Based on the preprocessed brightness data of any type of probe, multiple isolation trees are generated. The average path length from each probe scatter point in any type of probe to all isolation trees is calculated. Based on the average path length corresponding to each probe scatter point, the anomaly score corresponding to each probe scatter point is calculated. The brightness data of probe scatter points with anomaly scores greater than the anomaly threshold are regarded as abnormal brightness data and deleted. The brightness values of each target probe scatter point in any type of probe on each base channel are obtained.
[0044] In this embodiment, a four-dimensional base channel space is established based on the brightness of probe scatter points on each base channel, with each base representing a spatial dimension. Within this four-dimensional base channel space, for probe scatter data of a class of probes, a multidimensional Gaussian distribution is used to estimate the brightness distribution of the probe class. For example, the EllipticEnvelope library from Python is used to remove outliers. Its principle is that the data is a multivariate Gaussian distribution, and an ellipse is fitted to the data so that the ellipse covers the central region (i.e., the high-density region) of the data, while points outside the ellipse are considered outliers.
[0045] In this embodiment, an isolation forest can also be used to remove outliers from the four-dimensional scatter plot data. An isolation forest is a tree-based anomaly detection method that leverages the ease with which outliers can be isolated to identify anomalies. Outliers are relatively few in number and differ significantly from normal points, making them easy to randomly partition (isolate). By constructing multiple random trees, the path length (or average path length) required to isolate each point is calculated; the shorter the path, the more likely it is to be an outlier. The mapping relationship between path length and anomaly score can be configured, converting the average path length into anomaly score, thereby identifying and deleting the brightness data of probe scatter plots with anomaly scores greater than an anomaly threshold as anomalous brightness data. Figure 5 As shown, Figure 5 This is a comparative diagram of the process before and after removing abnormal data points in one embodiment; after removing abnormal brightness data, the noise is reduced.
[0046] In the above embodiments, for the probe scatter data of each type of probe, abnormal probe scatter points are removed to improve the overall quality of the data and make subsequent analysis more reliable.
[0047] In some embodiments, the scatter data features include at least one of the following: prior base sequence features, principal component features of base channel brightness, and Gaussian distribution features of base brightness. The extraction of scatter data features for each type of probe based on the brightness values of each target probe scatter point in each base channel includes at least one of the following: For any type of probe, principal component analysis is performed based on the brightness values of each target probe scatter point on each base channel to obtain the principal component features of the base channel brightness. For any type of probe, Gaussian distribution fitting is performed based on the brightness values of each target probe scatter point on each base channel to obtain the Gaussian distribution characteristics of the base brightness.
[0048] In this embodiment, for a type of probe, the prior base sequence features represent the sequence attributes of the probe sequence, and their core function is to provide a reference for the base recognition model to correct system bias interference. Optionally, the prior base sequence features include at least one of the following: the ratio of C to G bases in the probe sequence of any type of probe, and the number of bases at the ends of the probe sequence of any type of probe. For a type of probe, the ratio of C to G bases refers to the proportion of G and C bases in the entire probe sequence. This ratio can cause sequence bias interference by affecting the binding stability of the probe to the detection site and the signal response intensity of the base channel. Using it as model input allows the model to learn and calibrate this signal deviation caused by the difference in GC content, improving the consistency of base recognition for probes with different GC ratios. For a type of probe, the number of bases at the ends refers to a specific number of base combinations at the ends of the probe sequence, reflecting the sequence features near the binding region of the probe and the detection site. Due to the binding bias of enzymes during sequencing, these bases can cause adjacent interference to the signal detection results of the detection site. Encoding these bases digitally (e.g., A:0, C:1, G:2, T:3) and inputting them into the model helps the model identify and offset detection bias caused by adjacent sequences, optimizing the accuracy of target site base discrimination. Based on sequencing experience, due to enzyme bias, the first few bases (bp) of the detection site can interfere with the response. Therefore, using the first few bp as input helps correct this bias. For example, A:0, C:1, G:2, T:3, and ATGC would be the four numbers 0, 3, 2, and 1.
[0049] In this embodiment, for a certain type of probe, the base channel brightness principal component feature represents multiple features of that type of probe obtained based on principal component analysis.
[0050] Optionally, the principal component features of the base channel brightness include at least one of the following: a first principal component direction vector, the explained variance and the proportion of explained variance of the first principal component direction vector, a second principal component direction vector, the explained variance and the proportion of explained variance of the second principal component direction vector, the brightness mean and brightness variance of each base channel, wherein the first principal component direction vector represents the central axis direction of the four-dimensional distribution shape formed by the probe scatter points of any type of probe in the four-dimensional base channel space, and the second principal component direction vector represents the secondary central axis direction of the four-dimensional distribution shape formed by the probe scatter points of any type of probe in the four-dimensional base channel space.
[0051] In this embodiment, within the four-dimensional base channel space of four-base (A, C, G, T) brightness, for the probe scatter data of a certain type of probe, the distribution of these target probe scatter points exhibits a four-dimensional distribution shape, such as a spindle-shaped distribution. The spindle-shaped distribution is essentially an ellipsoidal cluster of sample points in the four-dimensional base channel space, and the major axis direction of the spindle-shaped distribution is the first principal component direction vector. The first principal component direction vector is the direction in the four-dimensional base channel space that maximizes the variance after projection of all target probe scatter points of this type of probe. The first principal component direction vector captures the most significant variation information among all target probe scatter points. The first principal component direction vector includes the proportional relationship of the brightness of each base channel, that is, the relative magnitude of the coefficients of the four base channels (A, C, G, T), which represents the brightness contribution ratio of this type of probe in the four base channels. For example, if the first principal component direction vector is (0.1, 0.85, 0.03, 0.02), it means that the C channel's brightness contribution is dominant, while the A / G / T channels' contributions are minimal. This directly corresponds to the ACGT brightness ratio characteristic of the probe. The first principal component direction vector describes the overall trend of all target probe scatter points in this type of probe.
[0052] For scatter plot data of a class of probes, the second principal component direction vector, after removing the first principal component direction vector and capturing the core trend, represents the direction of the second largest source of variation in the four-dimensional base channel space. It indicates secondary trend directions besides the scatter plot distribution trend represented by the first principal component direction vector. These secondary trend directions can be noise trends, signal preference trends, etc. These secondary trends precisely reflect the non-core characteristics of the probes, such as the noise preference of different probes (e.g., random noise concentrated in a certain channel in that direction); or the difference in signal-to-noise ratio (SNR). For example, the larger the variance of the second principal component direction vector, the greater the fluctuation of the data around the main trend, and the lower the SNR may be. As a supplement to the first principal component direction vector, the second principal component direction vector can fill in the details not covered by the first principal component direction vector, making the base recognition model's classification of probes more accurate. For example, if the first principal component direction vectors of two types of probes are similar, but their second principal component direction vectors are significantly different, they can be distinguished by the second principal component direction vector.
[0053] For probe scatter data of a certain type of probe, the explained variance of the first principal component direction vector represents the degree of dispersion of probe scatter points along the first principal component direction vector, which is related to the brightness of the probe scatter points in a sense. Generally, the larger the explained variance, the greater the brightness distribution of the probe scatter points, such as a longer spindle shape; for example, an explained variance of 0.5. The explained variance of the second principal component direction vector represents the degree of dispersion of probe scatter points along the second principal component direction vector, which reflects the thickness of the spindle shape in a sense. The larger the explained variance, the more noise the probe has and the lower the signal-to-noise ratio; for example, an explained variance of 0.3 for the second principal component direction vector. The explained variance proportion of the first principal component direction vector represents the proportion of probe scatter point information that the first principal component direction vector can express; a larger value is better (maximum value is 1), and a larger value indicates a thinner spindle shape of the scatter points. The explained variance proportion of the second principal component direction vector represents the proportion of probe scatter point information that the second principal component direction vector can express; a smaller value is better (maximum value is 1), and a larger value indicates a thicker spindle shape formed by the probe scatter points.
[0054] For scatter plot data of a class of probes, the brightness mean in a single base channel is calculated by summing and averaging the brightness of all target probes in that channel. Then, using the variance formula, the brightness variance in that channel can be calculated. Similar calculations are performed for other base channels.
[0055] Optionally, for any type of probe, based on the brightness values of each target probe scatter point on each base channel, principal component analysis is performed to obtain the principal component features of the base channel brightness, including: The brightness values of each target probe scatter point in each base channel of any type of probe are used to form a brightness data matrix, wherein the first dimension of the brightness data matrix represents the probe scatter point identifier and the second dimension represents the brightness of each base channel. Based on the brightness data matrix, a brightness covariance matrix is calculated, wherein the brightness covariance matrix represents the linear correlation between the brightness of the four base channels; Based on the brightness covariance matrix, eigenvalue decomposition is performed to obtain multiple eigenvalues. The eigenvector corresponding to the largest eigenvalue among the multiple eigenvalues is taken as the first principal component direction vector, and the eigenvector corresponding to the second largest eigenvalue among the multiple eigenvalues is taken as the second principal component direction vector. Based on the first principal component direction vector, calculate the explained variance and the proportion of explained variance of the first principal component direction vector; based on the second principal component direction vector, calculate the explained variance and the proportion of explained variance of the second principal component direction vector.
[0056] In this embodiment, for a type of probe, the luminance covariance matrix can be obtained based on the luminance data matrix and its transpose. The four components of the first principal component direction vector are the weighting coefficients of the four base channels A, C, G, and T, and their relative proportions directly reflect the base luminance characteristics of this type of probe. The explained variance of the first principal component direction vector is the largest eigenvalue. The second largest eigenvalue is the largest eigenvalue after removing the largest eigenvalue. The explained variance of the second principal component direction vector is the second largest eigenvalue. The proportion of the explained variance of the first principal component direction vector is equal to the ratio of the largest eigenvalue to all eigenvalues. The proportion of the explained variance of the second principal component direction vector is equal to the ratio of the second largest eigenvalue to all eigenvalues.
[0057] Using the first principal component correlation feature and the second principal component correlation feature as input to the base identification model has the core advantage of dimensionality reduction and purification of the four-dimensional base brightness data, preserving the core signal and key interference features while reducing model complexity. The first principal component correlation feature is the direction of maximum variance in the four-dimensional base channel space, directly corresponding to the central axis of the four-dimensional distribution shape of the probe scatter points. Its physical meaning is the core proportional feature of the bases at the detection site, such as A:C:G:T=0.1:0.85:0.03:0.02. Using the first principal component correlation feature as input is equivalent to directly representing the core trend of the bases at the detection site in one dimension, replacing the original four-dimensional brightness data. This retains the key information for determining the base type while avoiding redundant correlations between channels in the original data, allowing the model to quickly focus on the core features that determine the base type.
[0058] Secondly, by utilizing the second principal component correlation features, minor fluctuations can be captured, incorporating key features related to noise and bias. Second principal component correlation features focus on minor trends beyond the overall trend, such as the signal-to-noise ratio differences between different probe types and channel noise bias. Jointly inputting the first and second principal component correlation features is equivalent to supplementing the base identification model with an interference feature dimension. The base identification model can learn that the first principal component correlation features represent the core base signal, and it can also learn the correspondence between noise / bias interference in the signal represented by the second principal component correlation features. Furthermore, the second principal component correlation features can be used to calibrate the judgment bias of the first principal component correlation features, such as distinguishing sites with similar first principal component correlation features but different second principal component correlation features, thus improving the accuracy of base typing.
[0059] In the above embodiments, statistical analysis of the target probe scatter data for each type of probe yields first principal component correlation features and second principal component correlation features. The first principal component correlation features directly characterize the core trend of the bases at the detection site in one dimension, replacing the original four-dimensional brightness data. This retains the key information for determining the base type while avoiding redundant correlations between channels in the original data, allowing the model to quickly focus on the core features that determine the base type. The second principal component correlation features can focus on smaller trends beyond the overall trend, such as the signal-to-noise ratio differences between different probes and channel noise bias. Furthermore, the second principal component correlation features can be used to calibrate the judgment bias of the first principal component correlation features, thereby improving the accuracy of base type identification.
[0060] In some embodiments, the Gaussian distribution feature of base brightness includes at least one of the following: a brightness mean vector based on a Gaussian distribution, and a covariance matrix based on a Gaussian distribution.
[0061] Optionally, the step of performing Gaussian distribution fitting based on the brightness values of each target probe scatter point in each base channel to obtain the Gaussian distribution features of the base brightness includes: For any type of probe, a multidimensional Gaussian distribution is used to fit the brightness values of each target probe scatter point in each base channel to obtain the center position of the multidimensional Gaussian distribution and the covariance matrix of the multidimensional Gaussian distribution. The center position of the multidimensional Gaussian distribution is used as the brightness mean vector based on the Gaussian distribution.
[0062] In this embodiment, for a certain type of probe, the target probe scatter points of this type of probe are fitted using the following multidimensional Gaussian distribution formula, as follows: in This represents a vector, specifically the brightness value across the four base channels A, C, G, and T. This represents the density value corresponding to each base channel. Represents the mean brightness vector. Let represent the covariance matrix.
[0063] For a given type of probe, the mean brightness vector based on a Gaussian distribution represents the central distribution location of the probe scatter points, for example, a mean brightness vector of 0.5, 0.2, 0.1, and 0.8. The covariance matrix of the multidimensional Gaussian distribution represents the four-dimensional distribution shape of the probe scatter points and the correlation between the brightness of each channel, such as the stretching direction and the flattening direction (e.g., 4×4, a total of 16 values).
[0064] In this embodiment, the mean vector and covariance matrix of the multidimensional Gaussian distribution fitted in the four-dimensional base channel space are used as model inputs. For a certain type of probe, the core advantage is that it transforms the overall distribution characteristics of the probe scatter points of this type of probe into structured and interpretable quantitative indicators, providing a more accurate global reference for the base identification model than the original scatter data.
[0065] The Gaussian-distributed mean brightness vector provides the central benchmark feature of the base signal at the detection site. The four components of the Gaussian-distributed mean brightness vector correspond to the mean brightness values of the A, C, G, and T channels, directly representing the central position of the four-dimensional distribution shape of the probe scatter points and reflecting the average signal intensity ratio of the bases at the detection site. Compared to the chaotic fluctuations of the original brightness data, the Gaussian-distributed mean brightness vector is the core benchmark after noise reduction. The base identification model can quickly determine the base signal bias of the detection site using this benchmark; for example, if the mean value of the C channel is significantly higher than other channels, it directly points to the dominance of the C base. This avoids the interference of random noise in the original data on signal judgment, allowing the model to focus on the essential signal characteristics of the detection site and improving the stability of base typing. The covariance matrix of a multidimensional Gaussian distribution provides the correlation and discrete characteristics of base channel signals. The diagonal elements of the covariance matrix represent the variance of each base channel, reflecting the dispersion of the brightness of a single base channel; a larger variance indicates stronger signal fluctuations / noise in that channel. The off-diagonal elements represent the covariance between base channels, reflecting the degree of linear correlation between the brightness of two base channels. For example, if the covariance between channels A and C is positive, it means that if the brightness of channel A is high, the brightness of channel C is also likely to be high. By inputting the multidimensional Gaussian distribution covariance matrix into the model, the base recognition model can determine the reliability of the channel signal based on the variance magnitude and the interference relationship between channels based on the covariance, thereby calibrating base recognition biases caused by channel correlation or noise.
[0066] In the above embodiments, the mean brightness vector based on Gaussian distribution is used as the input of the base identification model. The base identification model can quickly determine the base signal bias of the detection site, avoiding the interference of random noise in the original data on signal judgment, allowing the model to focus on the essential signal features of the detection site and improving the stability of base typing. The covariance matrix of multidimensional Gaussian distribution is input into the model. The base identification model can judge the reliability of the channel signal based on the variance and judge the interference relationship between channels based on the covariance, thereby calibrating the base identification deviation caused by channel correlation or noise.
[0067] In some embodiments, the method includes: Obtain a training dataset, wherein each training sample in the training dataset includes scatter data sample features of a class of probes and base labels of the detection sites corresponding to a class of probes. The base recognition model is trained based on the training dataset.
[0068] In this embodiment, during the training process, the features of the scattered data samples are the scattered data features. In each iteration of training, the base recognition model outputs the base data of the detection sample site corresponding to the predicted training sample. The predicted base data is compared with the base label of the detection sample site, and the loss value of the current iteration is calculated. If the loss value of the current iteration is greater than the preset loss value, backpropagation is performed, and iterative training continues until the iteration termination condition is met, such as the loss value of the current iteration being less than or equal to the preset loss value.
[0069] Optionally, the method further includes: Obtain the base tags of the detection sample sites; The base tags for obtaining the detection sample sites include: Selecting qualified sample loci from the sequencing data in whole-genome sequencing as detection sample loci, wherein the sequencing data includes base identification data of each sample locus and base identification quality value of each sample locus, wherein the conditions include at least one of the following: the base identification quality value is greater than a preset quality value, the sequencing depth is greater than a preset sequencing depth, and the ratio of the sum of the most and second most bases is greater than a first preset ratio value. When the maximum base percentage of the detected sample site exceeds the second preset percentage value, the detected sample site is determined to be a homozygous sample, and the label of the detected sample site is the homozygous sample label of the base type corresponding to the maximum base percentage; When the maximum base percentage of the detection sample site is less than or equal to the second preset percentage value, the detection sample site is determined to be a heterozygous sample, and the label of the detection sample site is a heterozygous sample label, which is formed by the base type corresponding to the maximum base percentage and the base type corresponding to the second largest base percentage.
[0070] In this embodiment, the sequencing data includes base identification data for multiple sample sites and the base identification quality value for each sample site. However, some sample sites may not meet the requirements for tagging and need to be deleted. For example, sample sites with a base identification quality value below 25, sequencing depth less than 30, and sample sites where the sum of the most and second most identified bases is less than 90% are removed. The maximum base percentage represents the ratio of the most frequent base to the sequencing depth of the sample site. The most frequent base represents the number of bases that appear most often in the sequencing depth of the sample site. The second most frequent base represents the number of bases that appear less frequently than the most frequent base in the sequencing depth of the sample site. The second most frequent base percentage represents the ratio of the second most frequent base to the sequencing depth of the sample site. For example, if the sequencing depth of a certain sample locus is 50 (meaning a total of 50 sequencing results were obtained for that locus), and 25 of these results are A, 20 are C, 3 are G, and 2 are T, then the base tag for that sample locus is AC, where 25 is the maximum number of bases and 20 is the second maximum number of bases. Alternatively, if 45 of the 50 results are C, 3 are A, and 2 are G, then the base tag for that sample locus is CC.
[0071] Taking the lightGBM model as an example, the training and usage process is explained. There are a total of 10,000 probe sites, of which 8,200 detection sample sites have labels. Data was collected from ten chips (each chip has 24 blocks, and each block can be considered a data unit), resulting in 24 × 10 = 240 data units. Therefore, there are 8,200 × 240 = 196,800 training samples. The training samples are divided into training and testing sets in an 8:2 ratio, meaning the training set has 157,440 training samples and the testing set has 39,360 testing samples. Therefore, the model input is N × 45, and the output is 1, indicating the specific category of the detection sample site. The specific parameters of the lightGBM model are: number of leaf nodes num_leaves = 16, maximum tree depth max_depth = 8, learning rate learning_rate = 0.01. It is also worth noting that the first principal component direction vector can be used. The variance of the first principal component direction vector is used as the confidence level assessment for the detected sample. Alternatively, the maximum value of the predicted class probability using the LightGBM model (neural network) can be used as the confidence level.
[0072] Each training sample is strictly bound to specific scatter plot features of a probe type (such as principal component features, Gaussian distribution features, and prior sequence information) and the base label of the detection site of that probe type, for example, homozygous CC and heterozygous AC. This binding makes the model's learning objective highly specific, no longer a generalized base signal recognition, but a precise mapping of features to bases for a specific probe type; it avoids cross-interference of features from different probe types (for example, the scatter plot features of type A probes will not be misjudged by the model as base signals of type B probes), and significantly improves the model's specificity for recognizing different probes. The scatter plot features in the training samples, such as principal component features, Gaussian distribution features, and prior sequence information, are all systematic features strongly correlated with probe types, while the base label is the actual biological result corresponding to these features.
[0073] During training, the model automatically learns hidden patterns such as a high GC ratio in probes leading to enhanced C-channel signals and probes with G at the terminal causing adjacent interference to the target site. It then proactively corrects these signal biases caused by probe properties during the prediction phase, resolving systematic errors that traditional thresholding methods cannot eliminate. Training samples are constructed on a per-probe-class basis, and the model learns the association patterns between the common features of that probe class and the base tags. When new probe data of the same type is added, there is no need to retrain the model; simply inputting the scatter plot features of that probe allows the model to directly call upon the learned patterns to output base results, enabling rapid reuse across batches and experiments. For new probe types, only a small number of samples of the same probe's features—base tags—need to be added for fine-tuning, allowing for rapid adaptation and significantly reducing the model's iteration costs.
[0074] In the above embodiments, each training sample is strictly bound to the specific scatter features of a type of probe (such as principal component features, Gaussian distribution features, prior sequence information) and the base label of the detection sample site measured by that type of probe. This binding allows the model to be trained with the base label as the training target, which greatly improves the model's recognition specificity for different probes.
[0075] In another aspect, this application provides a computer program product, including a computer program that, when executed by a processor, implements the base recognition method based on a breeding chip as described in any embodiment of this application.
[0076] In the computer program product, the optional implementation form of the program module architecture of the computer program that implements each step of the base recognition method based on the breeding chip can be a base recognition device based on the breeding chip.
[0077] Please see Figure 6One embodiment of this application provides a base recognition device based on a breeding chip, comprising: an acquisition module 61, configured to acquire fluorescence images of the breeding chip including multiple different base channels; an extraction module 62, configured to acquire, based on the fluorescence images, the brightness values of each target probe scatter point in each base channel of each type of probe; the extraction module 62 is further configured to extract the scatter point data features of each type of probe based on the brightness values of each target probe scatter point in each base channel of each type of probe; and a recognition module 63, configured to form the input of a pre-trained base recognition model based on the scatter point data features of each type of probe, and output the base data of the target detection sites detected by each type of probe.
[0078] Optionally, the extraction module 62 is also used for: Based on the fluorescence images, preprocessed brightness data for each type of probe is obtained; Based on the preprocessed brightness data of each type of probe, abnormal brightness data in each type of probe are removed to obtain the brightness value of each target probe scatter point in each base channel of each type of probe.
[0079] Optionally, the extraction module 62 is also used for: Based on the preprocessed brightness data of any type of probe, the brightness distribution of the probe is estimated, the distance from each probe point in the probe to the center of the brightness distribution is calculated, and the brightness data of probe points with a distance greater than or equal to a preset distance threshold are regarded as abnormal brightness data and deleted, so as to obtain the brightness value of each target probe point in each base channel in the probe. Based on the preprocessed brightness data of any type of probe, multiple isolation trees are generated. The average path length from each probe scatter point in any type of probe to all isolation trees is calculated. Based on the average path length corresponding to each probe scatter point, the anomaly score corresponding to each probe scatter point is calculated. The brightness data of probe scatter points with anomaly scores greater than the anomaly threshold are regarded as abnormal brightness data and deleted. The brightness values of each target probe scatter point in any type of probe on each base channel are obtained.
[0080] Optionally, the scatter data features include at least one of the following: prior base sequence features, principal component features of base channel brightness, and Gaussian distribution features of base brightness. The extraction module 62 is further used for: For any type of probe, principal component analysis is performed based on the brightness values of each target probe scatter point on each base channel to obtain the principal component features of the base channel brightness. For any type of probe, Gaussian distribution fitting is performed based on the brightness values of each target probe scatter point on each base channel to obtain the Gaussian distribution characteristics of the base brightness.
[0081] Optionally, the prior base sequence features include at least one of the following: the ratio of base C to base G in the probe sequence of any type of probe, and the number of bases at the ends of the probe sequence of any type of probe. The principal component features of the base channel brightness include at least one of the following: a first principal component direction vector, the explained variance and the proportion of explained variance of the first principal component direction vector, a second principal component direction vector, the explained variance and the proportion of explained variance of the second principal component direction vector, the mean brightness and the brightness variance of each base channel, wherein the first principal component direction vector represents the central axis direction of the four-dimensional distribution shape formed by the probe scatter points of any type of probe in the four-dimensional base channel space, and the second principal component direction vector represents the secondary central axis direction of the four-dimensional distribution shape formed by the probe scatter points of any type of probe in the four-dimensional base channel space; The Gaussian distribution feature of base brightness includes at least one of the following: a brightness mean vector based on Gaussian distribution, and a covariance matrix based on Gaussian distribution.
[0082] Optionally, the extraction module 62 is also used for: The brightness values of each target probe scatter point in each base channel of any type of probe are used to form a brightness data matrix, wherein the first dimension of the brightness data matrix represents the probe scatter point identifier and the second dimension represents the brightness of each base channel. Based on the brightness data matrix, a brightness covariance matrix is calculated, wherein the brightness covariance matrix represents the linear correlation between the brightness of the four base channels; Based on the brightness covariance matrix, eigenvalue decomposition is performed to obtain multiple eigenvalues. The eigenvector corresponding to the largest eigenvalue among the multiple eigenvalues is taken as the first principal component direction vector, and the eigenvector corresponding to the second largest eigenvalue among the multiple eigenvalues is taken as the second principal component direction vector. Based on the first principal component direction vector, calculate the explained variance and the proportion of explained variance of the first principal component direction vector; based on the second principal component direction vector, calculate the explained variance and the proportion of explained variance of the second principal component direction vector.
[0083] Optionally, the extraction module 62 is also used for: For any type of probe, a multidimensional Gaussian distribution is used to fit the brightness values of each target probe scatter point in each base channel to obtain the center position of the multidimensional Gaussian distribution and the covariance matrix of the multidimensional Gaussian distribution. The center position of the multidimensional Gaussian distribution is used as the brightness mean vector based on the Gaussian distribution.
[0084] Optionally, a training module 64 is also included for: Obtain a training dataset, wherein each training sample in the training dataset includes scatter data sample features of a class of probes and base labels of detection sample sites corresponding to a class of probes. The base recognition model is trained based on the training dataset.
[0085] Optionally, training module 64 is also used for: Obtain the base tags of the detection sample sites; The base tags for obtaining the detection sample sites include: Selecting qualified sample loci from the sequencing data in whole-genome sequencing as detection sample loci, wherein the sequencing data includes base identification data of each sample locus and base identification quality value of each sample locus, wherein the conditions include at least one of the following: the base identification quality value is greater than a preset quality value, the sequencing depth is greater than a preset sequencing depth, and the ratio of the sum of the most and second most bases is greater than a first preset ratio value. When the maximum base percentage of the detected sample site exceeds the second preset percentage value, the detected sample site is determined to be a homozygous sample, and the label of the detected sample site is the homozygous sample label of the base type corresponding to the maximum base percentage; When the maximum base percentage of the detection sample site is less than or equal to the second preset percentage value, the detection sample site is determined to be a heterozygous sample, and the label of the detection sample site is a heterozygous sample label, which is formed by the base type corresponding to the maximum base percentage and the base type corresponding to the second largest base percentage.
[0086] Those skilled in the art will understand that the structure of the base recognition device based on the breeding chip does not constitute a limitation on the device itself. Each module can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in hardware or independently of the controller in the scanning device, or stored in software in the memory of the scanning device, so that the controller can invoke and execute the operations corresponding to each module. In other embodiments, the base recognition device based on the breeding chip may include more or fewer modules than those shown in the figures.
[0087] Please see Figure 7In another aspect of this application, a scanning device 20 is also provided, including a memory 3011 and a processor 3012. The memory 3011 stores a computer program, which, when executed by the processor, causes the processor 3012 to perform the steps of the base recognition method based on breeding chips provided in any of the above embodiments of this application. The scanning device may include a desktop computer, laptop computer, tablet computer, handheld computer, smart speaker, server, etc., mobile phone (e.g., smartphone, cordless phone, etc.), wearable device (e.g., a pair of smart glasses or a smartwatch), gene sequencer, or similar device.
[0088] The processor 3012 is the control center, connecting various parts of the scanning device via various interfaces and lines. It executes software programs and / or modules stored in the memory 3011, and calls data stored in the memory 3011 to perform various functions and process data. Optionally, the processor 3012 may include one or more processing cores; the processor 3012 includes, but is not limited to, one or more combinations of the following: Central Processing Unit (CPU), Graphics Processing Unit (GPU), Field-Programmable Gate Array (FPGA), etc. Preferably, the processor 3012 may integrate an application processor and a modem processor, wherein the application processor mainly handles the operating system, user page, and applications, and the modem processor mainly handles wireless communication. It is understood that the aforementioned modem processor may not be integrated into the processor 3012.
[0089] The memory 3011 can be used to store software programs and modules. The processor 3012 executes various functional applications and data processing by running the software programs and modules stored in the memory 3011. The memory 3011 may mainly include a program storage area and a data storage area. The program storage area may store the operating system, application programs required for at least one function (such as sound playback function, image playback function, etc.), etc.; the data storage area may store data created according to the use of the scanning device, etc. In addition, the memory 3011 may include high-speed random access memory, and may also include non-volatile memory, such as at least one disk storage device, flash memory device, or other volatile solid-state storage device. Accordingly, the memory 3011 may also include a memory controller to provide the processor 3012 with access to the memory 3011.
[0090] In another aspect, this application also provides a storage medium storing a computer program, which, when executed by a processor, causes the processor to perform the steps of the base recognition method based on breeding chips provided in any of the above embodiments of this application.
[0091] Those skilled in the art will understand that all or part of the processes in the methods provided in the above embodiments can be implemented by a computer program instructing related hardware. The program can be stored in a non-volatile computer-readable storage medium, and when executed, it can include the processes of the embodiments of the above methods. Any references to memory, storage, databases, or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), dual data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.
[0092] The above description is merely a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the technical scope disclosed in the present invention should be included within the scope of protection of the present invention. The scope of protection of the present invention should be determined by the scope of the claims.
Claims
1. A base recognition method based on a breeding chip, characterized in that, include: Acquire fluorescence images of the breeding chip, including multiple different base channels; Based on the fluorescence image, the brightness value of each target probe scatter point in each base channel of each type of probe is obtained; Based on the brightness values of each target probe scatter point in each base channel in each type of probe, the scatter point data features of each type of probe are extracted. Based on the scatter data features of each type of probe, the input to the pre-trained base recognition model is formed, and the output is the base data of the target detection site detected by each type of probe.
2. The base recognition method based on breeding chips as described in claim 1, characterized in that, The step of obtaining the brightness value of each target probe scatter point in each base channel for each type of probe based on the fluorescence image includes: Based on the fluorescence images, preprocessed brightness data for each type of probe is obtained; Based on the preprocessed brightness data of each type of probe, abnormal brightness data in each type of probe are removed to obtain the brightness value of each target probe scatter point in each base channel of each type of probe.
3. The base recognition method based on breeding chips as described in claim 2, characterized in that, The process of preprocessing brightness data for each probe type, removing abnormal brightness data from each probe type, and obtaining the brightness values of each target probe scatter point on each base channel for each probe type includes at least one of the following: Based on the preprocessed brightness data of any type of probe, the brightness distribution of the probe is estimated, the distance from each probe point in the probe to the center of the brightness distribution is calculated, and the brightness data of probe points with a distance greater than or equal to a preset distance threshold are regarded as abnormal brightness data and deleted, so as to obtain the brightness value of each target probe point in each base channel in the probe. Based on the preprocessed brightness data of any type of probe, multiple isolation trees are generated. The average path length from each probe scatter point in any type of probe to all isolation trees is calculated. Based on the average path length corresponding to each probe scatter point, the anomaly score corresponding to each probe scatter point is calculated. The brightness data of probe scatter points with anomaly scores greater than the anomaly threshold are regarded as abnormal brightness data and deleted. The brightness values of each target probe scatter point in any type of probe on each base channel are obtained.
4. The base recognition method based on breeding chips as described in claim 1, characterized in that, The scatter data features include at least one of the following: prior base sequence features, principal component features of base channel brightness, and Gaussian distribution features of base brightness. The extraction of scatter data features for each type of probe based on the brightness values of each target probe scatter point in each base channel includes at least one of the following: For any type of probe, principal component analysis is performed based on the brightness values of each target probe scatter point on each base channel to obtain the principal component features of the base channel brightness. For any type of probe, Gaussian distribution fitting is performed based on the brightness values of each target probe scatter point on each base channel to obtain the Gaussian distribution characteristics of the base brightness.
5. The base recognition method based on breeding chips as described in claim 4, characterized in that, The prior base sequence features include at least one of the following: the ratio of base C to base G in the probe sequence of any type of probe, and the number of bases at the end of the probe sequence of any type of probe at a predetermined number of positions. The principal component features of the base channel brightness include at least one of the following: a first principal component direction vector, the explained variance and the proportion of explained variance of the first principal component direction vector, a second principal component direction vector, the explained variance and the proportion of explained variance of the second principal component direction vector, the mean brightness and the brightness variance of each base channel, wherein the first principal component direction vector represents the central axis direction of the four-dimensional distribution shape formed by the probe scatter points of any type of probe in the four-dimensional base channel space, and the second principal component direction vector represents the secondary central axis direction of the four-dimensional distribution shape formed by the probe scatter points of any type of probe in the four-dimensional base channel space; The Gaussian distribution feature of base brightness includes at least one of the following: a brightness mean vector based on Gaussian distribution, and a covariance matrix based on Gaussian distribution.
6. The base recognition method based on breeding chips as described in claim 5, characterized in that, For any type of probe, based on the brightness values of each target probe scatter point in each base channel, principal component analysis is performed to obtain the principal component features of the base channel brightness, including: The brightness values of each target probe scatter point in each base channel of any type of probe are used to form a brightness data matrix, wherein the first dimension of the brightness data matrix represents the probe scatter point identifier and the second dimension represents the brightness of each base channel. Based on the brightness data matrix, a brightness covariance matrix is calculated, wherein the brightness covariance matrix represents the linear correlation between the brightness of the four base channels; Based on the brightness covariance matrix, eigenvalue decomposition is performed to obtain multiple eigenvalues. The eigenvector corresponding to the largest eigenvalue among the multiple eigenvalues is taken as the first principal component direction vector, and the eigenvector corresponding to the second largest eigenvalue among the multiple eigenvalues is taken as the second principal component direction vector. Based on the first principal component direction vector, calculate the explained variance and the proportion of explained variance of the first principal component direction vector; based on the second principal component direction vector, calculate the explained variance and the proportion of explained variance of the second principal component direction vector.
7. The base recognition method based on breeding chips as described in claim 5, characterized in that, The Gaussian distribution fitting of the base brightness characteristics obtained by performing Gaussian distribution fitting based on the brightness values of each target probe scatter point in each base channel includes: For any type of probe, a multidimensional Gaussian distribution is used to fit the brightness values of each target probe scatter point in each base channel to obtain the center position of the multidimensional Gaussian distribution and the covariance matrix of the multidimensional Gaussian distribution. The center position of the multidimensional Gaussian distribution is used as the brightness mean vector based on the Gaussian distribution.
8. The base recognition method based on breeding chips as described in claim 1, characterized in that, The method includes: Obtain a training dataset, wherein each training sample in the training dataset includes scatter data sample features of a class of probes and base labels of detection sample sites corresponding to a class of probes. The base recognition model is trained based on the training dataset.
9. The base recognition method based on breeding chips as described in claim 8, characterized in that, The method further includes: Obtain the base tags of the detection sample sites; The base tags for obtaining the detection sample sites include: Selecting qualified sample loci from the sequencing data in whole-genome sequencing as detection sample loci, wherein the sequencing data includes base identification data of each sample locus and base identification quality value of each sample locus, wherein the conditions include at least one of the following: the base identification quality value is greater than a preset quality value, the sequencing depth is greater than a preset sequencing depth, and the ratio of the sum of the most and second most bases is greater than a first preset ratio value. When the maximum base percentage of the detected sample site exceeds the second preset percentage value, the detected sample site is determined to be a homozygous sample, and the label of the detected sample site is the homozygous sample label of the base type corresponding to the maximum base percentage; When the maximum base percentage of the detection sample site is less than or equal to the second preset percentage value, the detection sample site is determined to be a heterozygous sample, and the label of the detection sample site is a heterozygous sample label, which is formed by the base type corresponding to the maximum base percentage and the base type corresponding to the second largest base percentage.
10. A scanning device, characterized in that, It includes a memory and a processor, the memory storing a computer program that, when executed by the processor, causes the processor to perform the steps of the method as described in any one of claims 1 to 9.
11. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by a processor, the processor performs the steps of the method as described in any one of claims 1 to 9.