A method for detecting abnormal ploidy in pear ploidy breeding
By combining Hough circle detection and areal density stability assessment with flow cytometry, the problem of high dependence on traditional clustering algorithms has been solved, achieving high accuracy and wide applicability in detecting ploidy anomalies in pear ploidy breeding.
Patent Information
- Application Number
- CN202310091457.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-02-09
- Publication Date
- 2025-12-02
- Estimated Expiration
- 2043-02-09
AI Technical Summary
In existing pear ploidy breeding methods, the clustering results of traditional clustering algorithms are highly dependent on parameters, resulting in low accuracy in identifying ploidy in breeding.
By combining Hough circle detection and areal density stability assessment with flow cytometry, the cluster core region and the region with the highest density in the flow cytometry scatter plot were obtained, and the fluorescence intensity of the scatter plot was used to determine whether the breeding ploidy was abnormal.
It improves the accuracy and applicability of breeding ploidy identification, eliminates the influence of subjective factors in parameter settings, and enhances the credibility of identification results.
Smart Images

Figure CN115984248B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of breeding ploidy abnormality detection technology, specifically to a method for detecting ploidy abnormalities in pear ploidy breeding. Background Technology
[0002] New pear varieties possess excellent traits that meet the needs of consumers and production. Therefore, pear breeding programs primarily aim to improve quality and appearance. Pluripotency breeding is the science that studies the laws governing ploidy variation in plant chromosomes and utilizes this variation to select new varieties. To improve the efficiency of pear breeding and select more superior, high-yielding, high-quality, and highly resistant new varieties, it is necessary to establish certain standards for identifying whether the mutated plant tissues have become homopolyploids.
[0003] Existing ploidy identification methods utilize flow cytometry to simultaneously measure multiple parameters of light scattering and different fluorescence in cells, enabling rapid and accurate qualitative and quantitative analysis of DNA content and chromosome ploidy in fruit tree cells. However, ploidy identification requires clustering, and the clustering results of traditional clustering algorithms are highly dependent on parameters. Since parameter settings are currently mostly subjective, their applicability and the reliability of clustering results are questionable, leading to low accuracy in breeding ploidy identification. Summary of the Invention
[0004] To address the issue of low accuracy in identifying ploidy in existing methods, this invention aims to provide a method for detecting ploidy anomalies in pear breeding. The specific technical solution adopted is as follows:
[0005] This invention provides a method for detecting ploidy abnormalities in pear ploidy breeding, the method comprising the following steps:
[0006] Extract leaf tissue from the pear tree to be tested, and process the leaf tissue to obtain a flow cytometry scatter plot corresponding to the leaf tissue to be tested.
[0007] The Hough circle detection is performed on the streaming scatter plot to obtain each Hough circle. The linear density of each Hough circle is obtained based on the number of scatter points on the edge line and the perimeter of each Hough circle. The areal density of each Hough circle is obtained based on the number of scatter points inside each Hough circle and the area of each Hough circle. Based on the linear density and areal density of each Hough circle, each Hough circle is filtered to obtain the cluster core region and the region to be merged in the streaming scatter plot.
[0008] Based on the surface density of the cluster core region and the scattered points in the region to be merged, the corresponding region to be merged is merged with the cluster core region to obtain the region with the highest density. Taking each edge scattered point of the region with the highest density as the center point, a structural element corresponding to each edge scattered point is constructed. Based on the number of scattered points in the structural element and the edge scattered points of the region with the highest density, the target subgroup is obtained.
[0009] Obtain the fluorescence intensity of all scattered points within the target subpopulation, determine the ploidy of the target subpopulation based on the fluorescence intensity, and determine whether the breeding ploidy is abnormal based on the ploidy.
[0010] Preferably, the step of merging the corresponding regions to be merged with the cluster core region based on the areal density of the cluster core region and the scattered points in the regions to be merged to obtain the regions with the highest density includes:
[0011] Select any region in the streaming scatter plot that intersects with or is externally tangent to the cluster core region, and denote it as the first region to be merged. Merge the first region to be merged with the cluster core region and denote the resulting region as the first region to be analyzed. Calculate the areal density of the first region to be analyzed based on the number of scatter points and the area of the first region. Determine the areal density stability corresponding to the first merge based on the areal density of the first region and the areal density of the cluster core region. Determine if the areal density stability corresponding to the first merge is greater than a stability threshold. If it is, merge the first region to be merged with the cluster core region and denote the resulting region as... The first merged region is defined as the first merged region. If the density is less than or equal to the first merged region, the cluster core region is defined as the first merged region. Any merged region that intersects with or is externally tangent to the first merged region in the streaming scatter plot is defined as the second merged region. The stability of the areal density corresponding to the second merge is determined to be greater than the stability threshold. If it is greater, the second merged region is merged with the first merged region, and the merged region is defined as the second merged region. If the density is less than or equal to the first merged region, the first merged region is defined as the second merged region. This process continues until all merged regions that intersect with or are externally tangent to the merged region are evaluated, and the finally merged region is defined as the region with the highest density.
[0012] Preferably, the formula for calculating the surface density stability corresponding to the Nth time is:
[0013]
[0014] Among them, T N The surface density stability corresponding to the Nth time step is given by exp(), which is an exponential function with the natural constant as the base. ε S represents the number of scattered points within the region obtained after the ε-th merging process. εLet ε represent the number of pixels in the region obtained after the ε-th merging process, A represent the surface density of the cluster core region, and N represent the N-th merging process.
[0015] Preferably, the step of filtering each Hough circle based on its linear density and areal density to obtain the cluster core region and the region to be merged in the streaming scatter plot includes:
[0016] For any Hough circle: Calculate the mean of the linear density and surface density of the Hough circle, and use it as the value of the objective function corresponding to the Hough circle;
[0017] The Hough circle corresponding to the maximum value of the objective function is taken as the cluster core region in the streaming scatter plot.
[0018] The Hough circles in the streaming scatter plot with a radius less than or equal to the radius of the cluster core region are designated as regions to be merged.
[0019] Preferably, obtaining the target subgroup based on the number of scattered points in the structuring element and the edge scattered points of the region connected by the maximum density includes:
[0020] The scattered points on the edge of the region with the highest density are denoted as edge scattered points;
[0021] For any edge point in the maximum density connected region: if there are no other points within the structuring element corresponding to the edge point, then the edge point is eroded to shrink it into the maximum density connected region to obtain a new edge pixel; if there are other points within the structuring element corresponding to the edge point, then the points within the structuring element located inside the maximum density connected region are designated as first-type points, and the points within the structuring element located outside the maximum density connected region are designated as second-type points; the average Euclidean distance between all first-type points within the structuring element corresponding to the edge point and the edge point itself is calculated and designated as the first average Euclidean distance; the average Euclidean distance between all second-type points within the structuring element corresponding to the edge point and the edge point itself is calculated and designated as the second average Euclidean distance; when the first average Euclidean distance is greater than or equal to the second average Euclidean distance, the edge point is dilated, and the second-type point closest to it is taken as the new edge pixel.
[0022] The target subgroup is obtained based on the new edge pixels.
[0023] Preferably, the step of processing the leaf tissue of the pear tree to be tested to obtain the flow cytometry scatter plot corresponding to the leaf tissue includes:
[0024] The surface of the pear tree leaf tissue to be tested was washed sequentially with distilled water and deionized water, and then subjected to dissociation, centrifugation, and refrigeration. The treated leaf tissue was then analyzed using flow cytometry to obtain the corresponding flow cytometry scatter plot.
[0025] Preferably, the step of obtaining the linear density of each Hough circle based on the number of scattered points on the edge line of each Hough circle and the circumference of each Hough circle includes:
[0026] For any Hough circle: calculate the ratio of the number of scattered points on the edge of the Hough circle to the circumference of the Hough circle, and use this as the linear density of the Hough circle.
[0027] Preferably, the step of obtaining the areal density of each Hough circle based on the number of scattered points within each Hough circle and the area of each Hough circle includes:
[0028] For any Hough circle: calculate the ratio of the number of scattered points within the Hough circle to the area of the Hough circle, and use this ratio as the areal density of the Hough circle.
[0029] Preferably, determining the ploidy of the target subpopulation based on the fluorescence intensity and judging whether the breeding ploidy is abnormal based on the ploidy include:
[0030] Calculate the average fluorescence intensity of all scatter points within the target subpopulation; based on the average fluorescence intensity, the fluorescence intensity of the reference object, and the ploidy of the reference object, obtain the ploidy of the target subpopulation;
[0031] Based on the ploidy of the target subgroup and the ploidy of the breeding direction, it is determined whether the breeding ploidy is abnormal.
[0032] The present invention has at least the following beneficial effects:
[0033] This invention first obtains the flow cytometry scatter plot corresponding to the leaf tissue to be tested. Due to the different levels of development of cell subpopulations, the region with the densest distribution of scatter points in the flow cytometry scatter plot is more representative of the leaf tissue characteristics. Therefore, this invention analyzes the distribution of scatter points in the flow cytometry scatter plot to obtain the region with the densest distribution. Considering that the traditional DBSCAN clustering algorithm requires manual parameter setting and is highly subjective, inappropriate parameter settings can lead to inaccurate acquisition of the region with the densest distribution of scatter points in the flow cytometry scatter plot, thus affecting the accuracy of subsequent identification of breeding ploidy. This invention combines Hough circle detection... Based on the linear and areal densities of the Hough circle, the cluster core region in the flow cytometry scatter plot was obtained. Then, based on the areal density of the cluster core region and the scatter point distribution in the regions to be merged, the corresponding regions to be merged were merged with the cluster core region to determine the region with the densest scatter point distribution, i.e., the region with the highest density connection. Based on this, the clustering results of the target subpopulation were obtained. The scatter points in the target subpopulation are relatively dense, therefore, the characteristics of the scatter points in the target subpopulation better reflect the ploidy of the target subpopulation. This invention uses the fluorescence intensity of the scatter points within the target subpopulation for ploidy identification to determine whether breeding ploidy is abnormal. The method provided by this invention does not rely on parameter settings, does not require pre-setting the search radius, and does not need to consider whether the parameter settings are appropriate. Furthermore, the size, linear density, areal density, and other arbitrary features are essentially distribution characteristics of the scatter points themselves, eliminating subjective factors and improving the accuracy of breeding ploidy identification results. It also has a wider range of applications and higher reliability. Attached Figure Description
[0034] To more clearly illustrate the technical solutions and advantages in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0035] Figure 1 A flowchart of a method for detecting abnormal ploidy in pear ploidy breeding provided by the present invention;
[0036] Figure 2 This is a flow cytometry scatter plot of the leaf tissue to be tested. Detailed Implementation
[0037] To further illustrate the technical means and effects adopted by the present invention to achieve the intended purpose, the following detailed description of a method for detecting abnormal ploidy in pear ploidy breeding based on the present invention is provided in conjunction with the accompanying drawings and preferred embodiments.
[0038] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains.
[0039] The following describes in detail, with reference to the accompanying drawings, a specific scheme for detecting abnormal ploidy in pear ploidy breeding provided by the present invention.
[0040] Example of a method for detecting abnormal ploidy in pear ploidy breeding:
[0041] This embodiment proposes a method for detecting ploidy abnormalities in pear ploidy breeding, such as... Figure 1 As shown, the pear ploidy breeding ploidy abnormality detection method of this embodiment includes the following steps:
[0042] Step S1: Extract leaf tissue from the pear tree to be tested, and process the leaf tissue to obtain a flow cytometry scatter plot corresponding to the leaf tissue to be tested.
[0043] The specific scenario addressed in this embodiment is as follows: In order to improve the quality of pears and breed more excellent, high-yielding, high-quality, and highly resistant new pear varieties, it is necessary to detect the ploidy of the mutated plant tissues. Therefore, this embodiment will detect the mutated leaf tissues to determine their corresponding ploidy and then determine whether the ploidy is abnormal.
[0044] First, extract approximately 2 cm of leaf tissue from the pear tree to be tested. 2 The extracted leaf tissue was washed sequentially with distilled and deionized water, filtered, and placed in a pre-chilled culture dish. 2 mL of pre-chilled dissociation buffer was added, and the leaf tissue was immersed in the dissociation buffer throughout the process. The pear leaf tissue was then shredded with a blade in one go. The tissue was then filtered into centrifuge tubes and refrigerated at 4°C. After 5 minutes, it was centrifuged at 1000 rpm for 5 minutes at 4°C. 2 μL of pre-chilled dissociation buffer and 300 μL of staining agent were added, mixed well, and refrigerated at 4°C in the dark to allow for thorough staining. The mixture was then transferred to sample tubes. Flow cytometry is a biological technique used for counting and sorting tiny particles suspended in fluids. Special fluorescent dyes (including PI, EB, A0, etc.) bind to DNA bases within cells. Cells stained with these dyes emit fluorescence under laser irradiation. The resulting fluorescence is converted into an electrical signal by a photoelectric converter. By detecting the labeled fluorescence signal, a high-speed, one-to-one quantitative analysis and sorting technique for single cells or other biological particles in a suspension is achieved. Therefore, this embodiment utilizes flow cytometry to detect the leaf tissue to be tested in the sample tube. Flow cytometry is used to obtain a flow scatter plot corresponding to the leaf tissue to be tested, such as... Figure 2As shown, a flow cytometer has two scattered light channels: the FSC channel and the SSC channel. FSC, or forward scattering, represents cell size; the larger the cell volume, the higher the FSC value. SSC, or side scattering, represents cell granularity; the more irregular the cell and the more protrusions on its surface, the higher the SSC value. Therefore, the FSC and SSC values can be used to compare cell size and granularity, thereby enabling cell grouping and classification. Each point on the flow cytometry scatter plot has a different level of fluorescence; the fluorescence intensity represents the DNA content.
[0045] At this point, a flow cytometry scatter plot of the leaf tissue to be tested was obtained, which will be used for subsequent detection of abnormalities in pear ploidy breeding.
[0046] Step S2: Perform Hough circle detection on the streaming scatter plot to obtain each Hough circle. Based on the number of scatter points on the edge line of each Hough circle and the perimeter of each Hough circle, obtain the linear density of each Hough circle. Based on the number of scatter points inside each Hough circle and the area of each Hough circle, obtain the areal density of each Hough circle. Based on the linear density and areal density of each Hough circle, filter each Hough circle to obtain the cluster core region and the region to be merged in the streaming scatter plot.
[0047] Leaf tissues contain multiple different types of subpopulations of cells, such as mesophyll cells, upper and lower epidermal cells, spongy tissue, and palisade tissue. It is necessary to separate the different subpopulations by setting gating, and then obtain the average fluorescence intensity of each subpopulation to obtain the ploidy identification results.
[0048] In the flow cytometry scatter plots of the leaf tissue to be tested, some cells are more developed than others. In such cases, classifying each subpopulation into different phyla can easily lead to errors in the average fluorescence intensity, as the classification results for cells with fewer distributions are difficult to control. Therefore, this embodiment starts with the core cluster region and uses the unique cluster obtained as the target subpopulation. Only the average fluorescence intensity of the target subpopulation is calculated to identify ploidy, avoiding errors caused by excessive differences between subpopulations. The reason why this embodiment uses the core cluster region to define the target subpopulation is that only when the number of cells is sufficiently large and dense can the ploidy identification be more reliable. In this case, the scatter plots of other subpopulations can be regarded as noise points for ploidy identification.
[0049] The specific process of phylogenetic classification of the flow cytometry scatter plot corresponding to the leaf tissue to be detected is as follows:
[0050] This embodiment uses the DBSCAN density clustering algorithm to obtain accurate clustering results. Considering that the core parameter of the DBSCAN clustering algorithm is the search radius, which directly affects the clustering results, but the search radius of existing DBSCAN clustering algorithms is mostly set subjectively, its application range is greatly limited. Improper parameter settings can lead to deviations in the clustering results, and ultimately, abnormal ploidy identification results. This embodiment first performs Hough circle detection on the flow cytometry scatter plot corresponding to the leaf tissue to be detected. Flow cytometry scatter plots are two-dimensional, and a Hough circle can be detected at any point on it. Therefore, a finite number of Hough circles can be detected on the flow cytometry scatter plot corresponding to the leaf tissue to be detected. This embodiment will construct an objective function about the Hough circle to obtain the densest core region in the flow cytometry scatter plot corresponding to the leaf tissue to be detected. Since the search range of the DBSCAN clustering algorithm is a circular range, this embodiment uses the Hough circle detection algorithm to better fit the search rules of the DBSCAN clustering algorithm. Moreover, the Hough circle detected based on the scatter points on the streaming scatter plot can reflect the distribution characteristics of the scatter plot itself. Compared with subjectively setting a local range or window for blind traversal, it can more accurately and quickly converge to the core area with the densest scatter point distribution.
[0051] In the above steps, Hough circle detection was performed on the flow cytometry scatter plot corresponding to the leaf tissue to be detected, obtaining multiple Hough circles in the flow cytometry scatter plot corresponding to the leaf tissue to be detected. The higher the proportion of scatter points on the edge line of the Hough circle and the higher the proportion of scatter points inside the Hough circle, the more likely the corresponding Hough circle is to be a region with a denser distribution of scatter points in the flow cytometry scatter plot. Therefore, the corresponding Hough circle is more suitable as the densest core region in the flow cytometry scatter plot. Linear density can limit the size of the Hough circle to a certain extent. When the scatter point distribution is more compact, both the linear density and the areal density will theoretically be larger. If only the areal density is considered, there may be Hough circles of different sizes with the same areal density. Based on this, the linear density of each Hough circle is obtained according to the number of scatter points on the edge line and the perimeter of each Hough circle. The areal density of each Hough circle is obtained according to the number of scatter points inside the Hough circle and the area of each Hough circle. Then, an objective function is constructed based on the linear density and areal density of each Hough circle. For any Hough circle, the value of its corresponding objective function is:
[0052]
[0053] Where E is the value of the objective function corresponding to the Hough circle, R is the radius of the Hough circle, π is pi, G is the number of scattered points on the edge of the Hough circle, W is the number of scattered points inside the Hough circle, and max() is the function to take the maximum value.
[0054] 2πR represents the circumference of the Hough circle. The density of the scattered points along the edge of the Hough circle, i.e., the linear density of the Hough circle; πR 2 This represents the area of the Hough circle. The density of the scattered points within the Hough circle is also known as the surface density of the Hough circle. This represents the mean of the linear density and areal density of the Hough circle. The larger this mean, the denser the scatter point distribution at the Hough circle. When both the linear density and areal density of the Hough circle are larger, it indicates that the scatter point distribution at the location of the Hough circle is denser.
[0055] Using the above method, the value of the objective function corresponding to each Hough circle can be obtained. The larger the value of the objective function, the more likely the corresponding Hough circle is to be a core region. Therefore, in this embodiment, the Hough circle corresponding to the maximum value of the objective function is taken as the cluster core region in the streaming scatter plot, where the scatter points are most densely distributed.
[0056] This embodiment obtains the cluster core region in the streaming scatter plot. Next, the cluster core region will be used as the initial cluster center, and a search will be performed in all directions of the cluster core region to obtain the region with the highest density of connected regions. In this embodiment, the search process is set as the merging of Hough circles. The radius of the Hough circle merged each time is less than or equal to the radius of the cluster core region. The Hough circles in the streaming scatter plot with a radius less than or equal to the radius of the cluster core region are recorded as regions to be merged, that is, multiple regions to be merged are obtained.
[0057] Step S3: Based on the surface density of the cluster core region and the scattered points in the region to be merged, merge the corresponding region to be merged with the cluster core region to obtain the region with the highest density. Using the scattered points at the edges of the region with the highest density as the center point, construct the structural element corresponding to each scattered point. Based on the number of scattered points in the structural element and the scattered points at the edges of the region with the highest density, obtain the target subgroup.
[0058] Any region in the streaming scatter plot that intersects with or is externally tangent to the cluster core region is selected and designated as the first region to be merged. Next, it is determined whether the first region to be merged can be merged with the cluster core region. Considering that if the areal density of the merged region is much smaller than that of the cluster core region, it indicates that the first region to be merged should not be merged with the cluster core region; therefore, this embodiment first designates the region after merging the first region to be merged with the cluster core region as the first region to be analyzed. Then, based on the number of scatter points in the first region to be analyzed and the area of the first region to be analyzed, the areal density of the first region to be analyzed is obtained. Based on the areal density of the first region to be analyzed and the areal density of the cluster core region, the areal density stability corresponding to the first merge is determined. The greater the areal density stability, the more suitable the first region to be merged is for merging. Therefore, a stability threshold T0 is set to determine whether the areal density stability corresponding to the first merge is greater than T0. If it is greater, the first region to be merged and the cluster core region are merged, and the region obtained after merging is recorded as the first merged region. If it is less than or equal to T0, the cluster core region is directly recorded as the first merged region. In this embodiment, the value of T0 is 0.75. In specific applications, the implementer can set it according to the specific situation. Next, select any region in the scatter plot that intersects with or is externally tangent to the first region to be analyzed. Record the region containing the Hough circle as the second region to be merged. Then, determine whether the second region to be merged can be merged with the first region. The determination method is the same as that for the first region, i.e., determine the areal density stability corresponding to the second merge. Check if the areal density stability corresponding to the second merge is greater than T0. If it is greater, merge the second region to be merged with the first region, i.e., perform the second merge. Record the resulting region as the second merged region. If it is less than or equal to T0, directly record the first merged region as the second merged region. Continue this process, calculating the areal density stability corresponding to the next merge after each merge. Based on the relationship between the areal density stability and the stability threshold, determine whether to perform the next merge, until all regions intersecting with or externally tangent to the merged region are evaluated. The specific expression for the areal density stability corresponding to the Nth merge is:
[0059]
[0060] Among them, T N The surface density stability corresponding to the Nth time step is given by exp(), which is an exponential function with the natural constant as the base. ε S represents the number of scattered points within the region obtained after the ε-th merging process. ε Let ε represent the number of pixels in the region obtained after the ε-th merging process, and A represent the surface density of the cluster core region.
[0061] Since the cluster core region is a Hough circle in the flow scatter plot, and the method for obtaining the areal density of each Hough circle has already been explained in step S2, it will not be repeated here. ε represents the ε-th merging process. Characterizing the surface density of the region obtained after the ε-th merging process, The difference between the areal density of the region obtained after the ε-th merging process and the areal density of the cluster core region reflects the change in the areal density of the region obtained after the ε-th merging process compared to the areal density of the cluster core region. This represents the variance of the areal density of the regions obtained from each merge up to the Nth merge process. The formula uses the density of the cluster core region instead of the mean term in the original variance formula to avoid the mean term deviating from the cluster center as the number of merges increases. This method utilizes a negative exponential function with a base of the natural constant to inversely normalize the surface density difference fluctuations. When the variance of the surface density of the regions obtained from each merge in the Nth merging process is larger, it indicates greater fluctuations in the surface density differences of the merged regions, meaning the surface density stability corresponding to the Nth merging process is lower.
[0062] Each time a merging process is performed, the surface density stability is calculated using the above method. The greater the surface density stability, the smaller the surface density difference fluctuation in the merging region, meaning that merging can be performed. Then, the search continues and other adjacent regions to be merged are judged until all Hough circles intersecting or externally tangent to the merged regions are judged. The finally obtained merged region is recorded as the region with the highest density.
[0063] This embodiment uses the Hough circle detection algorithm to obtain the maximum density connected region in the streaming scatter plot. The edge of the maximum density connected region is a relatively smooth arc-shaped edge, which may differ from the shape of the actual clustered region. Therefore, this embodiment will use morphology to trim the edge of the maximum density connected region. The streaming scatter plot is on a planar coordinate system, so it can be regarded as a grid. A structuring element of size k*k is selected, and the scatter points on the edge of the maximum density connected region are traversed as centers to obtain the structuring elements corresponding to each edge scatter point of the maximum density connected region. The edge scatter points are the scatter points on the edge. In this embodiment, the value of k is 5. In specific applications, the implementer can set it according to the specific situation.
[0064] For any edge point of the region with the highest density:
[0065] If there are no other scattered points within the structural element corresponding to the edge point, meaning the target point within the structural element is isolated, then the edge point is eroded, which means the center point of the structural element is eroded. The edge of the maximum density connected region at the center point shrinks inward to form a new edge. If there are other scattered points within the structural element corresponding to the edge point, then there may be two types of scattered points within the structural element: one type located inside the maximum density connected region, and the other type located outside the maximum density connected region. Scattered points located inside the maximum density connected region are denoted as the first type of scattered points, and scattered points located outside the maximum density connected region are denoted as the second type of scattered points. Based on the Euclidean distance between each first type of scattered point within the structural element corresponding to the edge point, the edge point is calculated. The average Euclidean distance between all first-type scattered points within the structuring element corresponding to a point and the edge scattered point is denoted as the first average Euclidean distance. Simultaneously, based on the Euclidean distance between each second-type scattered point within the structuring element corresponding to the edge scattered point and the edge scattered point, the average Euclidean distance between all second-type scattered points within the structuring element corresponding to the edge scattered point and the edge scattered point is calculated and denoted as the second average Euclidean distance. When the first average Euclidean distance is greater than or equal to the second average Euclidean distance, the edge within the structuring element corresponding to the edge scattered point is expanded outward to the nearest second-type scattered point within the structuring element; that is, the nearest second-type scattered point is taken as the new edge pixel. When the first average Euclidean distance is less than the second average Euclidean distance, no processing is performed. The method for calculating the Euclidean distance between points is a well-known technique and will not be elaborated further here.
[0066] For other edge pixels in the maximum density connected region, the above method is applied. Thus, the edge pixels in the maximum density connected region are corrected using the above method, and the corrected maximum density connected region is recorded as the target subgroup for subsequent anomaly detection in breeding ploidy.
[0067] Step S4: Obtain the fluorescence intensity of all scattered points within the target subpopulation, determine the species ploidy of the target subpopulation based on the fluorescence intensity, and determine whether the breeding ploidy is abnormal based on the species ploidy.
[0068] In this embodiment, the target subpopulation was obtained in step S3. Next, the scatter plots within the target subpopulation will be analyzed to determine the ploidy of the target subpopulation and then to determine whether the breeding ploidy is abnormal.
[0069] The fluorescence generated by the binding of the staining agent to DNA bases can characterize the DNA content, regardless of the unit used for the control calculation. The formula for calculating the ploidy of the target species is: Target species ploidy = DNA content of the test sample / DNA content of the reference sample * Reference sample ploidy. In this embodiment, the ploidy of the target subpopulation will be determined based on the fluorescence intensity of scatter plots. First, the fluorescence intensity of each scatter plot within the target subpopulation is obtained. Based on the fluorescence intensity of each scatter plot within the target subpopulation, the ploidy of the target subpopulation is calculated, i.e.:
[0070]
[0071] Where Q represents the ploidy of the target subgroup, M represents the number of scatter points within the target subgroup, and J... i Let K be the fluorescence intensity at the i-th scatter point within the target subgroup, and K be the fluorescence intensity of the reference analyte. The plurality of the reference object.
[0072] The average fluorescence intensity of all scatter points within the target subpopulation is used; the formula for calculating ploidy is an existing formula and will not be elaborated further here. Using the above method, the ploidy of the target subpopulation can be obtained. For example, assuming the DNA content or average fluorescence intensity of the target subpopulation is 100, the DNA content or fluorescence intensity of the reference is 50, and the ploidy of the reference is 2n, then the ploidy of the target subpopulation is...
[0073] The above steps yielded the ploidy of the target subpopulation. By comparing the ploidy of the target subpopulation with that of the breeding direction, it can be determined whether the breeding ploidy is abnormal. This completes the detection of ploidy abnormalities in pear breeding.
[0074] This embodiment first obtains the flow cytometry scatter plot corresponding to the leaf tissue to be tested. Since cell subpopulations have different levels of development, the region with the densest distribution of scatter points in the flow cytometry scatter plot better characterizes the leaf tissue. Therefore, this embodiment analyzes the distribution of scatter points in the flow cytometry scatter plot to obtain the region with the densest distribution. Considering that the traditional DBSCAN clustering algorithm requires manual parameter setting and is highly subjective, inappropriate parameter settings can lead to inaccurate acquisition of the region with the densest distribution of scatter points in the flow cytometry scatter plot, thus affecting the accuracy of subsequent identification of breeding ploidy. This embodiment combines Hough circle detection... Based on the linear and areal densities of the Hough circle, the cluster core region in the flow cytometry scatter plot was obtained. Then, based on the areal density of the cluster core region and the scatter point distribution in the regions to be merged, the corresponding regions to be merged were merged with the cluster core region to determine the region with the densest scatter point distribution, i.e., the region with the highest density connection. Based on this, the clustering results of the target subpopulation were obtained. The scatter points in the target subpopulation are relatively dense, so the characteristics of the scatter points in the target subpopulation can better reflect the ploidy of the target subpopulation. In this embodiment, ploidy is identified by the fluorescence intensity of the scatter points in the target subpopulation to determine whether the breeding ploidy is abnormal. The method provided in this embodiment does not rely on parameter settings, does not require pre-setting the search radius, and does not need to consider whether the parameter settings are appropriate. Furthermore, the essence of any feature such as the size, linear density, and areal density of the Hough circle is the distribution characteristic of the scatter points themselves, without subjective factors. This eliminates the influence of subjective factors, improves the accuracy of breeding ploidy identification results, and has a wider range of applications and higher reliability.
Claims
1. A method for detecting abnormal ploidy in pear ploidy breeding, characterized in that, The method includes the following steps: Extract leaf tissue from the pear tree to be tested, and process the leaf tissue to obtain a flow cytometry scatter plot corresponding to the leaf tissue to be tested. The Hough circle detection is performed on the streaming scatter plot to obtain each Hough circle. The linear density of each Hough circle is obtained based on the number of scatter points on the edge line and the perimeter of each Hough circle. The areal density of each Hough circle is obtained based on the number of scatter points inside each Hough circle and the area of each Hough circle. Based on the linear density and areal density of each Hough circle, each Hough circle is filtered to obtain the cluster core region and the region to be merged in the streaming scatter plot. Based on the surface density of the cluster core region and the scattered points in the region to be merged, the corresponding region to be merged is merged with the cluster core region to obtain the region with the highest density. Taking each edge scattered point of the region with the highest density as the center point, a structural element corresponding to each edge scattered point is constructed. Based on the number of scattered points in the structural element and the edge scattered points of the region with the highest density, the target subgroup is obtained. Obtain the fluorescence intensity of all scattered points within the target subpopulation, determine the species ploidy of the target subpopulation based on the fluorescence intensity, and determine whether the breeding ploidy is abnormal based on the species ploidy. The process of merging the corresponding regions to be merged with the cluster core region based on the areal density of the cluster core region and the scattered points in the regions to be merged, to obtain the connected regions with the highest density, includes: Select any region in the streaming scatter plot that intersects with or is externally tangent to the cluster core region, and denote it as the first region to be merged. Merge the first region to be merged with the cluster core region and denote the resulting region as the first region to be analyzed. Calculate the areal density of the first region to be analyzed based on the number of scatter points and the area of the first region. Determine the areal density stability corresponding to the first merge based on the areal density of the first region and the areal density of the cluster core region. Determine if the areal density stability corresponding to the first merge is greater than a stability threshold. If it is, merge the first region to be merged with the cluster core region and denote the resulting region as... The first merged region is defined as follows: if the density is less than or equal to the first merged region, the cluster core region is defined as the first merged region. Any merged region that intersects with or is externally tangent to the first merged region in the streaming scatter plot is defined as the second merged region. The stability of the areal density corresponding to the second merge is determined to be greater than the stability threshold. If it is greater, the second merged region is merged with the first merged region, and the merged region is defined as the second merged region. If the density is less than or equal to the first merged region, the first merged region is defined as the second merged region. This process continues until all merged regions that intersect with or are externally tangent to the merged region are evaluated, and the finally merged region is defined as the region with the highest density. No. The formula for calculating the surface density stability of the corresponding level is: , in, For the first The degree of stability of the corresponding surface density. It is an exponential function with the natural constant as its base. For the first The number of scattered points in the region obtained after the second merging process For the first The number of pixels in the region obtained after the second merging process The areal density of the cluster core region. For the first Secondary merge processing.
2. The method for detecting abnormal ploidy in pear ploidy breeding according to claim 1, characterized in that, The process of filtering each Hough circle based on its linear and areal density to obtain the cluster core region and the region to be merged in the streaming scatter plot includes: For any Hough circle: calculate the mean of the linear density and the areal density of the Hough circle, and use it as the value of the objective function corresponding to the Hough circle; The Hough circle corresponding to the maximum value of the objective function is taken as the cluster core region in the streaming scatter plot. The Hough circles in the streaming scatter plot with a radius less than or equal to the radius of the cluster core region are designated as regions to be merged.
3. The method for detecting abnormal ploidy in pear ploidy breeding according to claim 1, characterized in that, The step of obtaining the target subgroup based on the number of scattered points in the structuring element and the edge scattered points of the region connected by the maximum density includes: The scattered points on the edge of the region with the highest density are denoted as edge scattered points; For any edge point in the maximum density connected region: if there are no other points within the structuring element corresponding to the edge point, then the edge point is eroded to shrink it into the maximum density connected region to obtain a new edge pixel; if there are other points within the structuring element corresponding to the edge point, then the points within the structuring element located inside the maximum density connected region are designated as first-type points, and the points within the structuring element located outside the maximum density connected region are designated as second-type points; the average Euclidean distance between all first-type points within the structuring element corresponding to the edge point and the edge point itself is calculated and designated as the first average Euclidean distance; the average Euclidean distance between all second-type points within the structuring element corresponding to the edge point and the edge point itself is calculated and designated as the second average Euclidean distance; when the first average Euclidean distance is greater than or equal to the second average Euclidean distance, the edge point is dilated, and the second-type point closest to it is taken as the new edge pixel. The target subgroup is obtained based on the new edge pixels.
4. The method for detecting abnormal ploidy in pear ploidy breeding according to claim 1, characterized in that, The process of processing the leaf tissue of the pear tree to be tested to obtain the corresponding flow cytometry scatter plot includes: The surface of the pear tree leaf tissue to be tested was washed sequentially with distilled water and deionized water, and then subjected to dissociation, centrifugation, and refrigeration. The treated leaf tissue was then analyzed using flow cytometry to obtain the corresponding flow cytometry scatter plot.
5. The method for detecting abnormal ploidy in pear ploidy breeding according to claim 1, characterized in that, The method of obtaining the linear density of each Hough circle based on the number of scattered points on the edge line of each Hough circle and the circumference of each Hough circle includes: For any Hough circle: calculate the ratio of the number of scattered points on the edge of the Hough circle to the circumference of the Hough circle, and use this as the linear density of the Hough circle.
6. The method for detecting abnormal ploidy in pear ploidy breeding according to claim 1, characterized in that, The method of obtaining the areal density of each Hough circle based on the number of scattered points within each Hough circle and the area of each Hough circle includes: For any Hough circle: calculate the ratio of the number of scattered points within the Hough circle to the area of the Hough circle, and use this ratio as the areal density of the Hough circle.
7. The method for detecting abnormal ploidy in pear ploidy breeding according to claim 1, characterized in that, The determination of the species ploidy of the target subpopulation based on the fluorescence intensity, and the determination of whether the breeding ploidy is abnormal based on the species ploidy, include: Calculate the average fluorescence intensity of all scatter points within the target subpopulation; based on the average fluorescence intensity, the fluorescence intensity of the reference object, and the ploidy of the reference object, obtain the ploidy of the target subpopulation; Based on the ploidy of the target subgroup and the ploidy of the breeding direction, it is determined whether the breeding ploidy is abnormal.
Citation Information
Patent Citations
Method for quickly identifying pear black heart disease based on near-infrared diffuse transmission spectroscopy
CN110376159A
Detection method and equipment for measuring instability of genome, terminal equipment and computer readable storage medium
CN112669906A