Cluster segmentation and conditional base calling

By segmenting clusters into subpopulations based on specific conditions and using a Gaussian mixture model for base calling, the variation in inter-cluster intensity profiles is minimized, reducing errors and enhancing sequencing run efficiency.

JP2025532443APending Publication Date: 2025-10-01ILLUMINA INC
0 Cites 0 Cited by

Patent Information

Application Number
JP2024557195
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2022-09-16
Filing Date
2023-09-15
Publication Date
2025-10-01

AI Technical Summary

Technical Problem

The variation in inter-cluster intensity profiles during sequencing runs leads to decreased data throughput and increased error rates in base calling due to factors such as fragment length distribution, phase shifts, bleaching, cluster size, overlapping colonies, uneven illumination, and impurities on the flow cell.

Method used

Segmenting clusters into subpopulations based on specific conditions and performing base calling separately for each subpopulation using a Gaussian mixture model to minimize inter-cluster intensity variation.

Benefits of technology

This approach reduces base calling errors and improves throughput by ensuring clusters with similar conditions are grouped together, allowing for accurate fitting and base calling within each subpopulation.

✦ Generated by Eureka AI based on patent content.
Patent Text Reader

Abstract

The disclosed technology is directed to cluster segmentation and base calling. The disclosed technology describes a computer-implemented method that includes segmenting a population of clusters into multiple subpopulations of clusters based on one or more previous bases called in one or more previous sequencing cycles of a sequencing run. In a current sequencing cycle of the sequencing run, the method includes applying a mixture of four distributions to current sequenced data of each subpopulation of clusters in the multiple subpopulations of clusters, the four distributions corresponding to the four bases adenine (A), cytosine (C), guanine (G), and thymine (T), and the current sequenced data generated in the current sequencing cycle. The method further includes base calling a cluster within a particular subpopulation of clusters using a corresponding mixture of the four distributions.
Need to check novelty before this filing date? Find Prior Art