Cell confluency system and method

An image-based software application with machine learning for cell confluency estimation addresses the challenge of monitoring cell cultures in large vessels, ensuring GMP compliance and improving manufacturing efficiency.

WO2026111887A1PCT designated stage Publication Date: 2026-05-28BAYER HEALTHCARE LLC
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
BAYER HEALTHCARE LLC
Filing Date
2025-11-06
Publication Date
2026-05-28

AI Technical Summary

Technical Problem

Current manufacturing methods for cell therapies lack an end-to-end automated solution for monitoring cell confluency in large, stacked cultivation vessels, which is crucial for ensuring biomass and quality in adherent cell cultures, and existing software solutions are not GMP compliant.

Method used

An image-based software application integrated with a high-throughput microscopy system uses machine learning to estimate cell confluency via pixel classification, with image and metadata processing in a cloud environment, and displays results through an interactive web-based interface, ensuring GMP compliance.

Benefits of technology

The solution provides real-time, automated cell confluency measurement, enhancing process understanding and reducing manufacturing variability, thereby streamlining the development and commercial production of cell therapeutics.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure US2025054288_28052026_PF_FP_ABST
    Figure US2025054288_28052026_PF_FP_ABST
Patent Text Reader

Abstract

The present embodiments relate to cell confluency estimation and manufacturing. Subject matter of the present embodiments provides computer-implemented methods, computer systems and computer-readable storage media for predicting the results of image-based cell confluency estimation and manufacturing methods.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Docket No. BYP240194 WO

[0002] Cell Confluency System and Method

[0003] Field

[0004] Systems, methods, and computer-implemented programs disclosed herein relate to biopharmaceutical manufacturing.

[0005] Background

[0006] Cell therapies target the replacement or restoration of impaired or damaged cells in patients. Though cell therapies have the potential to revolutionize the treatment of many severe conditions, the manufacture of living cells as a therapeutic agent poses considerable challenge to the biopharmaceutical industry. For example, currently established manufacturing methods are insufficient to monitor biological parameters. Furthermore, the applied equipment and analytics are often retrofitted for cell therapies rather than specifically designed or optimized for these processes.

[0007] Some of these challenges can be addressed by investing in the development of new Process Analytical Technology (PAT). These methods can be leveraged throughout the process development of cell therapeutics to provide increased process understanding in real-time and subsequently during commercial production for process control to reduce manufacturing variability and risk of batch failure.

[0008] However, the variety of cellular products, coupled with the diverse array of cell cultivation vessels, presents significant challenges for the integration of in-line or on-line PAT sensors and technologies. Certain vessels or bioreactor configurations pose substantial difficulties for sensor access, while specific cell types can exhibit heightened sensitivity to analytical perturbations. Moreover, conventional bulk analytical methods fail to provide adequate insights when dealing with non-homogeneous cell populations. The advent of advanced imaging technologies, augmented by the substantial advancements in machine learning (ML) over recent years, has been recognized as a promising solution. These tools offer the versatility, minimal biological and process interference, and the capability for single-cell or aggregate / colony resolution, thereby addressing these analytical challenges effectively. Docket No. BYP240194 WO

[0009] One such analytic gap is a readout for biomass production in adherent cultivation vessels. Cell confluency (the extent of cell coverage in a culture dish) is widely used as a proxy readout for biomass. However, many analytical devices for estimating confluency are developed for small cultivation vessels and rely on contrast-enhancing methods (e.g. phase contrast), both of which are incompatible with the large, stacked cultivation vessels that can be used for industrial production.

[0010] In addition to the monitoring device, routine confluency estimation may employ an automated solution that integrates image acquisition, storage, pre-processing, analysis and reporting in near-real time. Although there are several software packages such as Cell Profiler and ImageJ that facilitate the establishment of automated image analysis workflows in a laboratory setting, they are often not readily applicable to a regulated industrial environment. This is due to the absence of an end-to-end automated pipeline for data collection, modeling, analysis, and display. Therefore, there is a need for an image-based software application for cell confluency estimation that can streamline the development and commercial manufacturing of cell therapeutics.

[0011] Further, there is a need for improved cell confluency systems and methods for application to cell therapeutics development and manufacturing. These and other issues are addressed by the present embodiments.

[0012] Summary

[0013] The embodiments provide a computer implemented system for image-based cell confluency estimation, comprising an image data acquisition module for acquiring image data, a data modeling and analysis module for modeling and analyzing image-based data; and an automated data transfer and storage module for storing acquired image-based data acquired from the acquisition module and transferring it to an image-based data modeling and analysis module; and a result module for displaying and reporting the image-based results.

[0014] The embodiments provide a computer-implemented system for image-based cell confluency estimation wherein the image data acquisition module serves as the initial unit within an integrated application framework and a high-throughput image capturing device (e.g., CM20 Docket No. BYP240194 WO incubation monitoring system) is employed to automatically capture images of cells growing within a cell culture vessel.

[0015] The embodiments provide a computer implemented system for image-based cell confluency estimation wherein after data transfer and storage, images and metadata can be extracted from a premise device associated with an image capturing instrument and then can be transferred to a cloud-based Simple Storage Service (S3) solution from Amazon Web Services (AWS) for storage.

[0016] The embodiments provide a computer-implemented system for image-based cell confluency estimation wherein during data processing and analysis, the acquired image and metadata may undergo pre-processing.

[0017] The embodiments provide a computer-implemented system for image-based cell confluency estimation wherein following pre-processing, a machine learning model, trained to estimate cell confluency through image pixel classification, is utilized to analyze the processed image data and predict cell confluency and store the analysis results in a relational database (RDS) for subsequent display.

[0018] The embodiments provide a computer-implemented system for image-based cell confluency estimation wherein the predicted cell confluency results, along with other relevant quality metrics, are displayed through an interactive web-based interface.

[0019] The embodiments provide a computer implemented method for image-based cell confluency estimation, comprising acquiring image data; storing the acquired image data; transferring the stored and acquired image data; modeling and analyzing the transferred data to produce results; and displaying and reporting the results.

[0020] The embodiments provide a computer implemented method for image-based cell confluency estimation wherein the image data acquisition module serves as the initial unit within an integrated application framework and a high-throughput image capturing device (CM20 incubation monitoring system) and is employed to automatically capture images of cells growing within a cell culture vessel. Docket No. BYP240194 WO

[0021] The embodiments provide a computer-implemented method for image-based cell confluency estimation wherein after data transfer and storage, images and metadata are extracted from a premise device associated with an image capturing instrument and then transferred to a cloudbased Simple Storage Service (S3) solution from Amazon Web Services (AWS) for storage.

[0022] The embodiments provide a computer implemented method for image-based cell confluency estimation wherein during data processing and analysis, the acquired image and metadata undergo pre-processing.

[0023] The embodiments provide a computer-implemented method for image-based cell confluency estimation wherein following pre-processing, a machine learning model, trained to estimate cell confluency through image pixel classification, is utilized to analyze the processed image data and predict cell confluency and store the analysis results in a relational database (RDS) for subsequent display.

[0024] The embodiments provide a computer-implemented method for image-based cell confluency estimation wherein the predicted cell confluency results, along with other relevant quality metrics, are displayed through an interactive web-based interface.

[0025] The embodiments provide a computer-implemented system for image-based cell confluency estimation, comprising an image acquisition and processing framework comprising a front-end and a backend wherein the backend system is orchestrated through a computer that interfaces with an imaging instrument via a connection and the computer hosts an API-server software that facilitates communication through a Representational State Transfer (REST) and in tandem, a SCADA system, orchestrates the data flow between the image instrument, the backend API- server, and a remote Relational Database Service (RDS); a data analysis and integration framework in connection with the image acquisition and processing framework wherein a scheduled task continuously monitors the RDS for new images awaiting analysis and the images, fetched from their AWS S3 buckets, are subjected to analysis using a confluency estimation model, wherein the model’s output, including confluency metrics and other relevant statistics, is then recorded back into the RDS; an image pre-processing framework in connection with the data analysis and integration framework wherein prior to classification, every image is subjected to a bank of predefined filters which capture a diversity of features assessed across multiple scales including intensity (simple Gaussian blur), gradient (gradient magnitude of the Gaussian), Docket No. BYP240194 WO edge / peak strength (Hessian eigenvalues and the Laplacian of the Gaussian), and texture (structure tensor eigenvalues), wherein the stack of filter outputs is transformed into a matrix in which each row corresponds to a pixel coordinate in a particular image, and each column corresponds to a filter applied at a defined scale; a modeling and analysis framework in communication with the pre-processing framework wherein the confluency estimation task can be formulated as a binary classification in which each pixel is classified as either ‘foreground’ or ‘background’ and wherein a machine learning model for classification can be employed with a random forest classifier (RFC) to predict the label of each class from the features extracted for each pixel coordinate and determine results; and a result reporting framework for outputting and displaying the results.

[0026] The embodiments provide a computer-implemented system for image-based cell confluency estimation, further comprising labels to provide high-quality training data for the model training, wherein selected is one image from each timepoint, to allow for the model to capture a diversity in cell morphology from the time of seeding to nearly 100% confluency and at each timepoint, an image at a random position is selected to capture variability, wherein image pixels are labeled as either “foreground” or “background” based on selection.

[0027] The embodiments provide a computer-implemented system for cell therapy manufacturing and automated real-time image-based cell confluency measurement, comprising a cell incubator system for cell expansion; an image data acquisition system in contact with the cell incubator system, comprising at least an imaging device for acquiring cell confluency image data; a data modeling and analysis system for modeling and analyzing image data and distinguishing background non-confluency image data from foreground confluency image data; an automated data transfer and storage system in contact with the data modeling and analysis system for compiling non confluency and confluency image data; and a dashboard for displaying a heat map of the results of the compiled non confluency and confluency image data wherein the heat map provides an automated real time image-based cell confluency measurement.

[0028] The embodiments provide a computer-implemented system for cell therapy manufacturing and automated real-time image-based cell confluency measurement wherein the cell confluency measurements are in the range of from 0 to 100%. Docket No. BYP240194 WO

[0029] The embodiments provide a computer-implemented system for cell therapy manufacturing and automated real-time cell confluency measurement wherein the manufacturing system does not need to be disconnected.

[0030] The embodiments provide a computer-implemented system for cell therapy manufacturing and automated real-time image-based cell confluency measurement wherein the image data acquisition system comprises a CM20.

[0031] The embodiments provide a non-transitory computer-readable storage medium having stored thereon software instructions that, when executed by a processor of a computer system, cause the computer system to execute the following method acquiring cell confluency image data; storing the acquired cell confluency image data; transferring the stored and acquired cell confluency image data; modeling and analyzing the transferred cell confluency data to produce cell confluency results; and displaying and reporting the cell confluency results.

[0032] These and other features of the present teachings are set forth herein.

[0033] Docket No. BYP240194 WO

[0034] Brief Description of the Drawings

[0035] The skilled artisan will understand that the drawings, described below, are for illustration purposes only. The drawings are not intended to limit the scope of the present teachings or claims in any way.

[0036] FIG. 1 shows a high-throughput image capturing device (CM20 incubation monitoring system) is employed to automatically capture images of cells growing within the cell culture vessel. Images and metadata are extracted from the premise device and then transferred to a cloudbased Simple Storage Service (S3) solution from Amazon Web Services (AWS) for storage. A machine learning model estimates cell confluency through image pixel classification, and results are transferred into a relational database for display via a web-based dashboard.

[0037] FIG. 2 shows a network data pipeline, from instrument to cloud application. The first layer includes the instrument connected via USB to an on-premises computer. This computer running different containerized back-end applications (Ignition, RDS, API-server). The second layer hosts the SCADA (Ignition) front-end that sends commands to the instrument via an API. This layer transfers image files and metadata from to AWS S3 buckets and a relational database, respectively. Images and metadata are ultimately processed and displayed in an AWS-hosted cloud application.

[0038] FIG. 3 shows a selected model predictions for “foreground”, masked in black (background) and white (foreground). The right side of each image is unmasked for illustration purposes. Estimated confluency (“C”) and segmentation uncertainty (“U”) are reported as percentages in the bottom-left of each panel, (a-c) Early (immediately after seeding), middling (46 h after seeding), and late (91 h after seeding) timepoints for a single position, (d-f) Defective images identified by their high segmentation uncertainty, (d) A blurry, somewhat out-of-focus image, (e) A more severely out-of-focus image, exhibiting a vertical “doubling” defect, (f) A somewhat dim image with visible “streaking” across most of the field.

[0039] FIG. 4 shows an estimated confluency (a) and segmentation uncertainty (b) over the course of an experiment. Black points represent individual images, connected by shared position on the plate. Diamonds and intervals represent the mean confluency across all positions for a time-point, with a 95% confidence interval. Docket No. BYP240194 WO

[0040] FIG. 5 shows an interactive, web-based dashboard for confluency estimation results. Interactive plots (a) of confluency and uncertainty trends enable users to visualize per-timepoint heatmaps (b) of confluency / uncertainty as well as individual images (c) masked by predicted pixel classes (i.e. “foreground” or “background”). Users can toggle various configuration options (d) to adjust the display to their needs.

[0041] FIG. 6 shows a block diagram of an exemplary computer system of the present disclosure suitable for determining one or more results for a method based on process parameters.

[0042] Detailed Description

[0043] The embodiments will be more particularly elucidated below without distinguishing between the aspects of the embodiments (method, computer system, computer-readable storage medium). On the contrary, the following elucidations are intended to apply analogously to all the aspects of the embodiments, irrespective of in which context (method, computer system, computer-readable storage medium) they occur.

[0044] If steps are stated in an order in the present description or in the claims, this does not necessarily mean that the embodiment is restricted to the stated order. On the contrary, it is conceivable that the steps can be executed in a different order or else in parallel to one another, unless one step builds upon another step, thus absolutely requiring that the building step be executed subsequently (this being, however, clear in the individual case). The stated orders are thus embodiments of the present disclosure.

[0045] As used herein, the articles “a” and “an” are intended to include one or more items and can be used interchangeably with “one or more” and “at least one.” As used in the specification and the claims, the singular form of “a”, “an”, and “the” include plural referents, unless the context clearly dictates otherwise. Where only one item is intended, the term “one” or similar language is used. As used herein, the terms “has”, “have”, “having”, or the like are intended to be open-ended terms. Further, the phrase “based on” is intended to mean “based at least partially on” unless explicitly stated otherwise. Further, the phrase “based on” can mean “in response to” and be indicative of a condition for automatically triggering a specified operation of an electronic device (e.g., a controller, a processor, a computing device, etc.) as appropriately referred to herein. Docket No. BYP240194 WO

[0046] The operations in accordance with the teachings herein can be performed by at least one computer system constructed for the desired purposes or general-purpose computer system configured for the desired purpose by at least one computer program stored in a typically non- transitory computer readable storage medium.

[0047] A “computer system” is a system for electronic data processing that processes data by means of programmable calculation rules. Such a system usually comprises a “computer”, that unit which comprises a processor for carrying out logical operations, and peripherals.

[0048] In computer technology, “peripherals” refer to all devices which are connected to the computer and serve for the control of the computer and / or as input and output devices. Examples thereof are monitor (screen), printer, scanner, mouse, keyboard, drives, camera, microphone, loudspeaker, etc. Internal ports and expansion cards are, too, considered to be peripherals in computer technology.

[0049] Computer systems of today are frequently divided into desktop PCs, portable PCs, laptops, notebooks, netbooks and tablet PCs and so-called handhelds (e g., smartphone); all these systems can be utilized for carrying out the embodiments.

[0050] The term “non-transitory” is used herein to exclude transitory, propagating signals or waves, but to otherwise comprise any volatile or non-volatile computer memory technology suitable to the system.

[0051] The term “computer” should be broadly construed to cover any kind of electronic device with data processing capabilities, including, by way of non-limiting example, personal computers, servers, embedded cores, computing system, communication devices, processors (e.g., digital signal processor (DSP), microcontrollers, field programmable gate array (FPGA), system specific integrated circuit (ASIC, etc.) and other electronic computing devices.

[0052] The term “process” as used above is intended to comprise any type of computation or manipulation or transformation of data represented as physical, e.g., electronic, phenomena which can occur or reside e.g., within registers and / or memories of at least one computer or processor. The term processor comprises a single processing unit or a plurality of distributed or remote such units. Docket No. BYP240194 WO

[0053] Some implementations of the present disclosure will be described more fully hereinafter with reference to the accompanying drawings, in which some, but not all implementations of the disclosure are shown. Indeed, various implementations of the disclosure can be embodied in many different forms and should not be construed as limited to the implementations set forth herein; rather, these example implementations are provided so that this disclosure will be thorough and complete, and will fully convey the scope of the disclosure to those skilled in the art.

[0054] Cell therapies (CTs) have emerged as a potent treatment modality for a range of diseases that remain intractable to traditional therapeutic interventions. Despite their tremendous potential, the development and commercial manufacturing of these therapies present significant challenges to the biopharmaceutical industry. These challenges are primarily due to the unique nature of living cells as therapeutic agents and the complexities associated with ensuring their proper expansion and differentiation.

[0055] One of the metrics in cell therapy manufacturing is cell confluency, which not only acts as a proxy for biomass in adherent cultivation vessels, but also as a direct process parameter to ensure quality of the intermediate cell product, such as preventing growth-inhibition to minimize lag-phases of subsequent cultivations and paving the way for automated harvesting by enabling data-driven decisions. However, the measurement of cell confluency poses its own set of challenges. Currently available devices for estimating confluency are typically unsuitable for the large, stacked cultivation vessels often used in industrial production. Existing software solutions lack the end-to-end automation for data collection, modeling, analysis, and display, in a regulated industrial environment with standards like Good Manufacturing Practices (GMP).

[0056] In this disclosure, an efficient and highly flexible solution is presented to address these challenges: an image-based software application for cell confluency estimation integrated with a high-throughput microscopy system. The software application employs a machine learning model trained to estimate cell confluency via pixel classification. It extracts images and metadata from the microscope and stores them in a cloud environment, allowing for efficient data processing and analysis. These results, along with relevant quality metrics, are displayed via an interactive web-based interface. Docket No. BYP240194 WO

[0057] This industry-level platform, developed through the integration of process analytical technologies (PAT), digitalization, data engineering and data science technologies, enables automated image acquisition, storage, pre-processing, analysis, and reporting in near-real time. The entire setup was designed with an underlying focus on future GMP-compliance. It is believed that such platform technologies can significantly streamline the development and commercial manufacturing of cell therapeutics, thereby addressing the current challenges and positively contributing to the faster and consistent delivery of promising cell therapeutics to patients.

[0058] Materials and methods

[0059] Software Application Overview

[0060] To streamline the development and commercial production of cell therapeutics, the image-based software application for cell confluency estimation comprises several functional components. These include image data acquisition, automated data transfer and storage, data modeling and analysis, and result reporting.

[0061] FIG. 1 shows an embodiment of a computer implemented system 90. As depicted in FIG. 1, the image data acquisition module 100 stands as the initial unit within the integrated application framework. Specifically, a high-throughput image capturing device (CM20 incubation monitoring system or similar type device) 110 can be employed to automatically capture images of cells growing within the cell culture vessel.

[0062] Subsequently, as part of the automated data transfer module 120 and storage pipeline, images and metadata are extracted from the premise device (an edge computer) 112 associated with the image capturing instrument and then transferred to a cloud-based Simple Storage Service (S3) solution from Amazon Web Services (AWS) for storage 122.

[0063] For data processing and analysis, the acquired image and metadata undergo preprocessing. Following this, a machine learning module 130 and model, trained to estimate cell confluency through image pixel classification, is utilized to analyze the processed image data and predict cell confluency and store the analysis results into a relational database (RDS) 122 for subsequent display. Docket No. BYP240194 WO

[0064] Finally, the predicted cell confluency results, along with other relevant quality metrics, are presented through a results reporting module 140 and interactive web-based interface, implemented using, e.g., Dash for Python. The specific details of the functional components and dashboard are further discussed in subsequent sections.

[0065] Cell cultivation and monitoring

[0066] Human induced pluripotent stem cells (hiPSCs, episomal) from Gibco™ (A 18945) were grown in Essential 8™ medium (Gibco™, A1517001) in TC-treated CellSTACK® culture vessels with 1-5 chamber layers (Coming®, 3268) coated with human recombinant laminin 521 (BioLamina, LN521). Cultures were incubated at 37°C in a humidified atmosphere of 5% CO2 and passaged weekly with TrypLE™ Express (Gibco™, 12604013) for dissociation and ROCK inhibitor Y-27632 (Sigma- Aldrich, SCM075) to prevent apoptosis within the first 24 h after seeding. Medium was exchanged every second day, starting one day after seeding. The number of passages was kept below 30.

[0067] Automated microscopy systems from Evident / Olympus (Provi CM20) inside the incubator were used for constant monitoring. Each CM20 monitoring platform (hereinafter referred to as head) was connected to a compact PC (Lenovo, M920 Tiny) and controlled via the CM20H API (version 1.1.1). An image acquisition protocol (API-script) was created to acquire images of 2048x1536px from 35 positions as an equally spaced 5x7 grid within the observation window of the CM20 heads. The autofocus function was used to find the optimal focus plane for each position. LEDs for sample illumination were switched, when (reaching boundaries of the observation window), default exposure times were used.

[0068] Cultivation vessels were placed onto the CM20 heads in a way to ensure representative monitoring of the growth area as well as a leveled surface to avoid inhomogeneous cell or medium distribution. A cycle of the API-script was performed with an interval of 4 h, starting within an hour after seeding until end of cultivation. Acquired raw images were cached locally and transferred as batch via Ignition Supervisory Control and Data Acquisition (SCAD A) to an AWS-hosted cloud storage. Docket No. BYP240194 WO

[0069] Data pipeline

[0070] FIG 1 shows a data pipeline of image acquisition from instrument via network layers to cloud applications. Layer PO includes the instrument connected via USB to a computer. This computer running different containerized back-end applications (Ignition, RDS, APLserver). In layer Pl the front-end is located that sending commands to the backend Ignition. In addition, this layer transfers image files and metadata from PO to AWS S3 buckets and RDS respectively. In an AWS cloud application, images and metadata are processed and displayed in a front-end.

[0071] Image Acquisition and Processing Framework

[0072] The image acquisition process is bifurcated into two main components: the containerized backend and the frontend (See FIG. 2). The backend system is orchestrated through an edge computer that interfaces with the imaging instrument via a USB connection. This edge computer hosts an API-server software that facilitates communication through the Representational State Transfer (REST) protocol. In tandem, a SCADA system, specifically Ignition, orchestrates the data flow between the instrument, the backend API-server, and the remote Relational Database Service (RDS).

[0073] Operators are provided the capability to select and initiate protocols, assign specific names, and set parameters such as interval frequency and overall duration. Cell culture imaging typically occurs at four-hour intervals over the course of a week. Operators have the flexibility to pause, and resume runs, e.g. to perform tasks such as media exchange.

[0074] Within the backend, an Ignition script engine is employed to load predefined procedures, execute the corresponding commands, and relay them to the imaging instrument via the REST API. Post-image capture, the files are stored locally, and associated metadata is recorded in the RDS. This metadata encompasses details such as the sample date, AWS S3 storage location, filename, information about the procedure protocol, capture positions, and specific instrument data.

[0075] The data pipeline's integrity and security are bolstered by the Pl network layer, where two independent applications are responsible for synchronizing the image files and RDS metadata with AWS cloud storage, ensuring a separation between the frontend and backend processes. Docket No. BYP240194 WO

[0076] Data Analysis and Integration

[0077] A scheduled task continuously monitors the RDS for new images awaiting analysis. These images, fetched from their AWS S3 buckets, are subjected to analysis using the confluency estimation model, details of which are elaborated later. The model's output, including confluency metrics and other relevant statistics, is then recorded back into the RDS. This integration allows for real-time visualization through a dedicated dashboard and facilitates the seamless transfer of data to the appropriate electronic lab notebook (ELN), thus streamlining the data management and analysis workflow.

[0078] Image labels

[0079] To provide high-quality training data for model training, one image was selected from each timepoint, for a total of 23 images. This allowed for the model to capture a diversity in cell morphology from the time of seeding out to nearly 100% confluency. At each timepoint, an image at a random position (from a total of 35) was selected to capture variability. The intention of this strategy is to provide our model with a training data set that is richly informed by the potential variability in the images, possibly due to factors such as changes in illumination, condensation or defects on the plate, and other difficult-to-control factors.

[0080] Image pixels were labeled as either "foreground" or "background" based on criteria determined by user inspection. In total, only 143,130 pixels were labeled for training, less than 0.20% of the pixels in the 23 selected images. The Results demonstrate that the model can achieve high accuracy with relatively little training data. Ambiguous regions (not identifiably “foreground” or “background”) were intentionally not labeled.

[0081] Image preprocessing

[0082] Prior to classification, every image was subjected to a bank of predefined filters. As with other techniques (CITE Ilastik, Trainable Weka Segmentation), the filters capture a diversity of features assessed across multiple scales. Broadly speaking, these filters extract features such as intensity (simple Gaussian blur), gradient (gradient magnitude of the Gaussian), edge / peak strength (Hessian eigenvalues and the Laplacian of the Gaussian), and texture (structure tensor eigenvalues). The stack of filter outputs was transformed into a matrix in which each row Docket No. BYP240194 WO corresponds to a pixel coordinate in a particular image, and each column corresponds to a filter applied at some scale (a blur radius of 0.5, 1, 2, 4, or 8 pixels).

[0083] Modeling approach

[0084] The confluency estimation task can be formulated as a binary classification problem in which each pixel is classified as either ‘foreground’ or ‘background’. A machine learning model for classification was employed for this task. Specifically, a random forest classifier (RFC) was used to predict the label of each class from the features extracted for each pixel coordinate. The default settings for scikit-learn’ s implementation was used, except for the maximum tree depth, which was restricted to 20 based on a preliminary evaluation of the trade-off between tree depth and accuracy. The model was trained on a personal laptop, then serialized (via Python's pickle module) and uploaded to an S3 bucket for later retrieval during evaluation.

[0085] Calculations

[0086] Following evaluation of the classifier, the confluency for a given image was calculated as the fraction of pixels classified as foreground. The uncertainty for a given predicted pixel label was calculated as the entropy of the classification “probabilities”, i.e.. the fraction of decision trees in the forest that assigns a pixel to a given class. For two classes, the base-2 entropy gives uncertainty values ranging from zero (total agreement across all decision tree classifiers in the random forest ensemble) to one (even split between foreground and background predictions). The uncertainty for an image is calculated as the average uncertainty across all pixels.

[0087] GMP strategy

[0088] Dashboard testing was performed using Pytest. Separate Docker containers were utilized for serving the Dash (Python) application and simulating dashboard interactions via Selenium (facilitated by the Dash testing extensions). Containers were used to create controlled testing image sets using LocalStack (to emulate S3) and metadata using a containerized PostgreSQL database. Graphical outputs of tests, captured using Selenium, were validated against references using the Pytest-Regressions extension. Testing of the confluency estimation model utilized Pytest and extensions. Docket No. BYP240194 WO

[0089] Results

[0090] To enable consistent cell therapy manufacturing, a data science application comprising an image analysis model, and a dashboard was developed to assess confluency throughout the course of cell expansion in adherent cell culture.

[0091] Modeling results

[0092] A random forest classifier was trained on 143,130 labeled pixels sampled from 23 images captured over approximately four days. The classifier achieved an accuracy of 99.42% on the training data and a mean 5-fold cross-validation accuracy of 97.89%. On a test set of 690 randomly selected and manually labeled pixels (30 from each image), the classifier achieved an accuracy of 97.53%, and a class-balanced accuracy of 97.82%.

[0093] On a 2019 MacBook Pro, feature extraction was performed in about 113 seconds (about 5 seconds per image), while training the model itself took roughly 11 seconds. Evaluation on images took less than a second per image.

[0094] Model predictions

[0095] FIG. 3 illustrates model predictions on select images. FIG. 3A shows the predicted foreground overlaid on the original image for an early timepoint, shortly after seeding. At this timepoint, cells are characteristically “round” and have yet to form sizable colonies. The same position after 46 hours can be seen in FIG. 3B. Here, cells, have begun to flatten out and form moderately sized colonies. After 91 hours (FIG. 3C), the cells have grown to cover most of the field. Despite the differences in confluency (16-83%) and morphology, these images all have relatively low uncertainty (12-18%).

[0096] Model outputs for images with various defects are shown in FIGS. 3D-F. FIG. 3D appears to be slightly out of focus. Consequently, the indistinct colony boundaries and uncharacteristic textures lead to a poor, “cloudy” segmentation. FIG. 3E demonstrates a more severe focus issue. In this instance, there is a visible “doubling” effect, where the prominent cell bodies (seen here as dark blobs) appear a second time as a less intense “shadow” above their original positions. FIG. 3F shows a different defect; while much of the segmentation is acceptable, many areas are poorly segmented due to the “streaks” visible on the original image. Docket No. BYP240194 WO

[0097] The original image appears to be poorly illuminated. Although all three of these images segmented poorly, they were readily identifiable by their relatively high uncertainty (31-59%).

[0098] FIG. 4 shows the trajectory of a fully analyzed experiment over roughly four days, comprising images taken at 35 different positions at each of 23 timepoints (roughly 4 hours apart). After a lag phase, confluency begins to steadily increase starting around 24 hours. Several strong outliers in uncertainty are visually apparent in the latter half of the experiment, all corresponding to defective images (like those depicted in FIGS. 3D-F). Apart from these outliers, the baseline uncertainty is low, though baseline uncertainty does increase somewhat over the course of the experiment. Pending new experiment, including z-position data: Some of these outliers could be attributed to and identified by the focus of the images, which can be observed in ‘z-position’ metadata captured during imaging. While some ‘tilt’ between the CM20 and the imaged plate is expected, significant excursions from the overall orientation of the focal plane across the captured positions serve as another control for image quality and subsequent overall confluency estimates.

[0099] Dashboard

[0100] Confluency estimation results are relayed to users via an interactive, web-based interface deployed in a cloud environment (FIG. 5). Timeseries trends are depicted by a pair of graphs (FIG. 5A). When a user clicks a point on the trend graphs, the dashboard will display the corresponding image overlaid with the model predictions (FIG. 5C) for inspection. Selecting an image will display a positional heatmap of confluency and uncertainty for each parameter (estimated confluency, prediction uncertainty, and z-position) (FIG. 5B) for users to inspect for visual trends across the analyzed plate. Various visualization options for the graphs, heatmaps, and images can be configured (FIG. 5D).

[0101] GMP preparedness

[0102] To satisfy regulatory commitments for GMP, the validation of the modeling efforts and the dashboard functionality were considered separately. This was due to the significant differences in both the strategy and requirements of validating these components of the overall application. Docket No. BYP240194 WO

[0103] Validation of the machine learning model comprised rigorous documentation of the modeling strategy, including the feature extraction methods, the choice of classifier algorithm, documentation of the model accuracy, and demonstration of model performance on representative images of various properties (such as different levels of confluency, or characteristic examples of out-of-focus images). Model consistency as well as the behavior of supporting code were enforced using automated testing.

[0104] In contrast with the model validation strategy, which predominantly focused on the documentation of design decisions and performance, validation of the dashboard was instead an exercise in identifying and testing behaviors of the user interface. Documented user requirements were user to identify desired behaviors, which in turn were tested by simulating user interactions in an isolated environment with controlled data.

[0105] The application’s development included regular interaction among members of a crossfunctional team comprising process experts (including production scientists and engineers), data scientists, and cloud- technology engineers.

[0106] Adopting an agile development methodology greatly facilitated routine engagement among process experts and other SMEs. Starting with a set of solid user requirements, a prototype was developed quickly, which was subsequently tested by stakeholders in biologies manufacturing. Their feedback was useful in driving a new cycle of development to implement further enhancements. An iterative technology-development methodology such as what was employed in this work can address evolving needs of end users while providing them with a functional application for confluency process-monitoring tasks. A noteworthy feature driven by close collaboration with SMEs was the addition of heatmaps for selected timepoints, which has enabled users to identify spatially correlated issues like uneven iPSC seeding and damaged (e.g. scratched) plates. Unobserved, these issues could lead to costly data and product quality issues. This collaboration facilitated the timely implementation of minor features like point jittering, log-scaling, and other visualization options that enhance the interpretability of the data and modeling results.

[0107] Software-engineering best practices such as object-oriented programming, version control, and in-line code documentation ensured agile application development. Open-source tools for interface design enabled the team to address custom stakeholder requirements that could Docket No. BYP240194 WO not be accommodated easily within the limitations of commercial, proprietary software. As discussed previously, the Dash (Python) framework for defining and serving web applications allowed for highly customizable interfaces while requiring developers to create little to no custom HTML, CSS, or JavaScript.

[0108] An image-based cell confluency application was described as an example of platform technology for streamlining the development and manufacturing of cell therapies. The application was developed leveraging three various components: (a) high-throughput microscopy-based image generation of cells growing on plates, (b) automation of image data acquisition from the instrument and subsequent storage on a cloud environment and (c) data science software for image analysis and interactive user interface for reporting results. The application was built on AWS cloud using open-source tools such as Python.

[0109] The results from this work showed that cell confluency was effectively estimated from the image-based model even with a relatively small amount of training data by using traditional computer vision machine learning algorithm for image segmentation. An uncertainty metric was implemented to effectively identify potential low-quality images. Testing using unseen images was conducted and showed the ability of the model to generalize results well. The design of both the Ignition interface and the fully featured interactive dashboard was primarily driven by SME user needs and requirements and allows for quantification of variation in cell confluency across the plate.

[0110] The work provides an industry-level example of a general fully automated solution that enables the effective and efficient quantification of cell confluency, which is a step in the development and manufacturing of cell therapeutics. It clearly illustrates how integration of PAT, digital and data science technologies enables consistent and faster evaluation of cells, which are both important variables for scalable and cost-efficient processes that produce safe and efficacious cell therapies.

[0111] Along with the development of the confluency application the GMP validation strategy was formulated to ensure the application works as intended elements of the GMP validation strategy include extensive documentation of the modeling objectives and approach in conjunction with software application development best practices such as version control, modular design, object-oriented programming and automated unit and integration testing. The Docket No. BYP240194 WO cell confluency estimation application is general in that it can be applied to different cell types and different process (i.e. cell expansion, cell differentiation) by requiring training a new model using a new set of images but reusing the same components and cloud pipeline.

[0112] The present embodiments can be carried out by using a computer system. FIG. 6 illustrates an exemplary computer system 200. In connection therewith, the computer system 200 can be configured, by executable instructions, to implement the various algorithms and other operations described herein.

[0113] The exemplary computer system 200 can comprise, for example, one or more servers, workstations, personal computers, laptops, tablets, smartphones, other suitable computing devices, combinations thereof, etc. In addition, the computer system 200 can comprise a single computing device, or it can comprise multiple computing devices located in proximity or distributed over a geographic region and coupled to one another via one or more networks. Such networks can comprise, without limitations, the Internet, an intranet, a private or public local area network (LAN), wide area network (WAN), mobile network, telecommunication networks, combinations thereof, or other suitable network(s).

[0114] With that said, the illustrated computer system 200 comprises a processing unit 202 and a memory 204 that is coupled to (and in communication with) the processing unit 202. The processing unit 202 can comprise, without limitation, one or more processors (e.g., in a multicore configuration, etc.), including a central processing unit (CPU), a microcontroller, a reduced instruction set computer (RISC) processor, and system specific integrated circuit (ASIC), a programmable logic device (PLD), a gate array, and / or any other circuit or processor designed for the functions described herein. The above listing is exemplary only, and thus is not intended to limit in any way the definition and / or meaning of the term processing unit.

[0115] The memory 204, as described herein, is one or more devices that enable information, such as executable instructions and / or other data, to be stored and retrieved. The memory 204 can comprise one or more computer-readable storage media, such as, without limitation, dynamic random-access memory (DRAM), static random-access memory (SRAM), read only memory (ROM), erasable programmable read only memory (EPROM), solid state devices, flash drives, CDROMs, thumb drives, tapes, hard disks, and / or any other type of volatile or nonvolatile physical or tangible computer-readable media. The memory 204 can be configured to store, Docket No. BYP240194 WO without limitation, process parameters, outputs, models and / or other types of data (and / or data structures) suitable for use as described herein, etc. In various embodiments, computerexecutable instructions can be stored in the memory 204 for execution by the processing unit 202 to cause the processing unit 202 to perform one or more of the functions described herein, such that the memory 204 is a physical, tangible, and non-transitory computer-readable storage media. It should be appreciated that the memory 204 can comprise a variety of different memories, each implemented in one or more of the functions or processes described herein.

[0116] In the exemplary embodiments, the computer system 200 comprises an output unit 206 that is coupled to (and is in communication with) the processing unit 202. The output unit 206 outputs, or presents, to a user of the computer system 200, by, for example, displaying and / or otherwise outputting information such as, but not limited to, output, process parameters, and / or any other type of data. It should be further appreciated that, in some embodiments, the output unit 206 can comprise a display device such that various interfaces (e.g., systems (network-based or otherwise), etc.) can be displayed at computer system 200, and at the display device, to display such information and data, etc. In some examples, the computer system 200 can cause the interfaces to be displayed at a display device of another computing device, including, for example, a server hosting a website having multiple webpages, or interacting with a web system employed at the other computing device, etc. Output unit 206 can comprise, without limitation, a liquid crystal display (LCD), a light-emitting diode (LED) display, an organic LED (OLED) display, an “electronic ink” display, combinations thereof, etc. In some embodiments, the output unit 206 can comprise multiple units.

[0117] The computer system 200 further comprises an input device 208 that receives input from a user. The input device 208 is coupled to (and is in communication with) the processing unit 202 and can comprise, for example, a keyboard, a pointing device, a mouse, a stylus, a touch sensitive panel (e.g.. a touch pad or a touch screen, etc.), another computing device, and / or an audio input device. Further, in some exemplary embodiments, a touch screen, such as that included in a tablet or similar device, can perform as both the output unit 206 and the input device 208. In at least one exemplary embodiment, the output unit 206 and the input device 208 can be omitted. In addition, the illustrated computer system 200 comprises a network interface 210 coupled to (and in communication with) the processing unit 202 (and, in some embodiments, to the memory 204 as well). The network interface 210 can comprise, without limitation, a wired Docket No. BYP240194 WO network adapter, a wireless network adapter, a telecommunications adapter, or other device designed for communicating to one or more different networks. The network interface (210) and / or the input device (208) are referred to as receiving unit herein.

[0118] While the present embodiments have been described with reference to the specific examples, various modifications and changes can be made, and equivalents can be substituted without departing from the true spirit and scope of the claims appended hereto. The specification and examples are, accordingly, to be regarded in an illustrative rather than in a restrictive sense.

Claims

Docket No. BYP240194 WOCLAIMSWe Claim:

1. A computer-implemented system for image-based cell confluency estimation, comprising:(a) an image data acquisition module for acquiring image data;(b) a data modeling and analysis module for modeling and analyzing image-based data;(c) an automated data transfer and storage module for storing acquired image-based data acquired from the acquisition module and transferring it to an image-based data modeling and analysis module; and(d) a result module for displaying and reporting the image-based results.

2. The computer-implemented system for image-based cell confluency estimation as recited in Claim 1, wherein the image data acquisition module serves as the initial unit within an integrated application framework and a high-throughput image capturing device and is employed to automatically capture images of cells growing within a cell culture vessel.

3. The computer-implemented system for image-based cell confluency estimation as recited in Claim 1, wherein after data transfer and storage, images and metadata are extracted from a premise device associated with an image capturing instrument and then transferred to a cloudbased Simple Storage Service (S3) solution from Amazon Web Services (AWS) for storage.

4. The computer-implemented system for image-based cell confluency estimation as recited in Claim 1, wherein during data processing and analysis, the acquired image and data undergo pre-processing.

5. The computer- implemented system for image-based cell confluency estimation as recited in Claim 4, wherein following pre-processing, a machine learning model, trained to estimate cell confluency through image pixel classification, is utilized to analyze the processed image data and predict cell confluency and store the analysis results in a relational database (RDS) for subsequent display.Docket No. BYP240194 WO6. The computer-implemented system for image-based cell confluency estimation as recited in Claim 5, wherein the predicted cell confluency results, along with other relevant quality metrics, are displayed through an interactive web-based interface.

7. The computer-implemented system for image-based cell confluency estimation as recited in Claim 1, wherein the high throughput image-capturing device is a CM20 incubation monitoring system.

8. A computer-implemented method for image-based cell confluency estimation, comprising the steps of:(a) acquiring image data;(b) storing the acquired image data;(c) transferring the stored and acquired image data;(d) modeling and analyzing the transferred image data to produce results; and(e) displaying and reporting the results.

9. The computer-implemented method for image-based cell confluency estimation as recited in Claim 8, wherein the acquiring step is accomplished by a data acquisition module, further wherein the data acquisition module is positioned within an integrated application framework, and wherein said data acquisition module is a high-throughput image capturing device.

10. The computer- implemented method for image-based cell confluency estimation as recited in Claim 8. wherein after the transferring and storing steps, an extracting step is performed wherein images and metadata are extracted from a premise device associated with an image capturing instrument and then transferred to a cloud-based Simple Storage Service (S3) solution from Amazon Web Services (AWS) for storage.

11. The computer-implemented method for image-based cell confluency estimation as recited in Claim 8, wherein during data processing and analysis, the acquired image and data undergo pre-processing.Docket No. BYP240194 WO12. The computer-implemented method for image-based cell confluency estimation as recited in Claim 8, wherein following pre-processing, further comprising an analyzing step in which a machine learning model, trained to estimate cell confluency through image pixel classification, is utilized to analyze the processed image data and predict cell confluency and store the analysis results in a relational database (RDS) for subsequent display.

13. The computer- implemented method for image-based cell confluency estimation as recited in Claim 8, wherein the displaying step is achieved through an interactive web-based interface and wherein the results comprise predicted cell confluency results.

14. A computer implemented system for image-based cell confluency estimation, comprising:(a) an image acquisition and processing framework comprising a front-end system and a backend system wherein the backend system is orchestrated through a computer that interfaces with an imaging instrument via a connection and the computer hosts an API-server software that facilitates communication through a Representational State Transfer (REST) and in tandem, a SCADA system, orchestrates data flow between the imaging instrument, the backend system, and a remote Relational Database Service (RDS);(b) a data analysis and integration framework in connection with the image acquisition and processing framework wherein a scheduled task continuously monitors the RDS for new images awaiting analysis and the images, fetched from their AWS S3 buckets, are subjected to analysis using a confluency estimation model, wherein an output of the model, including confluency metrics is then recorded back into the remote RDS;(c) an image pre-processing framework in connection with the data analysis and integration framework wherein prior to classification, every image is subjected to a bank of predefined filters which capture a diversity of features assessed across multiple scales including intensity, gradient, edge / peak strength, and texture, wherein the stack of filter outputs is transformed into a matrix in which each row corresponds to a pixel coordinate in a particular image, and each column corresponds to a filter applied at a defined scale;(d) a modeling and analysis framework in communication with the pre-processing framework wherein the confluency estimation task can be formulated as a binary classification in which each pixel is classified as either ‘foreground’ or ‘background’ and wherein a machine learning model for classification can be employed with a random forest classifier (RFC) toDocket No. BYP240194 WO predict the label of each class from the features extracted for each pixel coordinate and determine results; and(e) a result reporting framework for outputting and displaying the results.

15. The computer implemented system for image-based cell confluency estimation as recited in Claim 14, further comprising labels to provide high-quality training data for the model training, wherein selected is one image from each timepoint, to allow for the model to capture a diversity in cell morphology from a seeding time to nearly 100% confluency and at each timepoint, an image at a random position is selected to capture variability, wherein image pixels are labeled as either ’’foreground” or “background” based on selection.

16. A computer implemented system for cell therapy manufacturing and automated realtime image-based cell confluency measurement, comprising:(a) a cell incubator system for cell expansion;(b) an image data acquisition system in contact with the cell incubator system, comprising at least an imaging device for acquiring cell confluency image data;(c) a data modeling and analysis system for modeling and analyzing image data and distinguishing background non-confluency image data from foreground confluency image data;(d) an automated data transfer and storage system in contact with the data modeling and analysis system for compiling non confluency and confluency image data; and(e) a dashboard for displaying a heat map of results of the compiled non confluency and confluency image data wherein the heat map provides an automated real time image-based cell confluency measurement.

17. The computer-implemented system for cell therapy manufacturing and automated realtime image-based cell confluency measurement, as recited in Claim 16, wherein the cell confluency measurements are in a range of from 0 to 100%.Docket No. BYP240194 WO18. The computer-implemented system for cell therapy manufacturing and automated realtime cell confluency measurement, as recited in Claim 17, wherein the manufacturing system does not need to be disconnected.

19. The computer-implemented system for cell therapy manufacturing and automated realtime image-based cell confluency measurement, as recited in Claim 18, wherein the image data acquisition system comprises a CM20.

20. A non-transitory computer readable storage medium having stored thereon software instructions that, when executed by a processor of a computer system, cause the computer system to execute the following steps:(a) acquiring cell confluency image data;(b) storing the acquired cell confluency image data;(c) transferring the stored and acquired cell confluency image data; (d) modeling and analyzing the transferred cell confluency data to produce results; and(f) displaying and reporting the cell confluency results.

Citation Information

Patent Citations

  • Method and system for calculating confluence degree of adherent cells

    CN115601747A

  • Machine-learned cell counting or cell confluence for a plurality of cell types

    US20230108453A1