Digital pathology machine learning infrastructure

The digital pathology machine learning infrastructure addresses the impracticality of transmitting large images by using embeddings vectors, enhancing efficiency and reducing resource demands for training and inference.

WO2026055331A1PCT designated stage Publication Date: 2026-03-12PROSCIA INC
View PDF 4 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2025-09-04
Publication Date
2026-03-12

AI Technical Summary

Technical Problem

The transmission of large digital pathology images over computer networks is impractical and intractable due to their size, requiring significant time and resources, which hinders the efficient training and inference of machine learning models.

Method used

A digital pathology machine learning infrastructure that provides embeddings vectors instead of images, using a server computer to derive and transmit embeddings through an API, eliminating the need for clients to handle or transmit large image files.

Benefits of technology

This approach reduces storage and data transfer requirements by several orders of magnitude, allowing efficient machine learning operations without the need for specialized tooling or local deployment of large models, and lowers hardware and expertise barriers.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure US2025044865_12032026_PF_FP_ABST
    Figure US2025044865_12032026_PF_FP_ABST
Patent Text Reader

Abstract

Techniques for using a digital pathology machine learning model implementation without requiring transmission of digital pathology images to a location of the digital pathology machine learning model implementation are presented. The techniques may include: providing, on a server computer, a digital pathology image embeddings API; receiving, from a client computer, an embeddings job request including an identification of at least one digital pathology image, an image resolution instruction, and an identification of an embeddings network; passing, to an embeddings server, metadata characterizing the digital pathology image(s) resolved according to the resolution instruction; obtaining, from the embedding server, at least one embeddings vector after the embeddings server transforms a resolved set of digital pathology image(s) into at least one embeddings vector; and transmitting, by the server computer, the embeddings vector(s) to a storage location, without the client computer transmitting or receiving the digital pathology image(s).
Need to check novelty before this filing date? Find Prior Art

Description

Attorney Docket No. 0293.0012-PCTDIGITAL PATHOLOGY MACHINE LEARNING INFRASTRUCTURERelated Application

[0001] This application claims the benefit of U.S. Provisional Patent Application No. 63 / 691 ,075, entitled “Digital Pathology Machine Learning Infrastructure,” and filed September 5, 2024.Field

[0002] This disclosure relates generally to digital pathology, e.g., digital pathology machine learning training and inference.Background

[0003] For digital pathology, development of machine learning, especially deep learning, typically involve many steps. For example, medical professionals, scientists, and researchers may download large images (such as whole-slide images, which are each up to multiple gigabytes in size) from storage, manipulate them to detect tissue regions, tile those images to yield smaller image patches, and then use a large number of such images (or corresponding collections of tiles) to train a deep neural network.

[0004] Part of the training of a deep neural network for digital pathology may include training initial layers to extract features and provide a smaller, compressed set of features, which subsequent layers of the network or other downstream machine learning systems use to produce a prediction. Neural network layers that provide compressed sets of features may include convolutional layers or image transformer layers, by way of non-limiting examples. The extracted features, typically presented in the form of a low-dimensional numerical vector derived from an image, serve as compact representations of the image from which they are derived. Each number in a feature vector represents an amount of some property of the image, where a property is generally learned by a machine learning system during training and may relate to anything from a single pixel to the entire image. By way of non-limiting examples, a property may represent a presence of a vertical line in the image at a particular location, a corner in the image, a color in the image, a saturation at a portion of theimage, or a property that does not have a ready human interpretation. Typically, a unique deep learning model is used for each new classification task, requiring a large amount of training data, compute, time, and effort for each new machine learning project.

[0005] Foundation models (e.g., DinoV2, available at aodv.org / abs / 2304.07193, PLIP, available at arxiv.org / abs / 2201 .03545, and CtransPath, available at www.sciencedirect.com / science / article / pii / S1361841522002043) are a subject of current research. Foundation models generally include large neural networks that may be trained on very large image datasets to convert novel images into corresponding features, typically in the form of feature vectors. Foundation models include a series of layers, which may include an image transformer architecture and / or a convolutional neural network architecture, that ultimately culminate in a final output layer. The output layer typically produces a 1 *N vector, with N being the dimension of the embedding. Thus, foundation models encode or compress an input image into a single output vector. Many foundation models are designed in such a way that the features they generate - generally compressed “encodings” or “embeddings” of the input image - can be used as input to simple (or complex) downstream machine learning models that perform a variety of tasks.

[0006] Although speeds are increasing, transmitting large files over computer networks such as the internet remains onerous. For example, digital pathology typically uses whole-slide images, which can be multiple gigabytes in size. Even with an extremely fast internet connection of 100 megabits per second, transmitting a single whole-slide image may take well over a minute. A machine learning application may utilize thousands - or more - whole-slide images, and their transmission may very well take well over a day, or over a week for tens of thousands of such images. For a typical machine learning digital pathology application, transmission of the associated whole-slide images is highly impractical, if not outright intractable.Summary

[0007] Systems, methods, and non-transitory computer readable media are presented for using a digital pathology machine learning model implementation without requiring transmission of digital pathology images to a location of the digital pathologymachine learning model implementation. The techniques include: providing, on a server computer, a digital pathology image embeddings application program interface (API), where the server computer is communicatively coupled to: a client computer, an embeddings server, an image retrieval server, and a digital pathology image server; receiving, at the digital pathology image embeddings API and from the client computer, an embeddings job request, where the embeddings job request includes: (1 ) an identification of a set of at least one digital pathology image hosted by the digital pathology image server, (2) an image resolution instruction, and (3) an identification of an embeddings network; passing, by the server computer and to the embeddings server, (1 ) metadata characterizing the set of at least one digital pathology image resolved according to the resolution instruction and (2) an indication of the identification of the embeddings network; obtaining, by the server computer and from the embedding server, a set of at least one embeddings vector after the embeddings server transforms, by an embeddings network implementation corresponding to the identification of the embedding model, a resolved set of at least one digital pathology image into the set of at least one embeddings vector, where the resolved set of at least one digital pathology image corresponds to the set of at least one digital pathology image after having been resolved by the image retrieval server according to the resolution instruction; and transmitting, by the server computer, the set of at least one embeddings vector to a non-transitory embeddings storage location, for the client computer to use the digital pathology machine learning model implementation to process the set of at least one embeddings vector without the client computer transmitting or receiving the set of at least one digital pathology image.

[0008] Various optional features include the following. The image resolution instruction may be indicative of an image tiling, and where the set of at least one digital pathology image resolved according to the resolution instruction may include a set of image tiles. The techniques may include transmitting, by the server computer and to the client computer, image tile origination information indicative of image locations from which image tiles in the set of image tiles were extracted. The image resolution instruction may be indicative of a thumbnail image, and the set of at least one digital pathology image resolved according to the resolution instruction may include at least one downsampled image. The image resolution instruction may be indicative of a size of a pathology sample portion represented by each pixel, and the set of at least onedigital pathology image resolved according to the resolution instruction may include at least one digital pathology image with pixels each representing the size. The image resolution instruction may include an image segmentation instruction indicative of a request to segment, and the set of at least one digital pathology image resolved according to the resolution instruction may include at least digital pathology image segmented into at least one region of interest. The set of at least one digital pathology image resolved according to the resolution instructions may omit at least one region not of interest. The image resolution instruction may include an annotation of at least one downsampled digital pathology image. The non-transitory embeddings storage location may be remote from the server computer. The method may include selecting the embeddings server according to proximity to the digital pathology image server. The techniques may include: receiving, by the server computer and from the client computer, a request to upload a custom embeddings network, where the identification of the embeddings network includes an identification of the custom embeddings network, and where the embeddings network implementation includes an implementation of the custom embeddings network. The embedding job request may further include an image augmentation instruction, and the set of at least one digital pathology image resolved according to the resolution instruction may include images augmented according to the image augmentation instruction. The image augmentation instruction may include instructions for at least one of: rotation, color jitter, noise introduction, or stain augmentation. The client computer may train the digital pathology machine learning model implementation on the set of embeddings. The digital pathology machine learning model implementation may perform inference on the set of embeddings. The embedding server may include the image retrieval server. The digital pathology image server may include the image retrieval server.

[0009] Combinations, (including multiple dependent combinations) of the above-described elements and those within the specification have been contemplated by the inventors and may be made, except where otherwise indicated or where contradictory.Brief Description of the Drawings

[0010] Various features of the examples can be more fully appreciated, as the same become better understood with reference to the following detailed description of the examples when considered in connection with the accompanying figures, in which:

[0011] Fig. 1 is a schematic diagram of example architecture of a system for facilitating use of a digital pathology machine learning model implementation without requiring transmission of digital pathology images to a location of the digital pathology machine learning model implementation, according to various embodiments;

[0012] Fig. 2 depicts a flow diagram for a non-limiting example method of facilitating use of a digital pathology machine learning model implementation, without requiring transmission of digital pathology images to a location of the digital pathology machine learning model implementation, according to various embodiments; and

[0013] Fig. 3 illustrates an example network topology of a system for facilitating use of a digital pathology machine learning model implementation, without requiring transmission of digital pathology images to a location of the digital pathology machine learning model implementation, according to various embodiments.Description of the Examples

[0014] Reference will now be made in detail to example implementations, illustrated in the accompanying drawings. Wherever convenient, the same reference numbers will be used throughout the drawings to refer to the same or like parts. In the following description, reference is made to the accompanying drawings that form a part thereof, and in which is shown by way of illustration specific exemplary examples in which the invention may be practiced. These examples are described in sufficient detail to enable those skilled in the art to practice the invention and it is to be understood that other examples may be utilized and that changes may be made without departing from the scope of the invention. The following description is, therefore, merely exemplary.

[0015] Some embodiments present a digital pathology machine learning infrastructure that allows a client computer to obtain or use embeddings of digital pathology images without the client computer ever having to download the digital pathology images. Some embodiments provide an embeddings network implementation that a client computer can access using an Application ProgrammingInterface (API). (According to various embodiments, an “embeddings network” is a neural network, such as a foundation model, or a neural sub-network, trained to derive a compressed set of features, e.g., a feature vector, from image data. As used herein, the term “implementation” embraces at least one electronic processor coupled to non- transitory storage, where the storage includes instructions that are executable by the processor(s) to perform defined actions.) An embodiment may use an embeddings network implementation to generate embeddings from one or more digital pathology images that the embodiment retrieves from a storage location and serve the embeddings to the client computer via the API. The embeddings network may include, for example, a deep learning model, a foundation model, or a different embeddings / encoding model. The client computer that uses the API can use the embeddings to, e.g., train a digital pathology machine learning model implementation or to use a trained digital pathology machine learning model implementation to perform inference, such as classification or regression (such as on one or more of the embeddings). According to some embodiments, the client computer may be geographically and / or topologically remote from the storage location and never access the digital pathology images itself. Further, according to some embodiments, the storage location and the embeddings network implementation may be topologically and / or geographically close to each-other, e.g., co-located, such that the embeddings network may obtain and process the digital pathology images efficiently.

[0016] Non-limiting example embodiments may be described herein in reference to an “embedding API server.” (In general, the term “server” is used herein to refer to a computer that provides one or more services to another computer, which may be referred to as a “client” or “client computer.”) An embedding API server may operate in conjunction with, by way of non-limiting example, a computer-based whole slide Image Management System (IMS) and / or digital pathology solution. A client computer can make requests via the embedding API server for a single embedding per whole-slide image or multiple embeddings per whole-slide image (e.g. one embedding per whole-slide image patch or tile). Through the client computer, a user may specify an underlying embeddings network (e.g., a deep learning network or foundation model) to use to generate the embeddings, via the embedding API server, among some pre-specified embeddings networks. The user may also use theembedding API server to specify the location and format of their own trained embeddings network that the embedding API server is to use to generate embeddings.

[0017] Embodiments solve several significant problems of the prior art.

[0018] Some embodiments solve the internet-centric problem of training a digital pathology machine learning model implementation, or using a digital pathology machine learning model implementation for inference, when the associated digital pathology images are large. According to the prior art, such actions require transmission of huge amounts of image data, e.g., for training or inference, from a location at which the image data is acquired (e.g., a medical imaging facility) or stored (e.g., a medical records system) to a location of the digital pathology model implementation (e.g., a digital pathology diagnostic facility). Transmission of such voluminous image data is impractical or intractable. Some embodiments solve this problem by avoiding the need to transmit the image data itself to the location of the digital pathology model implementation. Instead, some embodiments provide embeddings, also referred to features or encodings, such as feature vectors, which occupy several orders of magnitude less space that the images themselves. For example, an RGB whole-slide image crop of 512 pixels on each side contains 512x512x3 = 786,432 unsigned 8-bit integers, or 786,432 bytes. By contrast, a Vision Transformer (ViT) feature vector (embedding) of such a crop contains 768 floats at 4 bytes apiece, for 3,072 bytes. The feature vector provides a compressed representation of the image, with a compression rate of 256:1. This means that instead of requiring the transmission of one or more 1 Gb whole-slide images, some embodiments may instead transmit one or more corresponding embeddings, which are each less than 4 Mb. Transmission of the embeddings derived from whole-slide images may take orders of magnitude less time than the transmission of the wholeslide images themselves (e.g., seconds instead of hours, or hours instead of weeks). Embodiments may solve the prior art digital pathology image transmission problem by transmitting a set of embeddings vector(s) to a storage location, for a client computer to use a digital pathology machine learning model implementation to process the embeddings vector(s) without the client computer transmitting or receiving any digital pathology images.

[0019] By transmitting, by a server computer, a set of embeddings vector(s) to an embeddings storage location for a client computer to process using a digitalpathology machine learning model implementation, without the client computer transmitting or receiving the associated digital pathology image(s), embodiments have profound advantages over prior art techniques that require transfer of the digital pathology images. For example, far less (e.g., by several orders of magnitude) storage is required to house and manipulate the embeddings than would be required to start from scratch and store / manipulate the corresponding whole-slide images. As another example, data transit requirements are far less than the alternative of operating directly on whole-slide image files, and therefore the process is less costly in terms of time and (depending on the setting) money. According to some embodiments, the client computer that executes the digital pathology machine learning model implementation does not have a need to download whole-slide images and instead downloads only embeddings, which are far smaller (e.g., by several orders of magnitude).

[0020] Thus, some embodiments solve the prior art problem of digital pathology image transmission by providing digital pathology machine learning infrastructure, which includes a server computer that provides a digital pathology image embeddings application program interface (API), where the server computer is communicatively coupled to a client computer, an image retrieval server, and a digital pathology image server, and an embeddings server that provides embeddings. This unconventional and non-generic combination of infrastructure components facilitates digital pathology machine learning without having to transmit image data to or from the client computer, and without the client computer having to manipulate image data. Further, for example, some embodiments provide the embedding API server at a specific topological and / or geographical location within a computer network in relation to one or more of the other infrastructure components, in order to increase the efficiency of interactions among the components.

[0021] Yet further, some embodiments solve the prior art problem of digital pathology image transmission by having the client computer send an embeddings job request to the digital pathology image embeddings API, instead of sending or receiving voluminous image data. The embeddings job request may include one or more of: (1 ) an identification of a set of one or more digital pathology images hosted by the digital pathology image server, (2) an image resolution instruction, and (3) an identification of an embeddings network. The server computer passes to the embeddings server: (1 ) metadata characterizing the set of digital pathology image(s) resolved accordingto the resolution instruction, and (2) an indication of the identification of the embeddings network to be used to obtain the embeddings. This alleviates the client computer from having to perform the embeddings itself. The image retrieval server, rather than the client computer, transforms the set of digital pathology image(s) into a resolved set of digital pathology image(s) according to the resolution instruction, freeing the client computer from having to manipulate voluminous image data. The embeddings server, rather than the client computer, then transforms the resolved set of digital pathology image(s) into the set of embeddings vector(s) using an embeddings network implementation corresponding to the identification of the embedding model. The server computer obtains this set of embeddings vector(s) from the embedding server, and provides them to a location at which the client computer can use them, e.g., for machine learning training or inference, without the image data itself ever being present at, or passing to or from, the client computer.

[0022] Not only are digital pathology images such as whole-slide images very large, making their storage, manipulation, handling, transfer, and retrieval expensive and difficult, but also, whole-slide images often contain a pyramid of multiple images at a wide range of resolutions, which require specialized tooling in order to read. Some embodiments solve this prior art problem regarding the need for specialized tooling through the use of an infrastructure that includes an image retrieval server, which resolves (e.g., performs image processing on) one or more digital pathology images into a resolved set of digital pathology image(s), and an embeddings server, which transforms the resolved set of digital pathology image(s) into a set of embeddings vector(s). These specialized infrastructure components, separate from the client computer that uses the embedding vector(s), permit the client computer to benefit from compact embeddings derived from whole-slide images, without having to store, manipulate, handle, transfer, or retrieve the whole-slide images themselves, and without requiring the client computer to implement the specialized tooling needed to process multi-resolution image pyramid data.

[0023] Yet another problem with the prior art is due to the fact that, because even single whole-slide images may be too large for real-world embeddings networks to operate on, workarounds such as image tiling and gridding are sometimes used. Prior art implementation of these workarounds requires complicated and error-prone careful data management to track the lineage and relationship of each tile back to theoriginal whole-slide image. Some embodiments solve this prior art data management problem by storing, e.g., in an embedding database, image tile origination information indicative of whole-slide image locations from which image tiles were extracted, in association with the embeddings derived from such image tiles. Some embodiments use a dedicated image retrieval server to perform the tiling, gridding, and associated data management. The data management information may be transmitted to the client computer, allowing the client computer to benefit from image tiling and gridding, without having to perform the attendant complicated and error-prone data management.

[0024] Yet another problem with the prior art is that the selection and implementation of a particular embeddings network for a particular digital pathology machine learning task is difficult. For example, foundation models that are performant on specific types of data are constantly evolving. The flexibility to experiment with different embeddings networks is important to success in this domain, particularly in the context of digital pathology, but prior art digital pathology machine learning techniques include only a fixed foundation model in a given context, for example. Some embodiments solve this prior art problem of lack of flexibility in the selection and implementation of embeddings networks by allowing a user, through a client computer, to send an embeddings job request that includes an identification of an embeddings network to use, and then passing, by the server computer that receives the request and to a dedicated embeddings server, an indication of the identification of the embeddings network. This infrastructure-based technique facilitates simple and essentially hot-swappable embeddings networks in a given digital pathology machine learning model implementation. Further, some embodiments make it easier to track the embeddings network that was used to generate a particular dataset and thereby ensure reproducible results, by storing an indication of the embeddings network that was used to generate the embeddings stored in an embeddings database in association with the embeddings themselves.

[0025] That some embodiments eliminate the need for the user to deploy embeddings networks themselves advantageously allows the user to save the time and capital that would otherwise be required to manage the underlying deployment infrastructure. For example, the actual serving of the embeddings networks is handled by the embedding API server, so it does not require any local or DIY deployment oflarge models on GPU compute resources, which can be prohibitive for some potential users of digital pathology machine learning model implementations.

[0026] That some embodiments eliminate the need for the end user of a digital pathology implementation to train their own embeddings network, by providing an embeddings server that includes or accesses pre-trained embeddings networks, itself has several advantages. According to some embodiments, the user does not need to collect a large amount of digital pathology images with which to train a digital pathology regression or classification machine learning implementation, and thus data requirements for their (downstream) digital pathology machine learning model implementation are reduced by several orders of magnitude. There are no Graphics Processing Unit (GPU) requirements to train a downstream digital pathology regression or classification machine learning implementation that utilizes the embeddings provided by an embodiment, and therefore training is easily accomplishable on a standard desktop or laptop CPU. This substantially reduces the computer hardware requirements for digital pathology machine learning. In general, downstream digital pathology machine learning training lowers the bar of expertise required, because it can be accomplished with much simpler machine learning systems and no sophisticated knowledge of deep learning.

[0027] Thus, some embodiments provide a technology-based solution to the problem of digital pathology machine learning implementation by obviating the transmission of digital pathology images to a location of the digital pathology machine learning model implementation.

[0028] These and other features and advantages are shown and described in detail in reference to the figures.

[0029] Fig. 1 is a schematic diagram of example architecture of a system 100 for facilitating use of a digital pathology machine learning model implementation without requiring transmission of digital pathology images to a location of the digital pathology machine learning model implementation, according to various embodiments. The system 100 includes a client computer 102, a digital pathology image server 120, and core digital pathology machine learning infrastructure 110 according to some non-limiting embodiments, as described in detail presently.

[0030] The client computer 102 represents a client computer with respect to an embedding API server 112. For example, and in brief, the client computer 102 maysend a request identifying image(s) and an embedding model to the embedding API server 112, and the client computer 102 (or a related computer) may consume a set of corresponding embeddings vector(s), e.g., by using them to train a machine learning model or by using a trained machine learning model to perform inference on the embedding vector(s). According to some embodiments, the client computer 102 may host the digital pathology machine learning model implementation.

[0031] The digital pathology image server 120 hosts the actual digital pathology images that the system 100 processes. The images may include whole-slide images stored in TIFF format, for example. Multiple sets of images may be stored, including images that depict various tissue types. For example, a set of images may include images of a particular tissue type, some of which including a particular characteristic, such as a tumor, artifact, or other trait, and some of which not including the particular characteristic. A set of images may include metadata in the form of labels that specify whether or not a specified characteristic is present in each image. Thus, a set of images stored at the digital pathology image server 120 may be used as a labeled (or unlabeled) digital pathology machine learning training corpus for any of a variety of characteristics. According to some embodiments, the digital pathology image server may be collocated with a clinical laboratory that acquires the images, by way of nonlimiting example.

[0032] The non-limiting example core digital pathology machine learning infrastructure 110 includes the embedding API server 112, an embedding database 113, an embedding server 114, embedding storage 116, and an image retrieval server 118. These components are described presently.

[0033] The embedding API server 112 hosts the embedding API and interacts with the client computer 102 to provide image embeddings for the client computer 102 (or a related computer that hosts a digital pathology machine learning model implementation) to use. The embedding API server 112 may coordinate between other components as described in detail herein. The embedding API server 112 may be selected to be in geographic and / or network topological proximity to the digital pathology image server 120. For example, the embedding API server 112 may be selected so that the network pipeline between it and the digital pathology image server meets or exceeds a specified threshold bandwidth requirement e.g., greater than 100 Mbps, such as 1 Tbps.

[0034] The embedding server 114 performs the actual derivation of embeddings from images. Typically, the embedding server 114 includes a high-power processor, which may include many Graphics Processing Units (GPU), so as to be capable of efficiently executing an embeddings network to process images into embeddings, e.g., feature vectors or safetensors. The embedding server 114 may include multiple embeddings network implementations, e.g., both a convolutional neural network embeddings network and a vision transformer embeddings network, according to some embodiments. According to some embodiments, a user may upload a custom embeddings network to the embedding server 114 by way of a request sent from the client computer 102. Note that the custom embeddings network upload may be directly from the client computer 102, or from a different computer, e.g., by the client computer providing a link to the customer embeddings network on the different computer. The embedding server 114 may implement the custom embeddings network, and the user may subsequently request that the custom embeddings network be used to generate embeddings.

[0035] The image retrieval server 118 may perform various image processing actions on the images prior to their conversion by the embedding server 114 into embeddings. Examples of such image processing actions include various resolving actions. For example, the image retrieval server 118 may perform one or more of: image tiling / patching (e.g., where an image is parsed into multiple parts, which may or may not overlap), downsampling (e.g., to convert original images to much smaller thumbnail images, which may be used for various purposes in addition, or in the alternative, to converting them into embeddings), segmenting (e.g., partitioning into one or more regions of interest and one or more other regions), augmenting (e.g., one or more of: rotation, color jitter, noise introduction, and / or stain augmentation), and / or determining pixel resolution (e.g., such that a specified size of a pathology sample portion is represented by each pixel). Note that some of these examples, such as tiling and determining pixel resolution, represent processing that does not generally occur outside of the field of machine learning for digital pathology. (As used herein, a “thumbnail image” of a particular image represents a downsampling of the particular image. The thumbnail image will occupy much less memory than the particular image from which it is derived. For example, the ratio of their sizes may be 1 :64 or smaller.)

[0036] The embedding storage 116 stores embeddings that the embedding server 114 derives from digital pathology images per jobs requested by the client computer 102. By way of non-limiting example, the embeddings for one or more whole-slide images processed according to a job request may be stored and provided in a file that includes metadata about the embeddings job and each whole-slide image processed. According to some non-limiting example embodiments, the file may or may not include lightweight thumbnail images of the tiles. Further by way of nonlimiting example, the embeddings themselves may be stored and provided as a dictionary of tensors with keys in “Y_X” format indicating the grid index of the embedded patch within a whole-slide image.

[0037] According to some embodiments, the embedding storage 116 may be implemented in a cloud server. According to some embodiments, the embeddings storage 116 provides both storage and data management capabilities to assist users with long term storage, query, and retrieval of their embeddings.

[0038] The embedding database 113 serves as a ledger for the embedding jobs. For example, the embedding database 113 stores copies of the metadata for each embedding job and the corresponding whole-slide images, as well as the location of the embeddings in the embedding storage 116. The embedding database may store image tile origination information indicative of whole-slide image locations from which the image tiles were extracted in association with the embeddings derived from such image tiles. When the client computer 102 requests an embedding job by way of the embedding API server 112, the embedding API server 112 queries the embedding server 114, and the embedding server 114 responds to the embedding API server 112 with data that indicates the location of the embeddings in the embedding storage 116, as well as any associated metadata (e.g., image labels), for the job. The embedding API server 112 provides data indicative of the location, as well as any requested job metadata (or its storage location) back to the client computer 102 so that the client computer 102, or a related computer, can download the embeddings and / or metadata.

[0039] Note that according to various embodiments, the architectures and / or functionalities of two or more of the above-referenced components may be merged or distributed. For example, one or more of the above-listed functions of the image retrieval server 118 may be performed by the embedding API server 112 and / orembedding server 114. Further, two or more of the above-listed functions of the image retrieval server 118 may be split among the embedding API server 112 and / or the embedding server 114. As another example, according to various embodiments, the digital pathology image server 120 may perform downsampling to provide thumbnail images to the embedding API server 112, which may perform image segmentation on the thumbnail images. According to such an embodiment, the segmented thumbnail images may be provided to the digital pathology image server 120, which may tile the original images based on the segmented thumbnail images, and the digital pathology image server 120 may provide only tiles that include a region of interest to the embedding server 114 to convert into embeddings. As yet another example, one or more of the above components may be implemented using multiple computers. For example, the embedding server 114 may include multiple computers, each of which may host one or more embeddings network implementations.

[0040] Note that in some embodiments, the digital pathology machine learning infrastructure 110 may train a digital pathology regression or classification machine learning implementations using the provided embeddings. For example, some embodiments allow a user to upload a digital pathology regression or classification machine learning model and model training parameters (e.g., number of epochs, optimizer, etc.), via the embedding API, to the embedding API server, or to a different computer in the digital pathology machine learning infrastructure 110, where it is trained using the provided embeddings according to the provided training parameters. Any of a variety of different computers may be used to train the uploaded digital pathology regression or classification machine learning model, including, by way of non-limiting examples, the embedding API server 112, the embedding database 113, the embedding server 114, the embedding storage 116, the image retrieval server 118, or even the digital pathology image server 120.

[0041] Fig. 2 depicts a flow diagram for a non-limiting example method 200 of facilitating use of a digital pathology machine learning model implementation, without requiring transmission of digital pathology images to a location of the digital pathology machine learning model implementation, according to various embodiments. The method 200 may be implemented using a system such as is shown and described herein in reference to Fig. 1 , by way of non-limiting example.

[0042] In general, the method 200 may include a user of a client computer directing the client computer to send one or more requests to an embedding API server (e.g., multiple requests in bulk) to provide embeddings for a single whole-slide image, a list of whole-slide images, or a whole repository of whole-slide images stored at a digital pathology image server. A non-limiting example use case for the method 200 is that the user may request a single embedding per whole-slide image, and a thumbnail image of the whole-slide image may be retrieved from the digital pathology image server and input to a foundation model, or other embeddings network, at an embeddings server. The user may request that the derived embeddings be sent to a computer that hosts a digital pathology model implementation. Such a computer may be the client itself, or a different computer that is typically, but not necessarily, colocated with, or on the same intranet as, the client computer.

[0043] At 202, an embedding API server receives a request from a client computer to create embeddings for one or more whole-slide images. In general, when the user makes an embedding request through the client computer, the client computer sends the request though a network such as the internet, and it arrives at a queue at the embedding API server. The request may include or be accompanied by one or more details of the requested embeddings job, e.g., a specific slide list and / or image repository list, an objective power / microns per pixel I magnification, a specific embeddings network to use, whether tissue segmentation should be performed or not, tile or thumbnail, etc.

[0044] Optionally, the request may include or be accompanied by annotations of regions of interest of the images, such that the embedding API server will ignore any regions outside the annotated area(s). According to some embodiments, the annotations may optionally be provided in the form of coordinates and dimensions of a bounding box that encloses one or more regions of interest. According to some embodiments, the annotations may optionally be provided in the form of one or more annotated thumbnail images. According to this optional feature, the client computer may request thumbnail images from the embedding API server, and the client computer or a related computer may annotate and return the thumbnail images to the embedding API server as part of, or associated with, the embedding request.

[0045] Thus, in general, the request may include one or more of: (1 ) an identification (e.g., a location, link, list, etc.) of a set of one or more digital pathologyimages hosted by the digital pathology image server, (2) an image resolution instruction (e.g., specifying parameters for one or more of: image tiling / patching, downsampling, segmenting, and / or magnification), (3) an identification of an embeddings network, and / or (4) an image augmentation instruction (e.g., specifying parameters for one or more of: rotation, color jitter, noise introduction such as Gaussian noise, and / or stain augmentation). The request may be sent via the embedding API of the embedding API server.

[0046] According to some embodiments, prior to making the request, the client computer may request an authentication token via the embedding API server. Once the client computer has this token, the client computer can provide it as an authenticator to the embedding API server to make requests via the API. When the user submits an authenticated request for embeddings, the embedding API validates the user authorization and requests whole-slide image data retrieval, e.g., from a digital pathology image server.

[0047] According to some embodiments, the user may request an estimation of cost for a specific embedding request. An owner of the embedding API server may charge users of the client based on number of tiles and resolution, according to some embodiments.

[0048] At 204, the embedding API server retrieves basic metadata (e.g., slide identification, storage location, labels, magnification, etc.) for the whole-slide image(s) of the request and forms embedding jobs for each such slide. Then, for each wholeslide image, the method 200 generates embeddings per 206, 208, and 210 as follows.

[0049] At 206, the image retrieval server extracts image data, e.g., from the digital pathology image server, at the magnification and patch sizes specified in the request, and submits tiled image data to the embedding server. According to some embodiments, as part of this process, the embedding API server obtains a token from the digital pathology image server, and provides it to the image retrieval server for the image retrieval server to obtain the digital pathology images.

[0050] Optionally, at 208, the images may be filtered and / or undesired content may be removed. For example, tiles may be filtered to leave or remove tiles containing tissue. As another example, tiles may be filtered to leave or remove tiles containing artifacts. As yet another example, tiles may be filtered to leave or remove tiles containing tumor. As yet another example, regions (e.g., tiles) of whole-slide imagesmay be excluded, e.g., regions that do or do not contain tissue, tumor, and / or artifacts. Any of the filtering described herein may be accomplished using a trained machine learning implementation that identifies background, tissue, tumor, and / or artifacts, as are known in the art. Any of the filtering described herein may be accomplished using detailed geometric annotations. In general, the disclosed filtering and / or content removal process advantageously minimizes network transfer costs by only transferring data that is to be used, in a readily useable format.

[0051] At 210, the embeddings server executes the specified embeddings network on the images (e.g., in one or more batches of patches). According to various embodiments, the embeddings server may handle one or more of: resource and embeddings network management for embeddings networks available to the embedding service, computation of embeddings given tiled image data, and / or returning embeddings for each tile requested. The embeddings server provides the completed embeddings to the embedding storage.

[0052] At 212, the system (e.g., the embedding API server) collates embeddings metadata and allows the client computer to download the embeddings from the embedding storage. According to some embodiments, the embeddings and / or any optional outputs (e.g., metadata) may be stored in either temporary or permanent storage (e.g., the embedding database). Optional outputs that may be so retrievably stored and provided to the user include one or more of:

[0053] • One or more thumbnail masks indicating the regions of the image containing tiles determined to be valid, having artifacts, and / or containing tissue.

[0054] • Coordinates in the grid over the image, and / or resolution, of tiles passed on to the embedding server. Resolution information for an entire whole-slide image may include tile size (e.g., including height and width), grid size (e.g., including number of rows and number of columns of tiles), tile overlap size (e.g., in the width and / or height directions), pad size (e.g., in terms of height and / or width), etc. The tile coordinates are particularly helpful for bookkeeping during the use of a downstream digital pathology machine learning model implementation, including during both training and inference.

[0055] • The valid tiles or whole-slide images (e.g., tile images showing a region of interest) themselves, e.g. as JPG, PNG, or BMP, format.

[0056] • An embedding model name and version used by the embedding server to produce the embeddings.

[0057] • Image labels (e.g., regarding the presence / absence of a particular characteristic).

[0058] • Image file names (e.g., including coded identifications of respective patients, slide numbers, etc.).

[0059] • Any other job metadata supplied at input, e.g. an indication of resolution (e.g., square microns per pixel) and / or an indication of the associated digital pathology images (e.g., an image list and / or a repository identification or list).

[0060] Thereafter, a user at the client computer may use the unique job token to retrieve the embeddings and any accompanying data or messages (e.g., error messages) from the embedding storage and / or embedding database. According to some embodiments, the system provides the user with a URL at which the embeddings and any associated metadata may be retrieved.

[0061] Fig. 3 illustrates an example network topology 300 of a system for facilitating use of a digital pathology machine learning model implementation, without requiring transmission of digital pathology images to a location of the digital pathology machine learning model implementation, according to various embodiments.

[0062] In general, according to some embodiments, the embedding API server and its associated subservices are selected or constructed in a way that facilitates efficient data transfer and minimizes wasted resources. In order to achieve this, the embedding server and image retrieval server may be selected or constructed such that they operate nearby the image data stored at the digital pathology image server both geographically and in terms of network connectivity. For example, the network pipeline between, on the one hand, the embedding server and image retrieval server, and, on the other hand, the digital pathology image server, may be required to meet or exceed a specified threshold bandwidth requirement e.g., greater than 100 Mbps, such as 1 Tbps. Such an arrangement allows for a reduction in required resources in comparison to a design that moves large image data between networks, e.g., prior art systems that operate directly on whole-slide image files that must be downloaded to servers that are capable of operating large foundation models.

[0063] As shown in Fig. 3, the client computer 302 may communicate with the embedding API server 310 through the internet 350. The embedding API server 310,including its RESTful API 312, may communicate with the other components 320 through the internet 350 or through a particular network 360 that provides high-speed connectivity (e.g., greater than 100 Mbps). The other components 320, such as the digital pathology image server 324 (which may be part of a greater digital image management service 322), the image retrieval server 326, and / or the embedding server(s) 328 may be co-located, in the same intranet, or otherwise communicatively coupled using a very high bandwidth channel (e.g., greater than 100 Mbps, such as 1 Tbps, for example using a Dedicated Internet Access - DIA - connection). Thus, Fig. 3 illustrates that the embedding server(s) 328 are able to operate close to the digital pathology image server 324 for improved performance relative to a decoupled system of networks.

[0064] Certain examples can be performed using a computer program or set of programs. The computer programs can exist in a variety of forms both active and inactive. For example, the computer programs can exist as software program(s) comprised of program instructions in source code, object code, executable code or other formats; firmware program(s), or hardware description language (HDL) files. Any of the above can be embodied on a transitory or non-transitory computer readable medium, which include storage devices and signals, in compressed or uncompressed form. Exemplary computer readable storage devices include conventional computer system RAM (random access memory), ROM (read-only memory), EPROM (erasable, programmable ROM), EEPROM (electrically erasable, programmable ROM), flash memory, and magnetic or optical disks or tapes.

[0065] Aspects of the present disclosure are described herein with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the disclosure. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented using computer readable program instructions that are executed by an electronic processor.

[0066] These computer readable program instructions may be provided to a processor of a general-purpose computer, special purpose computer, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the electronic processor of the computer or otherprogrammable data processing apparatus, create means for implementing the functions / acts specified in the flowchart and / or block diagram block or blocks. These computer readable program instructions may also be stored in a computer readable storage medium that can direct a computer, a programmable data processing apparatus, and / or other devices to function in a particular manner, such that the computer readable storage medium having instructions stored therein comprises an article of manufacture including instructions which implement aspects of the function / act specified in the flowchart and / or block diagram block or blocks.

[0067] In embodiments, the computer readable program instructions may be assembler instructions, instruction-set-architecture (ISA) instructions, machine instructions, machine dependent instructions, microcode, firmware instructions, statesetting data, configuration data for integrated circuitry, or either source code or object code written in any combination of one or more programming languages, including an object oriented programming language such as Smalltalk, C++, or the like, and procedural programming languages, such as the C programming language or similar programming languages. The computer readable program instructions may execute entirely on a user's computer, partly on the user's computer, as a stand-alone software package, partly on the user's computer and partly on a remote computer or entirely on the remote computer or server.

[0068] As used herein, the terms “A or B” and “A and / or B” are intended to encompass A, B, or {A and B}. Further, the terms “A, B, or C” and “A, B, and / or C” are intended to encompass single items, pairs of items, or all items, that is, all of: A, B, C, {A and B}, {A and C}, {B and C}, and {A and B and C}. The term “or” as used herein means “and / or.”

[0069] As used herein, language such as “at least one of X, Y, and Z,” “at least one of X, Y, or Z,” “at least one or more of X, Y, and Z,” “at least one or more of X, Y, or Z,” “at least one or more of X, Y, and / or Z,” or “at least one of X, Y, and / or Z,” is intended to be inclusive of both a single item (e.g., just X, or just Y, or just Z) and multiple items (e.g., {X and Y}, {X and Z}, {Y and Z}, or {X, Y, and Z}). The phrase “at least one of” and similar phrases are not intended to convey a requirement that each possible item must be present, although each possible item may be present.

[0070] The techniques presented and claimed herein are referenced and applied to material objects and concrete examples of a practical nature thatdemonstrably improve the present technical field and, as such, are not abstract, intangible or purely theoretical. Further, if any claims appended to the end of this specification contain one or more elements designated as “means for [perform]ing [a function]...” or “step for [performing [a function]...”, it is intended that such elements are to be interpreted under 35 U.S.C. § 112(f). However, for any claims containing elements designated in any other manner, it is intended that such elements are not to be interpreted under 35 U.S.C. § 112(f).

[0071] While the invention has been described with reference to the exemplary examples thereof, those skilled in the art will be able to make various modifications to the described examples without departing from the true spirit and scope. The terms and descriptions used herein are set forth by way of illustration only and are not meant as limitations. In particular, although the method has been described by examples, the steps of the method can be performed in a different order than illustrated or simultaneously. Those skilled in the art will recognize that these and other variations are possible within the spirit and scope as defined in the following claims and their equivalents.

Claims

What is claimed is:1 . A method of using a digital pathology machine learning model implementation without requiring transmission of digital pathology images to a location of the digital pathology machine learning model implementation, the method comprising: providing, on a server computer, a digital pathology image embeddings application program interface (API), wherein the server computer is communicatively coupled to: a client computer, an embeddings server, an image retrieval server, and a digital pathology image server; receiving, at the digital pathology image embeddings API and from the client computer, an embeddings job request, wherein the embeddings job request comprises: (1 ) an identification of a set of at least one digital pathology image hosted by the digital pathology image server, (2) an image resolution instruction, and (3) an identification of an embeddings network; passing, by the server computer and to the embeddings server, (1 ) metadata characterizing the set of at least one digital pathology image resolved according to the resolution instruction and (2) an indication of the identification of the embeddings network; obtaining, by the server computer and from the embedding server, a set of at least one embeddings vector after the embeddings server transforms, by an embeddings network implementation corresponding to the identification of the embedding model, a resolved set of at least one digital pathology image into the set of at least one embeddings vector, wherein the resolved set of at least one digital pathology image corresponds to the set of at least one digital pathology image after having been resolved by the image retrieval server according to the resolution instruction; and transmitting, by the server computer, the set of at least one embeddings vector to a non-transitory embeddings storage location, for the client computer to use the digital pathology machine learning model implementation to process the set of at least one embeddings vector without the client computer transmitting or receiving the set of at least one digital pathology image.

2. The method of claim 1 , wherein the image resolution instruction is indicative of an image tiling, and wherein the set of at least one digital pathology image resolved according to the resolution instruction comprises a set of image tiles.

3. The method of claim 2, further comprising: transmitting, by the server computer and to the client computer, image tile origination information indicative of image locations from which image tiles in the set of image tiles were extracted.

4. The method of claim 1 , wherein the image resolution instruction is indicative of a thumbnail image, and wherein the set of at least one digital pathology image resolved according to the resolution instruction comprises at least one downsampled image.

5. The method of claim 1 , wherein the image resolution instruction is indicative of a size of a pathology sample portion represented by each pixel, and wherein the set of at least one digital pathology image resolved according to the resolution instruction comprises at least one digital pathology image with pixels each representing the size.

6. The method of claim 1 , wherein the image resolution instruction comprises an image segmentation instruction indicative of a request to segment, wherein the set of at least one digital pathology image resolved according to the resolution instruction comprises at least digital pathology image segmented into at least one region of interest.

7. The method of claim 6, wherein the set of at least one digital pathology image resolved according to the resolution instructions omits at least one region not of interest.

8. The method of claim 6, wherein the image resolution instruction comprises an annotation of at least one downsampled digital pathology image.

9. The method of claim 1 , wherein the non-transitory embeddings storage location is remote from the server computer.

10. The method of claim 1 , further comprising selecting the embeddings server according to proximity to the digital pathology image server.11 . The method of claim 1 , further comprising: receiving, by the server computer and from the client computer, a request to upload a custom embeddings network, wherein the identification of the embeddings network comprises an identification of the custom embeddings network, and wherein the embeddings network implementation comprises an implementation of the custom embeddings network.

12. The method of claim 1 , wherein the embedding job request further comprises an image augmentation instruction, and wherein the set of at least one digital pathology image resolved according to the resolution instruction comprises images augmented according to the image augmentation instruction.

13. The method of claim 12, wherein the image augmentation instruction comprises instructions for at least one of: rotation, color jitter, noise introduction, or stain augmentation.

14. The method of claim 1 , wherein the client computer trains the digital pathology machine learning model implementation on the set of embeddings.

15. The method of claim 1 , wherein the digital pathology machine learning model implementation performs inference on the set of embeddings.

16. The method of claim 1 , wherein the embedding server comprises the image retrieval server.

17. The method of claim 1 , wherein the digital pathology image server comprises the image retrieval server.

18. A system for using a digital pathology machine learning model implementation without requiring transmission of digital pathology images to a location of the digital pathology machine learning model implementation, the system comprising: a server computer hosting a digital pathology image embeddings application program interface (API), wherein the server computer is communicatively coupled to: a client computer, an embeddings server, an image retrieval server, and a digital pathology image server, wherein the server computer is configured to perform operations comprising: receiving, at the digital pathology image embeddings API and from the client computer, an embeddings job request, wherein the embeddings job request comprises: (1 ) an identification of a set of at least one digital pathology image hosted by the digital pathology image server, (2) an image resolution instruction, and (3) an identification of an embeddings network; passing, by the server computer and to the embeddings server, (1 ) metadata characterizing the set of at least one digital pathology image resolved according to the resolution instruction and (2) an indication of the identification of the embeddings network; obtaining, by the server computer and from the embedding server, a set of at least one embeddings vector after the embeddings server transforms, by an embeddings network implementation corresponding to the identification of the embedding model, a resolved set of at least one digital pathology image into the set of at least one embeddings vector, wherein the resolved set of at least one digital pathology image corresponds to the set of at least one digital pathology image after having been resolved by the image retrieval server according to the resolution instruction; andtransmitting, by the server computer, the set of at least one embeddings vector to a non-transitory embeddings storage location, for the client computer to use the digital pathology machine learning model implementation to process the set of at least one embeddings vector without the client computer transmitting or receiving the set of at least one digital pathology image.

19. The system of claim 18, wherein the image resolution instruction is indicative of an image tiling, and wherein the set of at least one digital pathology image resolved according to the resolution instruction comprises a set of image tiles.

20. The system of claim 19, wherein the operations further comprise: transmitting, by the server computer and to the client computer, image tile origination information indicative of image locations from which image tiles in the set of image tiles were extracted.21 . The system of claim 18, wherein the image resolution instruction is indicative of a thumbnail image, and wherein the set of at least one digital pathology image resolved according to the resolution instruction comprises at least one downsampled image.

22. The system of claim 18, wherein the image resolution instruction is indicative of a size of a pathology sample portion represented by each pixel, and wherein the set of at least one digital pathology image resolved according to the resolution instruction comprises at least one digital pathology image with pixels each representing the size.

23. The system of claim 18, wherein the image resolution instruction comprises an image segmentation instruction indicative of a request to segment,wherein the set of at least one digital pathology image resolved according to the resolution instruction comprises at least digital pathology image segmented into at least one region of interest.

24. The system of claim 23, wherein the set of at least one digital pathology image resolved according to the resolution instructions omits at least one region not of interest.

25. The system of claim 23, wherein the image resolution instruction comprises an annotation of at least one downsampled digital pathology image.

26. The system of claim 18, wherein the non-transitory embeddings storage location is remote from the server computer.

27. The system of claim 18, wherein the operations further comprise selecting the embeddings server according to proximity to the digital pathology image server.

28. The system of claim 18, wherein the operations further comprise: receiving, by the server computer and from the client computer, a request to upload a custom embeddings network, wherein the identification of the embeddings network comprises an identification of the custom embeddings network, and wherein the embeddings network implementation comprises an implementation of the custom embeddings network.

29. The system of claim 18, wherein the embedding job request further comprises an image augmentation instruction, and wherein the set of at least one digital pathology image resolved according to the resolution instruction comprises images augmented according to the image augmentation instruction.

30. The system of claim 29, wherein the image augmentation instruction comprises instructions for at least one of: rotation, color jitter, noise introduction, or stain augmentation.31 . The system of claim 18, wherein the operations further comprise training the digital pathology machine learning model implementation on the set of embeddings.

32. The system of claim 18, wherein the operations further comprise the digital pathology machine learning model implementation performing inference on the set of embeddings.

33. The system of claim 18, wherein the embedding server comprises the image retrieval server.

34. The system of claim 18, wherein the digital pathology image server comprises the image retrieval server.

Citation Information

Patent Citations

  • Ensemble of machine learning models for automatic scene change detection

    US11776273B1

  • Distributed privacy-preserving computing on protected data

    US20230080780A1

  • Attention-based multiple instance learning

    US20240079138A1

  • Machine learning models for automated request processing

    US20240120080A1